For the complete documentation index, see llms.txt. This page is also available as Markdown.

Limits and Quotas

Rate limits, GPU quotas, and quota increase guidance for Parasail production workloads.

Parasail applies request rate limits and GPU quotas to protect shared capacity and help customers avoid unexpected usage.

Rate limits

Product class
RPM
Token limit

Serverless - Free

5

Not currently enforced

Serverless - User

500

Not currently enforced

Dedicated Serverless

1000

Not currently enforced

Dedicated Serverless Pro

4000

Not currently enforced

Enterprise

Unlimited

Not currently enforced

Contact Parasail if your workload needs higher request limits.

GPU quota

New customer organizations start with a quota of 4 GPUs shared across Batch and Dedicated services. The quota applies collectively across GPU types.

Examples:

  • If current usage is 2 H100 GPUs and you submit a Batch job that requires 8 H100 GPUs, Parasail rejects the submission because it exceeds the 4 GPU quota.

  • If a Dedicated deployment can auto-scale from 1 to 6 GPUs and the deployment already uses 4 GPUs, the quota prevents scaling to a fifth GPU.

If the deployment page shows Insufficient quota, the organization has reached its GPU quota.

Request a quota increase

Use the Quota Increase Form to request more GPU capacity. Quota increases up to 8 GPUs are available without additional justification. Requests for more than 8 GPUs require a short explanation of the workload.

Next steps

Last updated