> For the complete documentation index, see [llms.txt](https://docs.parasail.io/parasail-docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.parasail.io/parasail-docs/readme.md).

# Welcome

Parasail provides affordable, high-performance AI inference with shared Serverless, Elastic Endpoints, Dedicated Instances, and Batch.

Parasail is a cloud GPU platform for AI inference. We aggregate top AI hardware providers to give you scalable, on-demand access to powerful computing resources—without costly hardware investments or lengthy contracts.

We offer four ways to run inference, each designed for a different workload pattern:

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Serverless</strong></td><td>Shared, on-demand, pay-per-token inference for popular open-source models. No setup required—just call the API.</td><td></td><td></td><td><a href="/parasail-docs/products/overview.md">Serverless</a></td></tr><tr><td><strong>Elastic Endpoints</strong></td><td>Dedicated GPU performance with per-token billing and no GPU-hour commitment. Available for selected models; contact sales to set up your endpoint.</td><td></td><td></td><td><a href="/parasail-docs/products/capacity.md">Elastic Endpoints</a></td></tr><tr><td><strong>Dedicated Instances</strong></td><td>Private GPU endpoints with full control over model, hardware, and scaling, billed per GPU-hour.</td><td></td><td></td><td><a href="/parasail-docs/products/overview-1.md">Dedicated Instances</a></td></tr><tr><td><strong>Batch Processing</strong></td><td>Process millions of inferences at 50% off serverless pricing. OpenAI-compatible batch API.</td><td></td><td></td><td><a href="/parasail-docs/products/quickstart.md">Batch</a></td></tr></tbody></table>

## Not sure where to start?

Pick a product based on your workload:

| If you...                                                                   | Use                                                              | Why                                                                                        |
| --------------------------------------------------------------------------- | ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| Want to call a popular open-source model with no setup                      | [**Serverless**](/parasail-docs/products/overview.md)            | Pay-per-token, instant, OpenAI-compatible                                                  |
| Want dedicated GPU performance with per-token billing for an eligible model | [**Elastic Endpoints**](/parasail-docs/products/capacity.md)     | Private, autoscaling capacity with no idle-GPU bill; contact sales to set up your endpoint |
| Need a specific or private model, predictable latency, or production SLAs   | [**Dedicated Instances**](/parasail-docs/products/overview-1.md) | Private GPU endpoint with full control over model, hardware, and scaling                   |
| Process large volumes offline where real-time isn't required                | [**Batch**](/parasail-docs/products/quickstart.md)               | 50% off serverless pricing, up to millions of requests                                     |

Elastic Endpoints have one starting pricing tier at 25% above the corresponding Parasail Serverless input and output token rates. Availability is limited to selected models and expands model by model. [Contact sales](https://www.parasail.io/contact) to confirm model eligibility and set up your endpoint. After setup, use your endpoint's status page for capacity and rate limits. See [Elastic Endpoints](/parasail-docs/products/capacity.md) for details.

## Find your path

Choose the option that best matches what you're trying to do:

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Chat and text generation</strong></td><td>Build chatbots, assistants, and text generation pipelines</td><td></td><td></td><td><a href="/parasail-docs/use-cases/chat-text-generation.md">Chat and Text Generation</a></td></tr><tr><td><strong>RAG and embeddings</strong></td><td>Build retrieval-augmented generation and vector search systems</td><td></td><td></td><td><a href="/parasail-docs/use-cases/rag-embeddings.md">RAG and Embeddings</a></td></tr><tr><td><strong>Batch at scale</strong></td><td>Process large volumes of prompts, images, or text offline</td><td></td><td></td><td><a href="/parasail-docs/use-cases/batch-processing.md">Batch Processing at Scale</a></td></tr><tr><td><strong>Agents and tool calling</strong></td><td>Build agentic workflows with function calling and multi-step reasoning</td><td></td><td></td><td><a href="/parasail-docs/use-cases/agents-tool-calling.md">Agents and Tool Calling</a></td></tr></tbody></table>

## Quickstart

Make your first API call in under two minutes:

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Serverless quickstart</strong></td><td>Call a serverless model with the OpenAI SDK</td><td></td><td></td><td><a href="/parasail-docs/quickstart/serverless.md">Serverless</a></td></tr><tr><td><strong>Dedicated quickstart</strong></td><td>Deploy your own model on a dedicated GPU</td><td></td><td></td><td><a href="/parasail-docs/quickstart/dedicated.md">Dedicated Instances</a></td></tr><tr><td><strong>Batch quickstart</strong></td><td>Submit a batch job in five lines of Python</td><td></td><td></td><td><a href="/parasail-docs/quickstart/batch.md">Batch Processing</a></td></tr></tbody></table>

## API reference

Parasail's API is fully OpenAI-compatible. Point your existing OpenAI SDK at our base URL and start making requests.

* **Base URL**: `https://api.parasail.io/v1`
* **Auth**: `Authorization: Bearer <PARASAIL_API_KEY>`
* **Get an API key**: <https://www.saas.parasail.io/keys>

See the [API Reference](/parasail-docs/api-reference/authentication.md) for full details on authentication, endpoints, and parameters.

## For AI coding agents

Building on Parasail on behalf of a user? Read [`llms.txt`](https://github.com/parasail-ai/gitbook-doc/tree/main/llms.txt) for a flat index of every page, then see [For AI Coding Agents](https://github.com/parasail-ai/gitbook-doc/tree/main/resources/for-ai-coding-agents.md) for per-product setup notes. In short: use the OpenAI SDK against `https://api.parasail.io/v1` with the `PARASAIL_API_KEY` environment variable—every endpoint is OpenAI-compatible.

## Explore

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-cover data-type="files"></th><th data-hidden></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>Guides</strong></td><td>Step-by-step tutorials for common tasks</td><td></td><td></td><td><a href="/parasail-docs/guides/model-recommendations.md">Model Selection</a></td></tr><tr><td><strong>Billing</strong></td><td>Pricing and billing details</td><td></td><td></td><td><a href="/parasail-docs/billing/pricing.md">Pricing</a></td></tr><tr><td><strong>Operations</strong></td><td>Production readiness, retries, limits, and quotas</td><td></td><td></td><td><a href="/parasail-docs/operate-in-production/overview.md">Overview</a></td></tr><tr><td><strong>Security</strong></td><td>Security, privacy, Trust Center, and account controls</td><td></td><td></td><td><a href="/parasail-docs/security-and-account-management/overview.md">Security Overview</a></td></tr></tbody></table>
