For the complete documentation index, see llms.txt. This page is also available as Markdown.

Welcome

Parasail provides affordable, high-performance cloud GPUs for running demanding AI workloads—serverless inference, dedicated instances, and batch processing.

Parasail is a cloud GPU platform for AI inference. We aggregate top AI hardware providers to give you scalable, on-demand access to powerful computing resources—without costly hardware investments or lengthy contracts.

We offer three ways to run inference, each designed for a different workload pattern:

Not sure where to start?

Pick a product based on your workload:

If you...
Use
Why

Want to call a popular open-source model with no setup

Pay-per-token, instant, OpenAI-compatible

Need a specific or private model, predictable latency, or production SLAs

Private GPU endpoint with full control over model, hardware, and scaling

Process large volumes offline where real-time isn't required

50% off serverless pricing, up to millions of requests

Find your path

Choose the option that best matches what you're trying to do:

Quickstart

Make your first API call in under two minutes:

API reference

Parasail's API is fully OpenAI-compatible. Point your existing OpenAI SDK at our base URL and start making requests.

See the API Reference for full details on authentication, endpoints, and parameters.

For AI coding agents

Building on Parasail on behalf of a user? Read llms.txt for a flat index of every page, then see For AI Coding Agents for per-product setup notes. In short: use the OpenAI SDK against https://api.parasail.io/v1 with the PARASAIL_API_KEY environment variable—every endpoint is OpenAI-compatible.

Explore

Last updated