Welcome
Parasail provides affordable, high-performance cloud GPUs for running demanding AI workloads—serverless inference, dedicated instances, and batch processing.
Parasail is a cloud GPU platform for AI inference. We aggregate top AI hardware providers to give you scalable, on-demand access to powerful computing resources—without costly hardware investments or lengthy contracts.
We offer three ways to run inference, each designed for a different workload pattern:
Not sure where to start?
Pick a product based on your workload:
Want to call a popular open-source model with no setup
Pay-per-token, instant, OpenAI-compatible
Need a specific or private model, predictable latency, or production SLAs
Private GPU endpoint with full control over model, hardware, and scaling
Process large volumes offline where real-time isn't required
50% off serverless pricing, up to millions of requests
Find your path
Choose the option that best matches what you're trying to do:
Quickstart
Make your first API call in under two minutes:
API reference
Parasail's API is fully OpenAI-compatible. Point your existing OpenAI SDK at our base URL and start making requests.
Base URL:
https://api.parasail.io/v1Auth:
Authorization: Bearer <PARASAIL_API_KEY>Get an API key: https://www.saas.parasail.io/keys
See the API Reference for full details on authentication, endpoints, and parameters.
For AI coding agents
Building on Parasail on behalf of a user? Read llms.txt for a flat index of every page, then see For AI Coding Agents for per-product setup notes. In short: use the OpenAI SDK against https://api.parasail.io/v1 with the PARASAIL_API_KEY environment variable—every endpoint is OpenAI-compatible.
Explore
Last updated