Chutes
Serverless AI compute platform for running open-source LLMs, image, video, and audio models at scale
| What is it | Serverless AI compute platform for running open-source LLMs, image, video, and audio models at scale |
|---|---|
| Pricing | Paid — from $10/mo |
| Free tier | No |
| Platform | Web Application |
| API | Yes |
| Works with | n8n |
| Best for | deploying and scaling open-source LLMs for production, running image, video, and speech generation models |
| Domain registered | 2024 |
Data updated Aug. 1, 2026
What does Chutes do?
Chutes is a serverless compute platform built for deploying and scaling open-source AI models. It handles everything from large language models to image generation, video processing, speech synthesis, and music creation. The service is designed for developers and teams who need reliable, on-demand AI inference without managing their own infrastructure. Chutes runs models in ephemeral, hot instances that stay ready to scale, so cold starts aren't a problem.
The platform works by letting you choose from hundreds of pre-deployed models or bring your own code. You get per-token or per-hour pricing depending on the model type. Chutes also offers TEE (Trusted Execution Environment) compute for secure, isolated jobs. A unique selling point is how quickly new open-source models appear here — often within minutes of release. There's a chat interface and a search tool for end users, but the core value is for developers using the Python SDK or API.
Chutes is best for AI developers, researchers, and product teams who want to run open-source models without provisioning GPUs or managing Kubernetes. If you're building an AI-powered app and need scalable inference for LLMs, image generation, or content moderation, this is a solid option. The pay-as-you-go model and monthly subscription tiers make it flexible for different usage levels.
Key features
What makes it stand outWho is Chutes for?
Who benefits most from this toolPricing
Public Inference
- 0.0245 - 1.40 input per 1M tokens
- 0.0978 - 4.40 output per 1M tokens
- Per-token pricing
- No minimum
- No markup
Private Chute
- 1.8 hourly rate
- Per-second billing
- One-time deployment fee
- Confidential GPU
Plus
- unknown daily quota
- 6% discount on payg
- Bundled daily quota
- 6% off PAYG rates beyond the quota
Pro
- unknown daily quota
- 10% discount on payg
- Larger daily quota
- 10% off PAYG rates beyond the quota
Enterprise
- Volume discounts
- Custom rate limits
- Dedicated support
Trust & presence
Gallery
Click any image to enlargeAlternatives in AI inference
Serverless GPU platform for deploying machine learning models in minutes, with auto-scaling and pay-per-use pricing.
AI infrastructure platform — deploy, serve, and scale machine learning models in production with optimized inference
Renewable-powered cloud infrastructure and managed inference service for running large AI models.
Self-hosted AI inference engine for search and document processing — deploy models on your own cloud infrastructure.
High-speed AI inference API powered by purpose-built hardware, not repurposed GPUs.
A global GPU network for running AI models at scale — serverless, dedicated, or batch inference with pay-per-token pricing.
Rent high-performance GPUs on demand for AI, machine learning, and graphics rendering at significantly lower costs.
High-performance AI inference platform — deploy and scale open-source models like Llama and Gemma with a single API call.
Works with n8n
View all →AI platform offering ChatGPT for conversation, an API for developers, and business solutions — all powered by GPT models.
AI-powered translation tool delivering superior accuracy for text, documents, and real-time communication across 30+ languages.
AI research and product company building safe, capable assistants like Claude
AI-powered project management hub — manage tasks, documents, and collaboration with built-in AI assistants.
Online video editor with AI tools for subtitles, dubbing, avatars, and screen recording — all in your browser.
AI assistant that helps with writing, coding, analysis, and research — chat, generate content, or connect it to your tools
A suite of integrated development environments (IDEs) with built-in AI coding assistance for multiple programming languages.
AI-powered project management tool to organize tasks, track progress, and collaborate with teams using boards, lists, and cards.