Featherless LLM
Serverless API access to 22,700+ open-source AI models for coding, writing, and research.
| What is it | Serverless API access to 22,700+ open-source AI models for coding, writing, and research. |
|---|---|
| Pricing | Paid — from $10/mo |
| Free tier | No |
| Platform | API |
| API | Yes |
| Works with | n8n |
| Best for | Testing and fine-tuning AI models, Building AI-powered applications |
| Domain registered | 2024 |
Data updated Aug. 1, 2026
What does Featherless LLM do?
Featherless LLM is a serverless inference platform that provides instant API access to a massive library of over 22,700 open-source AI models. Instead of managing your own GPU infrastructure, you can deploy and use models for coding assistance, creative writing, deep research, and other AI tasks through a simple API. The service handles all the complex backend operations, from model loading to GPU orchestration, so developers can focus on building applications rather than managing servers.
What makes Featherless LLM stand out is its extensive model catalog and straightforward pricing. Unlike providers that charge per token, it offers flat monthly subscriptions with unlimited usage. The platform supports popular model architectures like Llama, Mistral, Qwen, and DeepSeek, and maintains compatibility with OpenAI's SDK, making it easy to integrate into existing projects. It also emphasizes privacy by not logging any prompts or completions.
This service is particularly valuable for AI developers and teams who need to test multiple models or run production workloads without the overhead of infrastructure management. Real-world use cases include powering AI writing tools like NovelCrafter, enhancing coding assistants like OpenHands, and supporting chat applications like WyvernChat with a wide selection of specialized models for different tasks and personalities.
Key features
What makes it stand outWho is Featherless LLM for?
Who benefits most from this toolPricing
Feather Basic
- Access to models up to 15B
- Up to 2 concurrent connections
- Up to 16K context
Feather Premium
- Access to DeepSeek, Kimi-K2 and GLM 4.6
- Access any model - no limit on size!
- Up to 4 concurrent connections
- Up to 32K context
Feather Scale
- Business plan that can scale to arbitrarily many concurrent connections
- Private, secure, and anonymous usage - no logs
Trust & presence
Alternatives in AI inference
Serverless GPU platform for deploying machine learning models in minutes, with auto-scaling and pay-per-use pricing.
High-throughput LLM inference engine for fast, memory-efficient AI model serving.
A global GPU network for running AI models at scale — serverless, dedicated, or batch inference with pay-per-token pricing.
High-performance AI inference API — deploy any HuggingFace LLM 3-10x faster with an OpenAI-compatible endpoint.
Local AI runtime for text, image, and speech — run models on your own hardware, free and private
Privacy-first AI inference stack — run 45+ open source models with flat monthly pricing and zero data retention
Plug-and-play local AI server — run LLMs and image generation on your own hardware with full data privacy.
High-performance AI inference platform — deploy and scale open-source models like Llama and Gemma with a single API call.
Similar tools
Platform for sharing, discovering, and running machine learning models, datasets, and AI apps.
Device-native AI foundation models that run on phones, laptops, and cars — fine-tune and deploy locally.
🤖 • Run LLMs on your laptop, entirely offline 📚 • Chat with your local documents 👾 • Use models through the in-app Chat UI or an OpenAI compatible local server
Works with n8n
View all →AI platform offering ChatGPT for conversation, an API for developers, and business solutions — all powered by GPT models.
AI-powered translation tool delivering superior accuracy for text, documents, and real-time communication across 30+ languages.
AI research and product company building safe, capable assistants like Claude
AI-powered project management hub — manage tasks, documents, and collaboration with built-in AI assistants.
Online video editor with AI tools for subtitles, dubbing, avatars, and screen recording — all in your browser.
AI assistant that helps with writing, coding, analysis, and research — chat, generate content, or connect it to your tools
A suite of integrated development environments (IDEs) with built-in AI coding assistance for multiple programming languages.
AI-powered project management tool to organize tasks, track progress, and collaborate with teams using boards, lists, and cards.