Doubleword

Cost-efficient AI inference platform — run open-weight models via async, batch, or realtime APIs

Visit Website
doubleword.ai
Verified API available
Quick facts
What is it Cost-efficient AI inference platform — run open-weight models via async, batch, or realtime APIs
Pricing Paid
Free tier No
Platform Web Application
API Yes
Best for running large-scale batch inference for data processing, powering AI agents and long-horizon reasoning tasks
Domain registered 2024

Data updated Sept. 19, 2026

What does Doubleword do?

Doubleword is an AI inference platform that gives developers access to multiple open-weight models at a fraction of the cost of closed-source APIs. You can run models like Kimi-K3, DeepSeek-V4, and Qwen through three pricing tiers: realtime for instant responses, async for background tasks that complete in minutes to hours, and batch for jobs that can wait up to 24 hours. The platform claims up to 90% lower costs compared to providers like OpenAI or Anthropic, with average cache hit rates of 93% to reduce redundant compute. The headline comparison shows Doubleword's Kimi-K3 model scoring 59.7 on an intelligence benchmark while costing $18,000 per billion tokens, against GPT-5.6 Sol max at $24,000 and same score, or Claude Fable 5 at $60,000 for a slightly higher score.

The service works through a standard OpenAI-compatible API. You point your existing code at Doubleword's endpoint, pick a model, and optionally set the service tier to 'flex' for async or batch processing. The platform handles queuing, caching, and scaling behind the scenes. What makes Doubleword interesting is its focus on the 'efficient frontier' — the best cost-to-intelligence ratio at each performance level. Seven of the nine models on that frontier run on Doubleword, according to their data from Artificial Analysis. This is not a flashy consumer tool; it's a backend service for teams that need a lot of tokens processed cheaply.

The target audience is engineering and data science teams building AI-powered products at scale. Common use cases include running thousands of agentic evaluations, generating training datasets, annotating large image collections (as OpenMed did with 119K medical images), and powering high-throughput extraction or ETL pipelines. If your project needs frontier-level intelligence but the cost of closed APIs is eating your budget, Doubleword is worth a look. It is straightforward infrastructure — no magic, just cheaper inference for open models.

Key features

What makes it stand out
01
Up to 90% lower cost than closed-source APIs for large-scale inference
02
Three latency tiers: realtime, async (minutes to hours), and 24-hour batch
03
Average cache hit rates of 93% across model calls
04
Access to multiple frontier-level open-weight models like Kimi-K3, DeepSeek-V4, Qwen, and more
05
Clean API with a unified interface for chat completions and agentic workloads

Who is Doubleword for?

Who benefits most from this tool
running large-scale batch inference for data processing
powering AI agents and long-horizon reasoning tasks
reducing costs for high-volume model calls in production

Pricing

Inference API

Custom
  • Input: $0.03-$3.00, Cache Read: $0.01-$0.30, Output: $0.00-$15.00 (varies by model and priority) pricing per 1M tokens
  • Usage-based pricing per 1M tokens
  • Three priority tiers: Realtime, Async, Batch (24h)

Trust & presence

Domain Domain registered 2024

Gallery

Click any image to enlarge

Alternatives in AI inference

ZeroGPU Verified AI inference

AI inference platform that routes high-volume tasks to specialized, cost-efficient models instead of expensive frontier LLMs.

fireworks.ai Verified AI inference

High-performance AI inference platform — deploy and scale open-source models like Llama and Gemma with a single API call.

Entrim AI Verified AI inference

API for running open-source LLMs — up to 80% cheaper than competitors, with high throughput and privacy-first handling.

General Compute Verified AI inference

High-speed AI inference API powered by purpose-built hardware, not repurposed GPUs.

SiliconFlow Verified AI inference

AI model inference platform — access multiple LLMs and multimodal models through a single API with predictable pricing

n8n
Run BiOS Verified AI inference

Enterprise AI inference platform — access top models via one API, with zero data retention and pay-per-token pricing

Baseten Verified AI inference

AI infrastructure platform — deploy, serve, and scale machine learning models in production with optimized inference

n8n Top 100k site
Groq Verified AI inference

High-speed, low-cost AI inference API for running large language models with minimal latency.

make · n8n+1 Top 100k site

Similar tools

Deep Infra Verified AI API

Access hundreds of AI models through a single API — text, image, video, and speech generation with pay-per-use pricing.

Top 100k site
Metatext Verified AI API

AI gateway that routes coding agent requests to the cheapest suitable model, cutting API costs by ~40%

n8n
OpenRouter Verified AI API

Single API to access 300+ AI models from 60+ providers with optimized pricing, uptime, and performance

zapier · n8n+1 Top 100k site
Experiential Labs Verified AI API

Open source AI gateway — access hundreds of models from one API endpoint at cost price

Share X LinkedIn Telegram
Doubleword Visit