Caveman
AI cost optimization stack that auto-caches, compresses, and routes LLM traffic to cut output tokens by ~65%
| What is it | AI cost optimization stack that auto-caches, compresses, and routes LLM traffic to cut output tokens by ~65% |
|---|---|
| Pricing | Freemium — from $29/mo |
| Free tier | Yes |
| Platform | API |
| API | Yes |
| Best for | cutting monthly LLM API costs in production, optimizing token use in AI coding agents |
| Domain registered | 2026 |
Data updated Aug. 15, 2026
What does Caveman do?
Caveman is an efficiency layer for teams that build on large language models. It sits between your application and the AI providers you use, watches the traffic, and applies cost-saving optimizations like caching, compression, and routing. The main promise is that it can cut roughly 65% of your AI costs by reducing the number of output tokens you pay for, then show you a breakdown of what was saved.
The tool is deliberately lightweight to install. There are four ways to run it: a Claude Code skill installed with a curl command, a global npm package called Caveman Code, a package called Cavemem, and a browser extension that works with ChatGPT, Claude, and Gemini. That range of entry points makes it useful inside AI-native coding agents, in backend services, and in everyday chat interfaces. The project reports about 97.9k GitHub stars and was #1 on Hacker News, which points to strong interest from developers.
Who benefits most? Engineering teams that pay significant LLM API bills, developers building AI agents, and companies that want to lower infrastructure costs without rewriting their code. Caveman is not an AI model itself; it is a cost-control layer for the AI tools you already use. The verified-savings reporting is especially useful for teams that need to justify a new tool to management.
Key features
What makes it stand outWho is Caveman for?
Who benefits most from this toolPricing
Free tier available — start without a credit cardFree
- Local wrap: recoverable compression on your machine (BYOK) — no account required
- MIT skill + extension for Claude Code (30+ agents)
- Inferred savings + Cave Score, computed locally
- Optional free account adds cloud sync and one dashboard seat
Indie
- Local wrap + hosted gateway (BYOK) — 1 seat
- Synced dashboard: inferred headroom + qualifying causal-cache ledger
- 50M optimized tokens / week
- Keep 100% of your savings — no gainshare
Team
- 10 seats included · $29 / extra seat
- Eval-gated rollout with automatic rollback
- Receipt export + Ed25519 verification (automatic signing disabled)
- Projects + metered $0.75 / M beyond plan
Enterprise
- Planned platform floor + gainshare on verified savings only
- OIDC SSO, audit log, RBAC, governance
- On-prem / BYOC deploy + OEM embedding
- Provider-invoice reconciliation (planned)
Trust & presence
Gallery
Click any image to enlargeAlternatives in Developer Tools
AI control plane that audits agent tool calls and caches LLM responses to cut costs by 50%
Drop-in proxy that monitors, optimizes, and protects your LLM spending across apps and coding agents
A macOS menu bar app that optimizes Claude Code and Codex prompts by compressing noisy output to reduce token costs by ~50%.
Persistent memory layer for AI coding agents — captures sessions, enables recall, and consolidates knowledge without external databases.
Run multiple AI coding agents in one shared repository and canvas — no merge conflicts.
AI gateway that routes, compresses, and governs requests to LLMs — cutting costs by 60–90%
Pre-built AI agent kits for Claude Code — specialized subagents automate coding, testing, and marketing workflows.
Count tokens and estimate API costs for 300+ LLM models including GPT-5, Claude, Gemini, and Llama