Caveman
AI cost optimization stack that auto-caches, compresses, and routes LLM traffic to cut output tokens by ~65%
| Qué es | AI cost optimization stack that auto-caches, compresses, and routes LLM traffic to cut output tokens by ~65% |
|---|---|
| Precios | Freemium — de $29/mo |
| Plan gratuito | Sí |
| Plataforma | API |
| API | Sí |
| Ideal para | cutting monthly LLM API costs in production, optimizing token use in AI coding agents |
| Dominio registrado | 2026 |
Datos actualizados 15 de agosto de 2026
¿Qué hace Caveman?
Caveman is an efficiency layer for teams that build on large language models. It sits between your application and the AI providers you use, watches the traffic, and applies cost-saving optimizations like caching, compression, and routing. The main promise is that it can cut roughly 65% of your AI costs by reducing the number of output tokens you pay for, then show you a breakdown of what was saved.
The tool is deliberately lightweight to install. There are four ways to run it: a Claude Code skill installed with a curl command, a global npm package called Caveman Code, a package called Cavemem, and a browser extension that works with ChatGPT, Claude, and Gemini. That range of entry points makes it useful inside AI-native coding agents, in backend services, and in everyday chat interfaces. The project reports about 97.9k GitHub stars and was #1 on Hacker News, which points to strong interest from developers.
Who benefits most? Engineering teams that pay significant LLM API bills, developers building AI agents, and companies that want to lower infrastructure costs without rewriting their code. Caveman is not an AI model itself; it is a cost-control layer for the AI tools you already use. The verified-savings reporting is especially useful for teams that need to justify a new tool to management.
Características principales
Qué la hace destacar¿Para quién es Caveman?
Quién saca más provecho de esta herramientaPrecios
Plan gratuito disponible — empieza sin tarjeta de créditoFree
- Local wrap: recoverable compression on your machine (BYOK) — no account required
- MIT skill + extension for Claude Code (30+ agents)
- Inferred savings + Cave Score, computed locally
- Optional free account adds cloud sync and one dashboard seat
Indie
- Local wrap + hosted gateway (BYOK) — 1 seat
- Synced dashboard: inferred headroom + qualifying causal-cache ledger
- 50M optimized tokens / week
- Keep 100% of your savings — no gainshare
Team
- 10 seats included · $29 / extra seat
- Eval-gated rollout with automatic rollback
- Receipt export + Ed25519 verification (automatic signing disabled)
- Projects + metered $0.75 / M beyond plan
Enterprise
- Planned platform floor + gainshare on verified savings only
- OIDC SSO, audit log, RBAC, governance
- On-prem / BYOC deploy + OEM embedding
- Provider-invoice reconciliation (planned)
Confianza y presencia
Galería
Haz clic en cualquier imagen para ampliarlaAlternativas en Herramientas para desarrolladores
Plano de control de IA que audita las llamadas a herramientas de agentes y almacena respuestas de LLM en caché para reducir costes en un 50 %
Proxy integrable que monitorea, optimiza y protege el gasto de tus LLM en aplicaciones y agentes de programación
Una app para la barra de menús de macOS que optimiza los prompts de Claude Code y Codex comprimiendo el ruido para reducir los costes de tokens en ~50%.
Capa de memoria persistente para agentes de programación de IA: captura sesiones, permite el recuerdo y consolida conocimientos sin bases de datos externas.
Ejecuta varios agentes de programación de IA en un mismo repositorio y lienzo compartido, sin conflictos de fusión.
AI gateway that routes, compresses, and governs requests to LLMs — cutting costs by 60–90%
Kits de agentes de IA preconfigurados para Claude Code: subagentes especializados que automatizan flujos de trabajo de programación, pruebas y marketing.
Cuenta tokens y estima costes de API para más de 300 modelos de IA, incluidos GPT-5, Claude, Gemini y Llama