Zro
Private inference endpoint for coding agents — zero data retention, EU-hosted, open-weight models.
| What is it | Private inference endpoint for coding agents — zero data retention, EU-hosted, open-weight models. |
|---|---|
| Pricing | Paid — from $20/mo |
| Free tier | No |
| Platform | Web Application |
| API | Yes |
| Best for | running coding agents with privacy, deploying open-weight models for code generation |
| Domain registered | 2025 |
Data updated Aug. 1, 2026
What does Zro do?
Zro is a private inference endpoint built specifically for coding agents. It lets you run open-weight models like MiniMax M3 and GLM-5.2 on EU-hosted infrastructure, with a promise of zero data retention and no training on your prompts or completions. If you're building or using AI coding tools that need to keep code and queries private, Zro gives you a fast, compatible API that works with the agents you already use.
Under the hood, Zro uses MoonMath's HyperQuant compression and custom attention kernels to handle long-context, multi-turn coding sessions efficiently. It supports both OpenAI-compatible and Anthropic-compatible API shapes, so you can point existing clients at Zro's base URL without rewriting code. The @moonmath-ai/zro npm package makes it easy to launch supported tools like Claude Code, Codex CLI, Cursor, and Cline with a single command. Inference runs in Finland and France, and billing starts at $20/month for $60 of spend.
This tool is for developers who need privacy-compliant inference for coding agents — whether you're building an internal assistant, running automated code reviews, or experimenting with open models. It's also useful for teams in regulated industries who can't send code to US-based cloud providers. Zro keeps your data in the EU and never uses it for training, so you get the speed of a modern inference endpoint without the privacy trade-offs.
Key features
What makes it stand outWho is Zro for?
Who benefits most from this toolPricing
Pro
- ~300M tokens expected monthly usage
- MiniMax M3 and GLM-5.2
- OpenAI-compatible API
- Zro CLI launcher
- $0.02 web search
- Usage packs always available
- Zero request retention
- No training on requests
Max
Everything in Pro, plus:
- ~1.5B tokens expected monthly usage
Enterprise
- Team and organization workflows
- Shared API-key planning
- Custom usage plan
- Custom SLA
- Direct support channel
- Security and procurement support
Trust & presence
Gallery
Click any image to enlargeAlternatives in AI inference
AI inference platform that routes high-volume tasks to specialized, cost-efficient models instead of expensive frontier LLMs.
Privacy-first AI inference stack — run 45+ open source models with flat monthly pricing and zero data retention
Unified API that routes LLM requests to the cheapest provider by analyzing cache behavior and pricing in real time
Distributed cloud platform for deploying and scaling AI inference and compute globally
Organize images, convert annotation formats, preprocess, augment, share, and ship more. We eliminate the boilerplate code every computer vision team has to write.
Serverless GPU platform for deploying machine learning models in minutes, with auto-scaling and pay-per-use pricing.
Helping millions of developers easily build, test, manage, and scale applications of any size faster than ever before.
Run open-source AI models on your own GPUs or UK-hosted hardware, with one OpenAI-compatible API and zero data retention.
Similar tools
Build custom AI agents and apps from your own data — no coding required.
Enterprise platform to build, deploy, and govern custom AI agents securely on your company's data and tools.