Selene 1
Atla is the only eval tool that helps you automatically discover the underlying issues in your AI agents. Understand step-level errors, prioritize recurring failure patterns, and fix issues fast–before your users ever notice.
| What is it | Atla is the only eval tool that helps you automatically discover the underlying issues in your AI agents. Understand step-level errors, prioritize recurring failure patterns, and fix issues fast–before your users ever notice. |
|---|---|
| Pricing | Freemium — from $199/mo |
| Free tier | Yes |
| Platform | API |
| API | Yes |
| Best for | testing AI response quality, evaluating AI agent performance |
| Domain registered | 2023 |
Data updated Aug. 1, 2026
What does Selene 1 do?
Selene 1 is a specialized AI model designed specifically for evaluating other AI systems. It acts as an automated judge, providing precise scores and detailed critiques on the performance of your AI applications, particularly large language models and AI agents. The tool analyzes responses based on criteria like accuracy, relevance, and safety, giving developers concrete feedback on where their AI might be failing or excelling.
The system comes in two versions: Selene 1 for comprehensive pre-production evaluations with industry-leading accuracy, and Selene 1 Mini optimized for speed and real-time inference evaluations. Both models are available through Hugging Face Transformers, Ollama, and GitHub, making them accessible to developers familiar with these platforms. What sets Selene apart is its focus specifically on evaluation tasks rather than general conversation, trained to provide consistent, reliable judgments that beat many frontier models from leading AI labs.
This tool is particularly valuable for AI developers, research teams, and companies building AI products who need to validate their system's reliability before deployment. Real-world use cases include testing chatbot responses for accuracy, evaluating AI agent performance in complex workflows, and building trust with customers by demonstrating rigorously tested AI capabilities. It's essentially a quality assurance specialist for your AI systems.
Key features
What makes it stand outWho is Selene 1 for?
Who benefits most from this toolPricing
Free tier available — start without a credit cardDeveloper Tier
- 2,000 traces
- Agent LLM-as-a-Judge inference
- Error patterns and improvement suggestions automatically generated
- Up to 3 custom LLM judge metrics
Startup Tier
Everything in Developer Tier, plus:
- 10,000 traces
- 60 data retention days
- 10 custom llm judge metrics
- Custom onboarding + dedicated Slack support
- Unlimited users
- Up to 10 custom LLM judge metrics
- 60-day data retention
- SOC2 report, BAA available (HIPAA)
Custom Tier
Everything in Startup Tier, plus:
- Everything included
- Self-hosted deployment
- Unlimited project workspaces
- Custom rate limits + SLA
- Custom data retention period
- Custom SSO and RBAC
- Access to deployed engineering team
- Team trainings & architectural guidance
Trust & presence
Gallery
Click any image to enlargeAlternatives in Testing
World’s first comprehensive evaluation, observability and optimization platform to help enterprises achieve 99% accuracy in AI applications across software and hardware.
Platform for evaluating and optimizing RAG pipelines and generative AI applications with automated and human-in-the-loop testing.
AI testing platform — evaluate, debug, and monitor AI agents with automated testing and guardrails
Open-source testing platform for AI agents. Run simulations, catch regressions, and ship autonomous agents with confidence. Built for developers who treat AI like software. Agent simulations are the new unit tests
AI agent monitoring platform — detects hallucinations and errors in real-time, enables human-in-the-loop remediation
Automated testing API for AI-generated answers — verify accuracy, completeness, and correctness against a source of truth.
AI agent testing platform — simulate thousands of conversations to find and fix issues before deploying to customers.
Replay and debug AI agent failures by forking the exact step that broke, then prove your fix before shipping
Similar tools
Agentic AI platform that automates IT and HR support, answering questions and resolving tickets across Teams, Slack, and email.
AI engineering agents automate coding, research, and review tasks so developers focus on high-level decisions.
AI service desk for IT, HR, and Operations teams — automates workflows and resolves employee requests directly inside Slack.
Enterprise AI platform that automates back-office workflows in HR, IT, Finance, and Procurement with pre-built AI agents.