Invocation-Level Reliability of Tool-Using Agents
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site …
Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Sear…
Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach f…
Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abdu…
A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Repres…
A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assist…
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Pr…
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
Toward a New Science of AI as Cognitive Infrastructure
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
Please stop flooding our projects with AI slop to furnish your CV
Lightspeed Sees Globally Competitive Indian AI Model by 2027
Anthropic was illegally blacklisted by the Trump administration, court rules
A Judge Has Blocked the Pentagon’s Attempt to Blacklist Anthropic
Yotta Plans IPO ‘Very Soon’ to Keep Up With AI Demand, CEO Says
Yotta Data Chairman on IPO Pipeline, Growth
Lightspeed's Mohapatra on India's AI Startups
Supporting Thailand’s next generation of AI startups
Anthropic Wins Court Challenge to US Supply-Chain Risk Label
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
The Open ASR Leaderboard Adds Its First Global South Language
Gemini Omni 1.1 🎬, Cohere Parse 📄, Codex persistent mode 👨💻
LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Pro…
Agent Seer: Synthesizing Scenarios from Specification Understanding
Build agentic creative workflows with Amazon Quick and fal
Can Martha Stewart convince you Waymo is a good thing?
Anthropic's new hardware standard lets AI agents control the physical world