Explorar

Noticias de IA

29336 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appeara…
arXiv cs.AI Research & Papers
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
arXiv cs.AI Research & Papers
Controlled Memory Interference in Continual LLM Agents
arXiv cs.AI Research & Papers
TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Sourc…
arXiv cs.AI Research & Papers
From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Or…
arXiv cs.AI Research & Papers
Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation…
arXiv cs.AI Research & Papers
Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD C…
arXiv cs.AI Research & Papers
An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadr…
arXiv cs.AI Research & Papers
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
arXiv cs.AI Research & Papers
Contextual Value Alignment via Multilayer Combinatorial Fusion
arXiv cs.AI Research & Papers
DarwinX: Evolving Agent Harnesses Through Natural Selection
arXiv cs.AI Research & Papers
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
arXiv cs.AI Research & Papers
Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Impr…
arXiv cs.AI Research & Papers
KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assi…
arXiv cs.AI Research & Papers
Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Sc…
arXiv cs.AI Research & Papers
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
arXiv cs.AI Research & Papers
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
arXiv cs.AI Research & Papers
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
arXiv cs.AI Research & Papers
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba…
arXiv cs.AI Research & Papers
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Lo…
arXiv cs.AI Research & Papers
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
arXiv cs.AI Research & Papers
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
arXiv cs.AI Research & Papers
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Gl…
arXiv cs.AI Research & Papers
AndroidReality: How Far Are Mobile Agents from the Real World?
arXiv cs.AI Research & Papers
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
arXiv cs.AI Research & Papers
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
arXiv cs.AI Research & Papers
Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven …
arXiv cs.AI Research & Papers
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Lang…
arXiv cs.AI Research & Papers
Probabilistic Circuits for Knowledge Graph Completion with Reduced Rule Sets
arXiv cs.AI Research & Papers
SiriusDeliver: Automating Data Warehouse Delivery at Tencent