See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appeara…
Explorar
Noticias de IA
29336 elementos — filtrados, clasificados y sin duplicados
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
Controlled Memory Interference in Continual LLM Agents
TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Sourc…
From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Or…
Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation…
Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD C…
An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadr…
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
Contextual Value Alignment via Multilayer Combinatorial Fusion
DarwinX: Evolving Agent Harnesses Through Natural Selection
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Impr…
KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assi…
Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Sc…
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba…
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Lo…
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Gl…
AndroidReality: How Far Are Mobile Agents from the Real World?
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven …
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Lang…
Probabilistic Circuits for Knowledge Graph Completion with Reduced Rule Sets
SiriusDeliver: Automating Data Warehouse Delivery at Tencent