AgentAtlas: Beyond Outcome Leaderboards for LLM Agents
Explorar
Noticias de IA
21271 elementos — filtrados, clasificados y sin duplicados
ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality…
On the Push-Based Asynchronous Federated Learning: A Bias-Correction Aggregation Approach
CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intellige…
Beyond Questions: Evaluating What Large Language Models (Actually) Know
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
Yes, Q-learning Helps Offline In-Context RL
Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Dat…
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
Mechanized Foundations of Structural Governance: Machine-Checked Proofs for Governed Inte…
Algebraic Semantics of Governed Execution: Monoidal Categories, Effect Algebras, and Cote…
Many Logics, One Methodology: A Plea for Logical Pluralism in Formalised Reasoning (prepr…
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models
LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop P…
How Chain-of-Thought Works? Tracing Information Flow from Decoding, Projection, and Activ…
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting …
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
Timestep-Aware SVDQuant-GPTQ for W4A4 Quantization of Wan2.2-I2V
Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Des…
MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering via Miscoverage Ri…
Grounding Text Embeddings in Stakeholder Associations
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical No…
From Feasible to Practical: Pareto-Optimal Synthesis Planning
GraphMind: From Operational Traces to Self-Evolving Workflow Automation
Declarative Data Services: Structured Agentic Discovery for Composing Data Systems