The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark
ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity S…
A Responsible Artificial Intelligence Framework for Groundwater Modeling
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Si…
Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Ins…
Small Models Scout Bottleneck Order for Large-Model Data Control
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physi…
Synchronized Logit Steering: Real-world Steganography
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Questio…
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
Incoherent by Design? On the Moral Self-Consistency of LLMs
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assis…
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
Position: AI Lock-In Is in Progress, and We Must Be Prepared
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integrat…
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detect…
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an E…
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Doc…
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving