Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based…
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an E…
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language…
LLMs Get Smarter from Targeted Synthetic Multilingual Data
A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network E…
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Docume…
Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated …
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Doc…
Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma …
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
JarvisBench: Always-on Intelligence Between Humans and Agents
Position: AI Lock-In Is in Progress, and We Must Be Prepared
Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Ne…
PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal…
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integrat…
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Mo…
AgentMV: A State-Guided Multi-Agent Framework for Budget-Aware Music Video Generation
Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Hi…
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-…