ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workpl…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Learning from the Test: Self-Referential Differential Testing for Deep RL Agents
Proxy reliance in large language model decisions is uncalibrated to predictive evidence
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents
Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Eva…
FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts
Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-…
Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays
Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization
On the Role of Citations in Preference Data
When Does AI for PDEs Yield Scientific Evidence?
More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language…
SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG
Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal S…
Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning
Evaluating Multimodal Narrative Understanding of Popular Hollywood Films
ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topolog…
Performance of a domain-specific large language model in answering patient questions in p…
MegaMem: A Retrieval Solution for Ultra-Large Context Windows
Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reco…
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Effici…
From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers
Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators…
StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models
SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Qu…
Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danm…
Towards a Densing Law for User Representation Learning at Billion-Scale Capacity
From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-T…
Radial Compensation: The Inverse Base-Distribution Problem for Chart-Based Generative Mod…