Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
FedRAW: Preserving Rare-Label Influence in Asynchronous Federated Learning
Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval
Generator-Independent Runtime Assurance under Partial Observation
Agentic Pressure: The Endogenous Entropy of Reliable Autonomy
DAREBench: Deployment-Aware and Reliable Evaluation of Models as Agents
Scratchy: Visual-Scratchpad Multimodal Reasoning for Cryptographic Proof Generation in Ea…
Generating Instance Generators in PDDL Planning
AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories
When and Why LLM Causal Priors Help: Closed-Loop Prior Selection for Amortized Causal Inf…
Modus Tollens and Counterfactuals and Counterfactual Reasoning Based on Three Types of Ne…
Exposing Weaknesses in Emotion Recognition in Conversations
SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores
Distilling Vision-Language Models for On-Device Fire Understanding
More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review
From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational As…
Beyond Final Decisions: A Process-Centric Benchmark for Transparent AI-Assisted Peer Revi…
The Normalization of Deviance in AI Development
Learning Counterfactual World Models for Embodied Reasoning under Partial Observability
Predicting Wind Turbine Power Using Machine Learning and Weather Forecasts
VST: Verifiable Structured Transport for Auditable Agent-to-Agent Alpha Discovery
Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval
The convergent laboratory: when AI reasoning, autonomous experiments, high performance an…
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record f…
CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning
Recovering Temporal and Geographic Signals from Language Model Embeddings
CIVI: A Framework for Diagnosing Search Agent Failures in Civic Information
The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies
Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in th…
Deep belief networks are exact