SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance
Explorar
Noticias de IA
29627 elementos — filtrados, clasificados y sin duplicados
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Inte…
MixReasoning: Switching Modes to Think
Capacity, Not Format: Rethinking Structured Reasoning Failures
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
RunAgent SuperBrowser: A Theory of Autonomous Web Navigation Grounded in Human Browsing B…
Discovering heuristics in a complex SAT solver with large language models
Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolvi…
TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encode…
MASS: Deep Research for Social Sciences with Memory-Augmented Social Simulation
Modeling the Diachronic Evolution of Legal Norms: An LRMoo-Based, Component-Level, Event-…
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs Rul…
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and …
TQA-Bench: Evaluating LLMs for Multi-Table Question Answering
Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
PTL-Diffusion: Manifold-Aware Diffusion with Periodic Terminal Laws
Topological Neural Operators
Bandits for Efficient Experimentation: Adapting to Control Group, Preferences, and Contex…
Data Synthesis and Parameter-Efficient Fine-Tuning for Low-Resource NMT: A Case Study on …
MemToolAgent overview with a simple restaurant booking scenario where the agent retrieves…
Vision Language Model Helps Private Information De-Identification in Vision Data
The CIFAR Synthetic Evidence Corpus for Detecting AI-Generated Evidence
Difference-Aware Retrieval Policies for Imitation Learning
Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
Observability for Delegated Execution in Agentic AI Systems
MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation
Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents
End-to-End Context Compression at Scale