OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for …
MemArbiter: Decision-Time Memory Arbitration for Long-Horizon LLM Agents
Beyond Solution-Centric Search: Adaptive Inquiry and Knowledge Revision for Autonomous ML…
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning
Auditing Data Provenance in LLM Fine-tuning via Intrinsic Distributional Fingerprints
PAC Approximation and DIRECT Optimization for Parametric Markov Models
MEGRAG: Multi-Granular Evidence Graphs for Answer-Aware Multi-Hop RAG
Homebot: A Personal AI Agent for Conversational Home Assistance and Automation
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning
KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement
Faster-WAM: Do World Action Models Need Deep Action Modules?
Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcem…
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on …
Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+
Real-Time Detection and Repair of LLM Agent Failures
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum O…
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientifi…
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluatio…
RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Re…
AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategi…
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelli…
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs