VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
A game theory for foundation models shows new paths to rational cooperation through simil…
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reaso…
SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Informatio…
Interpretable Adaptive Sampling for LLM Test-Time Scaling
TraceCAD: Trace-Guided Repair for Agentic CAD Generation
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for…
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agen…
MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows
CUADebug: Diagnosing and Repairing Computer-Use Agent Failures
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Expli…
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
Traceable Multi-Agent System for Knowledge-Based Forecasting
MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
Social Pressure Breaks Majority Voting in LLM Safety Panels
muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Ha…
Training-Free Hashing-Based Attention via Binary Principal Components
Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Perform…
EdgeLM: Edge Demonstrations for Language Models' Table Understanding
Image Classification Using CNN-QNN Hybrid Model with Optimized Correlated Features
The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Sc…
HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework …
Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning
Right Reset: Chunking by Prefix Removal