Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideolo…
When Writing Style Drifts: Benchmarking Authorship Verification under Distribution Shifts…
LinCa: Accelerating Diffusion Models via Learnable Decomposed Feature Caching
Understanding the Surprising Generalization Properties of Tabular Foundation Models
An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code W…
EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection
Adaptive Policy Portfolios for Robust Markov Decision Processes
Dynamic Compression in Recurrent Networks
A Residual Learning Approach for Unsteady Aerodynamic Load Prediction
CFB-GBM v2.0: An Augmented Longitudinal Dataset for Multi-Modal Glioblastoma Segmentation…
DistillPath: An Efficient 22M Distilled Pathology Encoder Approaching Large Foundation Mo…
Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical …
BayesPrompt: human readable prompts that make sense
Toward the Optimal Regret-Instability Trade-off in Multi-Armed Bandits
From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sect…
Spatially explicit feature importance for building height estimation using research-acces…
Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It
Scale Matters: Adaptive Granularity Selection for Cross-Species 3D Plant Organ Segmentati…
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific…
D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory
The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting
TINA+: Probing Residual Visual Knowledge in Unlearned Diffusion Models via Diffusion-Cons…
Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guaran…
Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals
Benchmarking Automated Security Patch Backporting: How Far Are We?
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for E…
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation