Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Analogy as Nonparametric Bayesian Inference over Relational Systems
Learning When to Trust via Selective Context Preference Optimization
An Optimal Agnostic PAC Algorithm
Text Steganography with Dynamic Codebook and Multimodal Large Language Model
Ge$^\text{2}$mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking…
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
Does FLAIR super-resolution erase or hallucinate small white-matter lesions?
BaKron: Efficient Quantization with Kronecker-Factored Hessians
Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for …
Depth-Guided Video Object Counting in Crowded Scenes
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implication…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New S…
Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architect…
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Si…
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Tr…
Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Hum…
GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal C…
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Age…
The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
Hierarchical Latent Prediction for Language Models
A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target D…
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
Multivariate Time Series Forecasting needs Cross Variable Loss