UIC-AIHealth4All at ArchEHR-QA 2026: Answer-First Evidence Grounding for Clinical Questio…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Expert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic …
Optimal Adversarial Testing: Extracting Honest Test Results from Dishonest Test Takers
VISTA: Verifier-Informed Student-to-Teacher Adaptation for On-Policy Self-Distillation
Multimodal Collaborative Debate for Zero-Shot Time Series Reasoning
Talk in Pieces, See in Whole: Disentangled and Hierarchical Representation Learning in La…
Learning a Size-Weight Frontier for Synthetic-Augmented Inference
Blog: Survey of Optimizers
Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss
Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failu…
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situatio…
A Probabilistic Interpretation of KV Cache Eviction
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
A comprehensive and trustworthy benchmark of AI methods for change detection in Earth obs…
Video Generative Models as Geometry Learner
Conformal Uncertainty Quantification Guarantees for Neural Operators
On the Maintenance and Co-evolution of Agent Plugins: An Empirical Study of Claude Code P…
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
Conformal Risk-Averse Decision Making with Optimized Certainty Equivalent Risk Control
When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded…
How Proper Scoring Rules Shape LLM Forecasting
SEGRA: A Structured Experience Guided Reasoning Agent for Property Graph Question Answeri…
Understanding and Enforcing Weight Disentanglement in Task Arithmetic
CheXtriev: Anatomy-Centered Representation for Case-Based Retrieval of Chest Radiographs
Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Register…
PanelShield: Verifiable Closed-Loop Safe Planning for Robotic Industrial Panel Operation
When Linguistic and Internal Confidence Diverge in Large Language Models
VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learn…
Embedding Models for Stance-Aware Argument Retrieval