Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strate…
Explorar
Noticias de IA
21272 elementos — filtrados, clasificados y sin duplicados
GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human
Relational In-Context Learning via Synthetic Pre-training with Structural Prior
Post-Training Language Models for Crosslingual Consistency
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
Hallucination Detection-Guided Preference Optimization for Clinical Summarization
Steering at the Source: Style Modulation Heads for Robust Persona Control
First head-to-head comparison of agentic AI applied to the analysis of simulated data of …
CosmicFish-HRM: Adaptive Reasoning via Hierarchical Recurrent Mechanisms in Compact Langu…
AuthorMix: Modular Authorship Style Transfer via Layer-wise Adapter Mixing
LoRe: Adaptive Interaction-Evaluation Routing with Per-Step Interaction Budgets for Itera…
SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
The Cognitive Categorical Transformer: Category-Theoretic Inductive Biases for Language M…
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
Rubric-Guided Process Reward for Stepwise Model Routing
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reaso…
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Mode…
Causal Disentanglement-Inspired Degradation Representation Learning for Full-Reference Im…
FHRFormer: A Self-Supervised Masked Transformer Framework for Fetal Heart Rate Time-Serie…
Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems
Archon: A Unified Multimodal Model for Holistic Digital Human Generation
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
Small Agent Group is the Future of Digital Health
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
Intent-aligned Autonomous Spacecraft Guidance via Reasoning Models
Rel-MOSS: Towards Imbalanced Relational Deep Learning on Relational Databases
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating