Length-Adaptive Decoding for Masked Diffusion Machine Translation
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning
Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Datas…
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quant…
On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study
Composable Trust Infrastructure for Manufacturing Knowledge Graphs: Cross-System Provenan…
There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile I…
Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritag…
RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Si…
What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Th…
OVIBench: Benchmarking Online Video Question Answering under Interruption
Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under n…
One-Step Evolution for Long-Time Extrapolation: An Error-Bound-Informed and Prior-Guided …
WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and H…
FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows
The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning
SkillAlchemy: Open-World Agent Skill Creation
Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordi…
Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environm…
Future Querying: Can LLMs Serve as Implicit Medical World Models?
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?
How to Train a Critic Stably and Efficiently
How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles
ProBel: Propaganda Detection with Techniques, Spans, and Explanations
WAM-OPD: On-Policy Distillation for World Action Models
AdaR: A Framework for Equipping LLMs with Adaptive Reasoning
What LLMs explain is not what they believe: Evaluating explanation sufficiency under mode…
Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator
GenCoord: Skill-Path Commitments under Private Information