HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Ur…
Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured…
EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on C…
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Info…
ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One…
A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on …
Toward a New Science of AI as Cognitive Infrastructure
CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery
Recurrence Meets Transformers for Universal Multimodal Retrieval
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
No Plan, Yet Human: A Reactive Robotics Model Predicts Human Planning Failures on a Clini…
Fine-Tuning of Transformer models with Frames
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Predi…
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-…
Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Me…
Agentic AI for operating scientific instruments for nanoscale characterization
Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical R…
The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasti…
Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
Cartan flow matching
Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems
SynthCharge: An Electric Vehicle Routing Instance Generator with Feasibility Screening to…
From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Ca…
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
BPMN4CAI: A BPMN Extension for Modeling Dynamic Conversational AI