Safety from Honesty in a Disinterested AI Predictor
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-G…
SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large L…
Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Re…
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence
Interpreting Latent CoT Reasoning as Dynamical Systems
From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argument…
CMSL: Constructive Multi-Sequence Learning for Recommendation Systems
RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
Graph Construction and Matching for Imperative Programs using Neural and Structural Metho…
Stable On-Policy Distillation through Adaptive Target Reformulation
Enhancing Adversarial Transferability through Block Stretch and Shrink
Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game G…
GES-TSP: Graph Edge Sparsification for TSP
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness…
ECoLAD: Selecting Anomaly Detectors for Automotive Deployment via Compute-Reduction Evalu…
AutoMatBench: An Automatic Optimization Toolkit for the Acceleration of Material Properti…
Heuristic Learning for Active Flow Control Using Coding Agents
Extending LLM Context via Associative Recurrent Memory
PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour
Constrained Reinforcement Learning for Safe Heat Pump Control
Asynchronous Perception Machine For Efficient Test-Time-Training
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Lea…
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software …
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Of…
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains