AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Explorar
Noticias de IA
30588 elementos — filtrados, clasificados y sin duplicados
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Dist…
Active Inference as the Test-Time Scaling Law for Physical AI Agents
AI-Assisted Help-Seeking Trajectories in Programming Education from an SRL-Informed Persp…
A Formula-Driven Survey and Research Agenda for On-Policy Distillation
Learning Filters with Certainty
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via…
Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task …
Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Represe…
AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding …
Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs
SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement
Text2DSL: LLM-Based Code Generation for Domain-Specific Languages
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horiz…
Imagine to Ensure Safety in Hierarchical Reinforcement Learning
Grounded Scaling: Why Agentic AI Needs Deterministic Environments
SCOPE: Evolving Symbolic World for Planning in Open-Ended Environments
PRIME: Evaluating Prompt Resolution Under Incompatible Instructions in LLMs
The More the Merrier: Combining Properties for ABox Abduction under Repair Semantics in E…
FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR
CAOA -- Completion-Assisted Object-CAD Alignment
The Impact of AI Search on the Online Content Ecosystem: Evidence from Google and Reddit
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
TMPO: Trajectory Matching Policy Optimization for Diverse and Efficient Diffusion Alignme…
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimiza…