Abstract representational geometry supports inference in large language models
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs
SPADE: Structure-Prior Adaptive Decision Estimation
DART: Draft-Agreement Routing for Training-Free Adaptive Thinking Budgets in Hybrid Reaso…
Decomposing Financial Market Dynamics via Mechanism Analysis in an Evolutionary Multi-Age…
Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
TTFT-Aware Graph Chain-of-Thought:Distance-Indexed Neural A* for Low-Hallucination Multi-…
Cognitive Digital Twins: Ethical Risks and Governance for AI Systems That Model the Mind
Some Results about the Expressivity of Preference-Incomplete Structured Argumentation Fra…
IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Aut…
From numerical proportions to analogical proportions between probabilities
A Stackelberg Framework for Resource-Aware LLM Agents: Learning, Repair, and Conditional …
When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Mode…
The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language…
ENVS: Environment-Native Verified Search for Long-Horizon GUI Agents
When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents
AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Finding the Evidence: Discovering Decision-Supporting Tokens for On-Policy Reasoning Dist…
Active Inference as the Test-Time Scaling Law for Physical AI Agents
AI-Assisted Help-Seeking Trajectories in Programming Education from an SRL-Informed Persp…
A Formula-Driven Survey and Research Agenda for On-Policy Distillation
Learning Filters with Certainty
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via…
Safety-Aware Evaluation of LLM-Generated Driver Intervention Messages through Multi-Task …
Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Represe…
AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding …
Confident but Conflicted: Internal Uncertainty and Cognitive Dissonance Resolution in LLMs
SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement