HELIX: Model-Harness Co-evolution for Recursive Self-Improvement
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Joint Optimization of Memory and Computing Frequency for Energy-Efficient DNN Inference
MemoryLake on MemoryArena: A Matched Study of Agent Memory Backends
Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact
SDO: Subspace Deconflicting Operator for Multi-Adapter Composition
Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy
From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small La…
Explanation Multiplicity: Circuit-Level Interpretability Evidence Does Not Survive Defens…
Coverage Aware Active Evaluation for Failure Discovery with Paired Systems
A Calibrated Test of Internal Action Maps: State Signals Without Global Affine Closure
How Compliant is Sepsis Treatment? An Expert-Guided Neuro-symbolic Pipeline for Generatin…
Exploring ESC Winners with Nested Diagrams
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation o…
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
FLARE MCMC: Fidelity-based Layer-Adaptive REcursive proposals for MCMC
Reward Machines for Signal Temporal Logic
Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, a…
Active Perception for Embodied Disambiguation
No Universal Signal Predicts Sample-Level LLM Regression under Version Updates
SemPlan: Benchmarking Structured Semantic Planning for LLM-Based Queries over Enterprise …
Depth-Aware Sensitivity Analysis of Mixture-of-Experts Models via Magnitude-Based Expert …
Inducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent Evaluation
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents
NEURON: A Neuro-symbolic System for Grounded Clinical Explainability
AI Evaluation Should Work With Humans
Simulation-Driven Vehicular Traffic Data Augmentation: Extending Sensor Coverage Through …
Measuring Cross-Task Behavioral Consistency in Language Model Agents
Context Aware AI Assistant and AR Interface for Lunar Extravehicular Activity (EVA) Proce…