ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizo…
Explorar
Noticias de IA
21863 elementos — filtrados, clasificados y sin duplicados
TerraNova: A Foundation Model for the Anthropocene
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Gated Q-learning: Add Off-Policy Bias to Taste
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Exper…
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimiz…
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported …
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Ac…
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring…
ActionParty: Multi-Subject Action Binding in Generative Video Games
Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Sys…
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
Creative Integration: A Decidable Criterion of Creativity
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Oper…
Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement o…
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Tra…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift …
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
MOSAIC: Masked Outsourcing of Secure AI Computations
DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing
A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
SERUM: State Extraction and Refinement for User Modeling