Explorar

Noticias de IA

21863 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizo…
arXiv cs.AI Research & Papers
TerraNova: A Foundation Model for the Anthropocene
arXiv cs.AI Research & Papers
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
arXiv cs.AI Research & Papers
Gated Q-learning: Add Off-Policy Bias to Taste
arXiv cs.AI Research & Papers
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
arXiv cs.AI Research & Papers
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
arXiv cs.AI Research & Papers
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Exper…
arXiv cs.AI Research & Papers
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimiz…
arXiv cs.AI Research & Papers
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported …
arXiv cs.AI Research & Papers
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Ac…
arXiv cs.AI Research & Papers
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
arXiv cs.AI Research & Papers
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
arXiv cs.AI Research & Papers
HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring…
arXiv cs.AI Research & Papers
ActionParty: Multi-Subject Action Binding in Generative Video Games
arXiv cs.AI Research & Papers
Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
arXiv cs.AI Research & Papers
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Sys…
arXiv cs.AI Research & Papers
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
arXiv cs.AI Research & Papers
Creative Integration: A Decidable Criterion of Creativity
arXiv cs.AI Research & Papers
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Oper…
arXiv cs.AI Research & Papers
Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
arXiv cs.AI Research & Papers
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
arXiv cs.AI Research & Papers
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement o…
arXiv cs.AI Research & Papers
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Tra…
arXiv cs.AI Research & Papers
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift …
arXiv cs.AI Research & Papers
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
arXiv cs.AI Research & Papers
MOSAIC: Masked Outsourcing of Secure AI Computations
arXiv cs.AI Research & Papers
DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing
arXiv cs.AI Research & Papers
A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem
arXiv cs.AI Research & Papers
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
arXiv cs.AI Research & Papers
SERUM: State Extraction and Refinement for User Modeling