Explorar

Noticias de IA

37834 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament …
arXiv cs.AI Research & Papers
Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
arXiv cs.AI Research & Papers
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
arXiv cs.AI Research & Papers
When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Rea…
arXiv cs.AI Research & Papers
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
arXiv cs.AI Research & Papers
HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
arXiv cs.AI Research & Papers
Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Univer…
arXiv cs.AI Research & Papers
Large language models for partial differential equation workflows
arXiv cs.AI Research & Papers
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
arXiv cs.AI Research & Papers
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Fre…
arXiv cs.AI Research & Papers
UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
arXiv cs.AI Research & Papers
Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Do…
arXiv cs.AI Research & Papers
One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning
arXiv cs.AI Research & Papers
LiveEvalBench: Toward Open-World Evaluation for Web Generation
arXiv cs.AI Research & Papers
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchang…
arXiv cs.AI Research & Papers
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algo…
arXiv cs.AI Research & Papers
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Mult…
arXiv cs.AI Research & Papers
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
arXiv cs.AI Research & Papers
Implementing Causal Perception: Competing SCMs and Situated Fairness
arXiv cs.AI Research & Papers
Multi-Camera Trajectory Forecasting with Trajectory Tensors
arXiv cs.AI Research & Papers
KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluati…
arXiv cs.AI Research & Papers
When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
arXiv cs.AI Security & Safety
$S^3$: Improving Agent Safety through Multi-Stage Defense
arXiv cs.AI Security & Safety
A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
arXiv cs.AI Research & Papers
Learning Molecular Representations from Cellular Phenotypes with Structure Preservation
arXiv cs.AI Research & Papers
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physi…
arXiv cs.AI Research & Papers
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
arXiv cs.AI Research & Papers
Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures
arXiv cs.AI Research & Papers
ISEE: Interactive Semantic Enrichment for Database Fields
arXiv cs.AI Research & Papers
Self-Organising Digital Circuits