Explorar

Noticias de IA

27413 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
Metanormative Theory for RL-Based Moral Agents
arXiv cs.AI Research & Papers
SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
arXiv cs.AI Research & Papers
OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents
arXiv cs.AI Research & Papers
Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-wo…
arXiv cs.AI Research & Papers
Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders
arXiv cs.AI Research & Papers
Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump…
arXiv cs.AI Research & Papers
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoni…
arXiv cs.AI Research & Papers
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
arXiv cs.AI Research & Papers
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions a…
arXiv cs.AI Research & Papers
Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
arXiv cs.AI Research & Papers
RAG-Based Auto-Configuration for Industrial Fieldbus Devices
arXiv cs.AI Research & Papers
TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Perso…
arXiv cs.AI Research & Papers
Estimating Uncertainty in Galaxy Morphology Classification
arXiv cs.AI Research & Papers
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
arXiv cs.AI Research & Papers
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
arXiv cs.AI Research & Papers
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark a…
arXiv cs.AI Research & Papers
Adversarial Latent-State Training for Robust Policies in Partially Observable Domains
arXiv cs.AI Research & Papers
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
arXiv cs.AI Research & Papers
Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Comp…
arXiv cs.AI Research & Papers
NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
arXiv cs.AI Research & Papers
A Sobering Look at Tabular Data Generation via Probabilistic Circuits
arXiv cs.AI Research & Papers
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
arXiv cs.AI Research & Papers
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
arXiv cs.AI Research & Papers
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
arXiv cs.AI Research & Papers
Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control
arXiv cs.AI Research & Papers
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
arXiv cs.AI Research & Papers
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Al…
arXiv cs.AI Research & Papers
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimod…
arXiv cs.AI Research & Papers
Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of…
arXiv cs.AI Research & Papers
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?