Explorar

Noticias de IA

37834 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
arXiv cs.AI Research & Papers
Evaluating Multiple LLM Generations with Validated Task Coverage
arXiv cs.AI Research & Papers
Adaptive Influence Graphs for Failure Attribution in Multi-Agent Systems
arXiv cs.AI Research & Papers
Scalable Question-Centric Text-to-Image Evaluation: Reliable Ranking, Fine-Grained Diagno…
arXiv cs.AI Research & Papers
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
arXiv cs.AI Research & Papers
Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
arXiv cs.AI Research & Papers
SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Percepti…
arXiv cs.AI Research & Papers
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation
arXiv cs.AI Research & Papers
EMRB: A Multi-Level Benchmark for Evaluating LLM Reasoning over Raw Electromagnetic Signa…
arXiv cs.AI Research & Papers
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Visi…
arXiv cs.AI Research & Papers
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device La…
arXiv cs.AI Research & Papers
Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audi…
arXiv cs.AI Research & Papers
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human O…
arXiv cs.AI Research & Papers
From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
arXiv cs.AI Research & Papers
Method, Mind, and Morality: How People Make Sense of Artificial Intelligence
arXiv cs.AI Research & Papers
Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watche…
arXiv cs.AI Research & Papers
ToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Callin…
arXiv cs.AI Research & Papers
Ensemble of Convolutional Neural Networks for StrokePrediction: Towards Improved Diagnost…
arXiv cs.AI Research & Papers
ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory For…
arXiv cs.AI Research & Papers
Identifying Latent Declarative Representations of Code for Assisting Repository Migration
arXiv cs.AI Research & Papers
FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs
arXiv cs.AI Research & Papers
CoMMa: Contribution-Aware Medical Multi-Agents for Decentralized Oncology Decision Support
arXiv cs.AI Research & Papers
When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and th…
arXiv cs.AI Research & Papers
ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Le…
arXiv cs.AI Research & Papers
Restoring Without Forgetting: Continual Learning Across Image Degradations
arXiv cs.AI Research & Papers
PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understandi…
arXiv cs.AI Research & Papers
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Ba…
arXiv cs.AI Research & Papers
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect General…
arXiv cs.AI Research & Papers
Robust Motion Generation using Part-level Reliable Data from Videos
arXiv cs.AI Research & Papers
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving