Talking to Itself While Coding: What Makes Comments Help Code Generation?
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
MADS: Multi-Agent Dialogue Simulation for Diverse Persuasion Data Generation
RAU: Reference-based Anatomical Understanding with Vision Language Models
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Le…
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affe…
Show-Harness: Just a VLM Agent Can Play Robots
What Makes Adversarial Examples Transfer Across Deepfake Detectors?
MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
Albedo Estimation via Latent Bridge Matching
Predicting Estimated Times of Restoration for Electrical Outages Using Longitudinal Tabul…
Seven Sources of Physical AI Capability Formation
Safe Learning Under Irreversible Dynamics via Asking for Help
Strangers to Themselves: What Language Models Say About Themselves Is Generic
Efficient Diversity-based Experience Replay for Deep Reinforcement Learning
Query Brand Entity Linking in E-Commerce Search
Distributed Physical Layer Authentication and Collaborative RSMA in Non-Terrestrial Netwo…
Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA
A Taxonomy of Architecture Options for Foundation Model-based Agents: Analysis and Decisi…
Influence-Oriented Personalized Federated Learning
"What Are You Really Trying to Do?": Co-Creating Life Goals from Everyday Computer Use
Incentives to Offer Algorithmic Recourse
Equity Promotion in Online Resource Allocation
An Autonomous GeoAI Agent for Arctic Eco-Navigation
NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acou…
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Bench…
The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Tra…
Can AI Agents Detect and Repair Artifact Drift in Network Experiments?
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
FrontierChallenge: Evaluating Scientific Workflow Completion