Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
Explorar
Noticias de IA
29336 elementos — filtrados, clasificados y sin duplicados
Proportional Analogies on Probability Distributions via Bayesian Updating
Chain-of-Thought Shows the Path to a Tree: Realizing Branching Complexity
Nvidia Partner Hon Hai’s Profit Beats on Sustained AI Spending
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
From assistance to execution: How enterprises put AI to work
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance A…
Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic S…
Vacuum Pump Makers, Producers of Gases Are the New AI Winners
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperativ…
Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Moti…
Towards a Formal Definition of Agent Memory: Basis, Span, Optimality, and the Sequential …
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Enco…
Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study
Robustness of AI-Art Detectors under Generator Shift
CAM-Guided Saliency Cutout and Image-Based Malware Classification
Making AI-Generated Feedback Matter: From Provision to Student Enactment
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time serie…
Generative Video Compression Based on Hierarchical Referencing
KANResDiff: Learning Local Residual Diffusion via Kolmogorov-Arnold Network for Ambiguous…
AI Startup Cognition in New Funding Talks at $40 Billion Value
Counterpoint's Shah on Nvidia, Circular Deals
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
RECAST: A Machine-Learning Framework for Correction and Super-Resolution of Coarse-Grid P…
Company Offering '100% Human-Written, Never AI' Medical Research Is 100% AI
When Offline Evaluation Misleads: A Diagnostic Protocol for Reward and Policy Selection i…
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agen…