JarvisBench: Always-on Intelligence Between Humans and Agents
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma …
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Doc…
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated …
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detect…
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an E…
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Gene…
Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Langu…
ARENA: Automated Red-Teaming for Large Audio Language Models
ShadowNet for Data-Centric Quantum System Learning
Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spect…
A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network E…
Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpr…
Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair
ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity S…
Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusa…
Hoeffding adaptive splitting trees for data stream classification with concept drift and …
Beyond Single Object: Learning 3D Relations with Large Language Models