Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning
Explorar
Noticias de IA
22116 elementos — filtrados, clasificados y sin duplicados
OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Respon…
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Tes…
TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for M…
JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications
Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier
Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quanti…
CausalMoE: A Billion-Scale Multimodal Foundation Model for Granger Causal Discovery with …
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Orga…
Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in …
MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs
Constructing Evaluation Datasets for Procedural Reasoning: Balancing Naturalness, Groundi…
Token Complexity Theory for AI-Augmented Computing
A Machine Learning Framework for Real-Time Personalized Ergonomic Pose Analysis
Two-Layer Linear Auto-Regressive Models Estimate Latent States
Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism
From AGI to ASI
LLM-Powered Personalized Glycemic Assessment in Type 2 Diabetes with Wearable Sensor Data
Diffusion Transformer World-Action Model for AV Scene Prediction
Uncertainty-Aware Hybrid Retrieval for Long-Document RAG
OpenMedQ: Broad Open Pretraining for Medical Vision-Language Models
Zero-source LLM Hallucination Detection with Human-like Criteria Probing
AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages
HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasks
CRAFTIIF: Cross-Resolution Analytic Four-Type Interpretable Isolation Forest for Multivar…
Topical Phase Transitions in Artificial Intelligence Research: Large-Scale Evidence and a…
Agentic MPC for Semantic Control System Resynthesis