Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Intelligent Detection and Mitigation of Carpet-Bombing DDoS Attacks in SDN Using Retrieva…
APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented…
Decoupled Delay Compensation: Enhancing Pre-trained MARL Policies via Learned Dynamics Fi…
VesselSim: learning 3D blood vessel segmentation without expert annotations
Unified Neural Scaling Laws
AgentSociety: Incentivizing Agentic Social Intelligence
FLUIDSPLAT: Reconstructing Physical Fields from Sparse Sensors via Gaussian Primitives
Tool Calling is Linearly Readable and Steerable in Language Models
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagati…
When VLMs 'Fix' Students: Identifying and Penalizing Over-Correction in the Evaluation of…
BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning
Co-folding model guided by structural proteomics
HRVConformer: Neonatal Hypoxic-Ischemic Encephalopathy Classification from the Heart Rate…
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Langua…
SetupX: Can LLM Agents Learn from Past Failures in Functionality-Correct Code Repository …
The ATOM Report: Measuring the Open Language Model Ecosystem
From Static Context to Calibrated Interactive RL: Mitigating Distribution Shift in Multi-…
PolyFusionAgent: A Multimodal Foundation Model and Autonomous AI Assistant for Polymer Pr…
MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Explora…
MedGuideX: Internalizing Decision Logic from Executable Guidelines into Large Language Mo…
UnityMAS-O: A General RL Optimization Framework for LLM-Based Multi-Agent Systems
Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Ret…
It's Not the Capability: Harness Sensitivity Is Non-Monotone Across LLM Agent Tiers
Towards Feedback-to-Plan Decisions for Self-Evolving LLM Agents in CUDA Kernel Generation
Helicase: Uncertainty-Guided Supply Chain Knowledge Graph Construction with Autonomous Mu…