MARFT: Multi-Agent Reinforcement Fine-Tuning
Explorar
Noticias de IA
29655 elementos — filtrados, clasificados y sin duplicados
T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models
Deft Scheduling of Dynamic Cloud Workflows with Varying Deadlines via Mixture-of-Experts
Ideas in Inference-time Scaling can Benefit Generative Pre-training Algorithms
Skill-Based Mixture-of-Experts: Adaptive Routing for Heterogeneous Reasoning via Inferred…
Efficient LLM Moderation with Multi-Layer Latent Prototypes
ShapeLib: Designing a library of programmatic 3D shape abstractions with Large Language M…
Introduction to Graph Neural Networks for Machine Learning Engineers
Self-supervised Monocular Depth and Pose Estimation for Endoscopy with Latent Priors
Herculean: An Agentic Benchmark for Financial Intelligence
Implicit Regularization for Multi-label Feature Selection
Recent Advances in Multi-modal 3D Intelligence: A Comprehensive Survey and Evaluation
DeepIPCv2: LiDAR-powered Robust Environmental Perception and Navigational Control for Aut…
Can LLM Agents Sustain Long-Horizon Organizational Dynamics?
Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles …
c-TPE: Tree-structured Parzen Estimator with Inequality Constraints for Expensive Hyperpa…
Emergent Ordinal Geometry in Transformers Trained on Local Comparisons
Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Promp…
TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs
Versatile Framework with Semantic and Structural guidance for Image Reconstruction from B…
Hot-Start Chinese Language Modeling:Visual Glyphs Accelerate Sample-Efficient Learning
MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems
A Monosemantic Attribution Framework for Stable Interpretability in Clinical Neuroscience…
Towards a General Intelligence and Interface for Wearable Health Data
LLM-Guided Communication for Cooperative Multi-Agent Reinforcement Learning
Ethical Hyper-Velocity (EHV): A Hardware-Rooted Zero-Trust Runtime Enforcement Architectu…
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Ru…
Structured Visual Evidence Decomposition for Evidence-Grounded Multimodal Screening of Ob…