Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Arti…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
TIMEGATE: Sustainable Time-Boxed Promotion Gates for Continual ML Adaptation Under Resour…
FinVerBench: Benchmark Validity and Calibration in Large Language Model Financial Stateme…
Sustainable Metal-Organic Framework Water Harvesters in the Artificial Intelligence Era
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Ad…
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured …
From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaborati…
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
VikingMem: A Memory Base Management System for Stateful LLM-based Applications
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Ar…
ReasonOps: Operator Segmentation for LLM Reasoning Traces
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
Governing Technical Debt in Agentic AI Systems
Extreme dynamic symmetry enables omnidirectional and multifunctional robots
Stochastic Lifting for Generating Trajectories of Stochastic Physical Systems
Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers…
Beyond Consensus: Trace-Level Synthesis in Mixture of Agents
Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting
Unlocking the Working Memory of Large Language Models for Latent Reasoning
Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federa…
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Super…
Planning with the Views via Scene Self-Exploration
CB-SLICE: Concept-Based Interpretable Error Slice Discovery
Personalized Turn-Level User Conversation Satisfaction Benchmark
Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verifi…
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, a…