SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomi…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
LLMBDC: Language Model for Biological Domains Oriented Clustering of Gene Ontology
GRAIN: Molecules Are Not the Right Granularity -- Active-Ingredient Modeling for Safe Med…
Conservation laws determine what physical learning remembers
ELECTRIC: Evidential Learning-Enhanced CT Reconstruction via Iterative Correction
Neural Circuit Function Inference with LLMs
Fast Generation of Representative Synthetic Dataset with Salsa to Train ATR Models with E…
Role Steering of Language Models for Social Simulations
Uncertainty-Aware Simulation-Based Inference for Operations Research with Large Language …
CITBench: A Comprehensive Benchmark for Interactive Tabular Data Processing with LLMs
Obshazard-bench: Benchmarking Multimodal Foundation Models for Real-Time Disaster Intelli…
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
AgentMemBench: A Systematic Benchmark for Evaluating Long-Term Memory Management Strategi…
RubricReviewer: From Direct Critique to Objective and Comprehensive Rubric-Driven Peer Re…
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluatio…
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies
A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI
Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientifi…
CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum O…
Real-Time Detection and Repair of LLM Agent Failures
Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+
Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on …
Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training
Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcem…
Faster-WAM: Do World Action Models Need Deep Action Modules?
KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement
Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories