OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evalu…
Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories
It`s All About Speed: AI`s Impact on Workflow in Music Production
Meta-Programming for Linear-time Temporal Answer Set Programming
Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs
A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach
Scaling Small Agents Through Strategy Auctions
MOO: A Multi-view Oriented Observations Dataset for Viewpoint Analysis in Cattle Re-Ident…
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
P$^2$RAG: Efficient Privacy-Preserving RAG Service Supporting Arbitrary Top-$k$ Retrieval
Hierarchical Task Network Planning with LLM-Generated Heuristics
A Foundation Model for Zero-Shot Logical Rule Induction
Bridge-RAG: An Abstract Bridge Tree Based Retrieval Augmented Generation Algorithm
EvA: An Evidence-First Audio Understanding Paradigm for LALMs
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Mode…
Reducing Political Manipulation with Consistency Training
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
GiPL: Generative augmented iterative Pseudo-Labeling for Cross-Domain Few-Shot Object Det…
PhoneWorld: Scaling Phone-Use Agent Environments
Source-Grounded Semantic Reinforcement Learning for Low-Resource Target-Language Generati…
Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Mul…
Composing Non-Conjugate Factor Graphs with Closed-Form Variational Inference
How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignmen…
TRACER: Persistent Regularization for Robust Multimodal Finetuning
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning