LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
GenMatter: Perceiving Physical Objects with Generative Matter Models
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinem…
Parameter Efficient Hybrid Transformer (PEHT) for Network Traffic Prediction via Dynamic …
Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behav…
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
ELF: Embedded Language Flows
DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand
PreferThinker: Reasoning-based Personalized Image Preference Assessment
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
SEA-TS: Self-Evolving Agent for Autonomous Code Generation of Time Series Forecasting Alg…
Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing
ReFreeKV: Towards Threshold-Free KV Cache Compression
Ontology-Guided Evidence Path Inference for Multi-hop Knowledge Graph Question Answering
Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Mode…
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation i…
Deep Neural Networks Inspired by Differential Equations
Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generat…
Calibrating Biophysical Models for Grape Phenology Prediction via Multi-Task Learning
LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery
Hybrid coupling with operator inference and the overlapping Schwarz alternating method
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
Distribution-based deep multiple instance learning for tumor proportion scoring in NSCLC
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction