RUMBA: Russian User Memory Benchmark
Explorar
Noticias de IA
22115 elementos — filtrados, clasificados y sin duplicados
Benchmarking Unlearning for Vision Transformers
From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurre…
Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dim…
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
From Agent Failures to Text Policies: What Works and What Breaks
Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Lear…
A Counterfactual Cause in Situation Calculus
Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
StackingNet: Collective Inference Across Independent AI Foundation Models
Phonetic forced alignment for low-resource language varieties: Model training and evaluat…
Diagnosing Pathological Chain-of-Thought in Reasoning Models
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based D…
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Syn…
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligen…
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failu…
Synthetic minority data is redundant or invalid: a data-dependent validity theory and a d…
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language…
Equivariant Conditional Diffusion Model for Head and Neck CT Image Synthesis from CBCT
Simple Policy Gradients for Reasoning with Diffusion Language Models
On the Granularity of Causal Effect Identifiability
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine …