A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Voice "Cloning" is Style Transfer
NCSAM Noise-Compensated Sharpness-Aware Minimization for Noisy Label Learning
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Act…
Improving Evaluation of Recombination-based Cartesian Genetic Programming
Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
Do readers prefer AI-generated Italian short stories?
Entropy-aware Masking for Masked Language Modeling
Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation
Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection
On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning
Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refi…
Picid: A Modular Evaluation Infrastructure for Reproducible PHM Across Tasks and Domains
Cultural Fidelity in English-to-Hindi Translation: A Preservation-Fluency Frontier for Ge…
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
Adapting, Fast and Slow: On Few-Shot Transportability of Compositions
Generic Interpretation Approach for Transformer Models Incorporating Heterogenous Attenti…
From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constr…
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement …
Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning
You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal Stat…
Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language…
Measuring Progress Toward AGI: A Cognitive Framework
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoni…
SkillGrad: Optimizing Agent Skills Like Gradient Descent
GraD-IBD: Graph Representation Learning from Diagnosis Trajectories for Early Detection o…
RAGe: A Retrieval-Augmented Generation Evaluation Framework
CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diag…