Forget or Fine-tune? A Comparative Study of Machine Unlearning Strategies for Noisy Label…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
HoopMind: A Real-Time Neural Game-Tree System for Opponent-Aware Possession Planning
"Act Like a 5th Grader" is Not Enough: Bounding Knowledge in LLM-Based User Simulators
Flow-JEPA: Flow Matching for Robust Latent Dynamics in JEPA World Models
TPR-Attention for Combinatorial Generalization
The reach of a verification tool decides its value: A controlled study of verification su…
Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextuali…
Curvature Cryptanalysis of Smooth Transformer Feed-Forward Networks
Star-Fusion: A Multi-modal Transformer Architecture for Discrete Celestial Orientation vi…
AgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud Agent
ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation
A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction
R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Ali…
PRISM: Predictive Recomposition via Semantic Latent Decomposition for View-invariant Vide…
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
LCoT-GV: Graph Attention Networks for Verifying Long Reasoning Chains in Large Language M…
CoFiRec: Coarse-to-Fine Tokenization for Generative Recommendation
Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-R…
Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends
Toward Scalable Audio Description Quality Control: A Workflow for Evaluating Human and VL…
Stratified Consistency Distillation for Natural Language Formalization
GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning
On the Prospects of Dynamic LLM Conversations in Software Development
Beyond Dense States: Sparse Transcoders as Causally Testable Operators for LLM Latent Rea…
GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Infere…
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research …
Masked Distillation: Internalizing the Chain-of-Thought in Language Models
Repair or Resample? Rethinking Failure Debugging in LLM Multi-Agent Systems
SimGuide: Typed Multi-Context User Representations for Preference-Conditioned Agent Plann…
post-graph-rag: A PostgreSQL-Native Bi-Temporal Graph RAG Engine with Temporal Grounding …