APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented…
Explorar
Noticias de IA
21270 elementos — filtrados, clasificados y sin duplicados
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
Alignment Tuning for Large Language Models: A Data-Centric Lens on Alignment Data Pipelin…
Cross-scale Aligned Supervision for Training GANs
Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization
Demystifying Video Reasoning
Scalable GANs with Transformers
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable…
Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
HiSpec: Hierarchical Speculative Decoding for LLMs
Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
Mechanistic Interpretability of Antibody Language Models Using SAEs
SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge…
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuit…
Alignment Makes Language Models Normative, Not Descriptive
Where Code Meets Natural Language: Taxonomy-Driven Information Flow Analysis for LLM-Inte…
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigil…
Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Sp…
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Pos…
Constructing Industrial-Scale Optimization Modeling Benchmark
Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts …
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User Hist…
ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question Answering
FLUIDSPLAT: Reconstructing Physical Fields from Sparse Sensors via Gaussian Primitives
Tool Calling is Linearly Readable and Steerable in Language Models
GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation
Geometrically Constrained Outlier Synthesis