Ratio-Variance Regularized Policy Optimization
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent …
Scalable GANs with Transformers
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable…
Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction
HiSpec: Hierarchical Speculative Decoding for LLMs
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
Mechanistic Interpretability of Antibody Language Models Using SAEs
SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge…
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuit…
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigil…
Understanding Reasoning in LLMs through Strategic Information Allocation under Uncertainty
Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Sp…
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Pos…
Constructing Industrial-Scale Optimization Modeling Benchmark
Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts …
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User Hist…
GICDM: Mitigating Hubness for Reliable Distance-Based Generative Model Evaluation
Geometrically Constrained Outlier Synthesis
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
FedTreeLoRA: Reconciling Statistical and Functional Heterogeneity in Federated LoRA Fine-…
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
APEX-Searcher: Refining Credit Assignment with Subgoaling for Agentic Retrieval-Augmented…
Detached Skip-Links and $R$-Probe: Decoupling Feature Aggregation from Gradient Propagati…
Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argum…
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Ques…
Understanding the Challenges in Iterative Generative Optimization with LLMs
SenBen: Sensitive Scene Graphs for Explainable Content Moderation