Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Unveiling Multi-regime Patterns in SciML: Distinct Failure Modes and Regime-specific Opti…
SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search
A Minimal Bifurcation Model of Load Imbalance in a Softmax Mixture-of-Experts Router
When and How Long? The Readout-Mediator Angle in Temporal Reasoning
Test Time Training for Supervised Causal Learning
Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?
On the Geometry of Games and their Solvers
Make LLM Learn to Synthesize from Streaming Experiences through Feedback
No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
A Predictive Law for On-Policy Self-Distillation From World Feedback
Accelerating Constrained Decoding with Token Space Compression
RAISE: RAG Design as an Architecture Search Problem
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical …
Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Meth…
Anchorless Diversification for Parallel LLM Ideation
Temporal Stability and Few-Shot Prompting in Math Task Assessment
Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of …
Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance
Agora: Toward Autonomous Bug Detection in Production-Level Consensus Protocols with LLM A…
Double-Edged Sword or Sharp Tool? Designing and Evaluating Triadic LLM-Teacher Collaborat…
Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised …
mcp-proto-okn: Natural-language access to open scientific knowledge graphs through the Mo…
MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
Micro-Macro Retrieval: Reducing Long-Form Hallucination in Large Language Models
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report…
Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Models