Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for …
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Mechanism of Task-oriented Information Removal in In-context Learning
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Di…
Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player…
Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language M…
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
AI Assistance Reduces Persistence and Hurts Independent Performance
An empirical evaluation of the risks of AI model updates using clinical data: stability, …
CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross…
The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech …
Self-Organising Digital Circuits
ISEE: Interactive Semantic Enrichment for Database Fields
Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physi…
Learning Molecular Representations from Cellular Phenotypes with Structure Preservation
When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluati…
Multi-Camera Trajectory Forecasting with Trajectory Tensors
Implementing Causal Perception: Competing SCMs and Situated Fairness
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Mult…
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algo…
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchang…
LiveEvalBench: Toward Open-World Evaluation for Web Generation