Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
Explorar
Noticias de IA
22116 elementos — filtrados, clasificados y sin duplicados
Characterizing Software Aging in GPU-Based LLM Serving Systems
Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Pr…
Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
LASA: A Weak Supervision Method for Open-Vocabulary Scene Sketch Semantic Segmentation
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in M…
When Researchers Say Mental Model/Theory of Mind of AI, What Are They Really Talking Abou…
Causal Emotion Recognition in Conversation: Context Saturation and Discourse-Marker Evide…
To Intervene or Not: Guiding Inference-time Alignment with Probabilistic Model Blending
Sparsified Kolmogorov-Arnold Networks for Interpretable Quantum State Tomography
Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for S…
Dual-Stance Evaluation of Sycophancy: The Structure of Agreement and the Limits of Interv…
BioDivergence: A Benchmark and Evaluation Framework for Hidden Contextual Contradictions …
ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based P…
The Impossibility of Eliciting Latent Knowledge
The Environmental Cost of LLMs in AIED: Reporting and Practices
AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory
Making Models Unmergeable via Scaling-Sensitive Loss Landscape
Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in …
SkillJuror: Measuring How Agent Skill Organization Changes Runtime Behavior
Hey Chat, Can You Teach Me? Structuring Socratic Dialogue for Human Learning in the Wild
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning
MedCTA: A Benchmark for Clinical Tool Agents
Noise-Aware Framework for Correcting Corrupted Labels
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching
Internet of Everything in the 6G Era: Paradigms, Enablers, Potentials and Future Directio…
Estimating Tail Risks in Language Model Output Distributions
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
LaQual: An Automated Framework for LLM App Quality Evaluation
Cross-Layer Discrete Concept Discovery for Interpreting Language Models