Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Le…
Explorar
Noticias de IA
30313 elementos — filtrados, clasificados y sin duplicados
Does Distributed Training Undermine Compute Governance?
Training Deliberative Monitors for Black-Box Scheming Detection
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
A unified deeplearning framework for contrast-phase-specific virtual monochromatic imaging
Data filtering methods for training language models
ESPO: Early-Stopping Proximal Policy Optimization
Projectional Decoding: Towards Semantic-Aware LLM Generation
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EP…
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained…
Assessing Dutch Syllabification Algorithms and Improving Accuracy by Combining Phonetic a…
Specialty-Specific Medical Language Model for Immune-Mediated Diseases
Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization
Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents
Provably Secure Agent Guardrail
Harmonizing Real-Time Constraints and Long-Horizon Reasoning: An Asynchronous Agentic Fra…
Indexing the Unreadable: LLM-Native Recursive Construction and Search of Service Taxonomi…
ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role…
ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for …
The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions vi…
Xetrieval: Mechanistically Explaining Dense Retrieval
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, a…
Opt-Verifier: Unleashing the Power of LLMs for Optimization Modeling via Dual-Side Verifi…
Personalized Turn-Level User Conversation Satisfaction Benchmark
CB-SLICE: Concept-Based Interpretable Error Slice Discovery
Planning with the Views via Scene Self-Exploration
DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Super…