Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in …
Explorar
Noticias de IA
22116 elementos — filtrados, clasificados y sin duplicados
JSCGC: Joint Source-Channel-Generation Coding for Wireless Generative Communications
TimeROME-DLM: Temporal Causal Tracing and Low-Rank Inference-Time Knowledge Editing for M…
OCOO-T : A Simple and Scalable Virtual Cell Model for Transcriptional Perturbation Respon…
Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning
Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: …
Localizing Anchoring Pathways in Language Models
ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling
Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning
Agentic MPC for Semantic Control System Resynthesis
AfriSUD: A Dependency Treebank Collection for Evaluating Models on African Languages
LLM-Powered Personalized Glycemic Assessment in Type 2 Diabetes with Wearable Sensor Data
Two-Layer Linear Auto-Regressive Models Estimate Latent States
Token Complexity Theory for AI-Augmented Computing
Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
HybridCodeAuthorship: A Benchmark Dataset for Line-Level Code Authorship Detection
Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs
Graph Reduction in Multirelational Networks: A Spreading-Oriented Reduction Benchmark
A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constr…
Improving Crash Frequency Prediction from Simulated Traffic Conflicts Using Machine Learn…
Speculative Rollback Correction for Quality-Diverse Web Agent Imitation
ReCal: Reward Calibration for RL-based LLM Routing
GeoDial: A Multimodal Conversational Tutoring Dataset for Geometry Problem-Solving with V…
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibili…
Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning
Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight…
Reward Modeling for Multi-Agent Orchestration
Multiagent Protocols with Aggregated Confidence Signals
A Three-Layer Framework for AI in Scientific Discovery
Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Pe…