How Chain-of-Thought Works? Tracing Information Flow from Decoding, Projection, and Activ…
Explorar
Noticias de IA
21270 elementos — filtrados, clasificados y sin duplicados
LiPUP-MA: A Residential Experience-centric Multi-Agent Framework for Living-in-the-loop P…
Governed Capability Evolution: Lifecycle-Time Compatibility Checking and Rollback for AI-…
GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesis, Research, and Testing
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Det…
The ATOM Report: Measuring the Open Language Model Ecosystem
Risk Averse Alert Prioritization for IDS Using Subnormal Gaussian Fuzzy Models
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
Governed Evolution of Agent Runtimes through Executable Operational Cognition
Many Logics, One Methodology: A Plea for Logical Pluralism in Formalised Reasoning (prepr…
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Langua…
Algebraic Semantics of Governed Execution: Monoidal Categories, Effect Algebras, and Cote…
Mechanized Foundations of Structural Governance: Machine-Checked Proofs for Governed Inte…
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Mind the Tool Failures: Achieving Synergistic Tool Gains for Medical Agents
OmniToM: Benchmarking Theory of Mind in LLMs via Explicit Belief Modeling
Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation
BrainDINO: A Brain MRI Foundation Model for Generalizable Clinical Representation Learning
When VLMs 'Fix' Students: Identifying and Penalizing Over-Correction in the Evaluation of…
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Dat…
Bridging Classification and Reconstruction: Cooperative Time Series Anomaly Detection
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentat…
Yes, Q-learning Helps Offline In-Context RL
Self-Cascaded Diffusion Models for Arbitrary-Scale Image Super-Resolution
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
Tool Calling is Linearly Readable and Steerable in Language Models
Beyond Questions: Evaluating What Large Language Models (Actually) Know