When does a scaling result justify a different allocation? A critical review of resource-…
Browse
AI News
37834 items — filtered, classified, deduplicated
Beyond Scene Description: Multi-Agent Orchestration for Non-visual Access to Virtual Worl…
CodeTS: Verifiable Text-to-Time Series Generation via Executable Code
Tool Use Reduces Depth-Induced Collapse in OOD Reasoning
Attention Is All You Need (to Avoid Spurious Oscillations)
AcquireBound: Runtime Authorization for Resources Acquired by AI Agents
AI Persuasion as a Threat to Human Control
MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving
Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accou…
One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling
Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program …
A Voxel-Spacing-Aware Extension of PyRadiomics for Anisotropic Texture Analysis
MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
Geometric Flow enhanced Graph Coarsening
Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misali…
Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and To…
Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomed…
The average-farmer illusion in language-model simulations of agricultural decisions
STHMoE: Hypergraph-Enhanced Heterogeneous Dependency Coordination for LLM-Based Urban Tra…
VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries
Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generat…
ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cock…
Parameter-Efficient Adaptation of Pretrained Language Models for Time-Series Forecasting
Who Teaches Which Token? Verifier-Gated Multi-Expert On-Policy Distillation for Scientifi…
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary
RA-CoA: Training-free Fashion Image Captioning via Retrieval-Augmented Chain-of-Attributes
Oops, Not Now: PEARL, a RAG-Based Support Agent for Gameplay and What Players Want from A…
Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection