Safety Signals to Verify NetOps Agents with Action-Level Granularity
Browse
AI News
37834 items — filtered, classified, deduplicated
When does a scaling result justify a different allocation? A critical review of resource-…
Tool Use Reduces Depth-Induced Collapse in OOD Reasoning
Diffusion-Based Generation of Gait Trajectories
Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M
AcquireBound: Runtime Authorization for Resources Acquired by AI Agents
AI Persuasion as a Threat to Human Control
Another Blueprint In The Wall: How to Ask Frontier AI Like a Kid?
Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accou…
Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference
Attention Is All You Need (to Avoid Spurious Oscillations)
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary
Domain Generalization for Smartphone-Based Human Activity Recognition: A Systematic Analy…
GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems
Converting Sequenced Fuzzy Cognitive Maps to Causal Virtual Worlds with Large Video Gener…
Externalizing Requirement-to-Repair Artifacts as Observable Traces for LLM-Based Program …
MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
Geometric Flow enhanced Graph Coarsening
Semantic-TVM: Structure-Preserving Trustworthy Virtual Memory for Memory-Augmented and To…
Overflip: Repetition-Induced Label Flips in Guardrail Models
The average-farmer illusion in language-model simulations of agricultural decisions
Enabling Creative Exploration for Vibe Design Agents
STHMoE: Hypergraph-Enhanced Heterogeneous Dependency Coordination for LLM-Based Urban Tra…
VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries
Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generat…
ProIQA: A Process-Based Framework for Fine-Grained Math Item Quality Assessment
A Voxel-Spacing-Aware Extension of PyRadiomics for Anisotropic Texture Analysis
Mizan: A National Benchmark for Evaluating Large Language Models on Iraqi Arabic and the …
HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments
Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection