Agentic coding without the cloud: evaluating open-weight large language models on longitu…
Explorar
Noticias de IA
30660 elementos — filtrados, clasificados y sin duplicados
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle a…
Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
OpenForgeRL: Train Harness-native Agents in Any Environment
From Agent Failures to Text Policies: What Works and What Breaks
Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Lear…
Benchmarking Unlearning for Vision Transformers
Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for D…
From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying …
Agent-Guided Relational Concept Discovery: Toward Interpretable Surgical Margin Assessment
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamic…
Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integr…
DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
Multimodal Pretraining for Generalizable EEG Representation Learning
RE-AD: Real-Time Requirement Adherence for Data Labeling
OPOD: On-Policy Omni Distillation
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Constrained latent state modeling: A unifying perspective on representation learning unde…
Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions …
Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design
Probabilistic Residual Learning for Online Recommendations
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning