JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Explorar
Noticias de IA
30934 elementos — filtrados, clasificados y sin duplicados
RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
Interactive Task Alignment as a POMDP
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
Mechanistic Attention Guidance for Agent Memory Refinement
ProEvent: An Event-centric Benchmark for Proactive Agents
Boundary-Seeking GAN-Augmented TabTransformer for Adversarially Robust Intrusion Detection
The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior …
KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
Learning Structural Manipulability in Gate-Level Netlists Using Graph Neural Networks
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction
Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support
Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution
AEC-DS: Adaptive Erasure Coding with PDP-Triggered Reputation and QoS-Aware Migration for…
Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Searc…
ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Compan…
Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation M…
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Foo…
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional …
Neural Controlled Differential Equations for EMT-Level Surrogate Modeling of Grid-Forming…
Reducing Per-Sample Harm in Stochastic Optimization
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
Reliable Remediation Impact Prediction for Black-Box Security Ratings
A Survey on the Verification of Reinforcement Learning Policies
Accurate and Efficient Long-Term Memory for LLM Agents
PRISM: Multimodal Terrain Mapping for Rover Navigation in Unstructured Environments
AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers