VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Explorar
Noticias de IA
21010 elementos — filtrados, clasificados y sin duplicados
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
Forecasting Side Effects of Activation Steering
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Uni…
Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchi…
Every pooling rule has its world: matching probability combination rules to situations an…
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty…
Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Qu…
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discove…
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Lang…
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal …
Policy-as-logic for robust reasoning over rules
Representation Finetuning for Continual Learning
Cavity-Enhanced Collective Quantum Processing with Polarization-Encoded Qubits
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Sy…
CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classificat…
Probably Approximately Correct Maximum A Posteriori Inference
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invaria…
Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin D…
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Mode…
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Se…
MaSRead: Content-Addressed Reading of Replicated Latent Stores
Quantization-Aware Neuromorphic Architecture for Skin Lesion Classification on Resource-C…
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error…
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk