Evidence-Ledger Adjudication for Claim-Evidence Traceability
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Commu…
AI as Friction for Reflection Support in Ideation
Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking
Gated Adaptation for Continual Learning in Human Activity Recognition
Hearsay: Vision-Language Medical Diagnoses Without an Image
Adaptively Robust LLM Monitoring via Activation Watermarking
FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mut…
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrie…
Exact Symmetry as Algebra: A Machine-Verified Tensor Calculus that Enforces Physical Sele…
SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search
PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low…
The Rise of AI in Weather and Climate Information and its Impact on Global Inequality
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness…
ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language …
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Def…
Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
Human diversity fuels collective creativity that large language models cannot simulate or…
Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing
Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language…
Do Methods Support the Claims? Intra-Paper Verification for Peer Review
An Unofficial FastLAS Tutorial: A Programmer's Guide
Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents
Where Is the Cost of Third-Party API Routers in Agentic Software Development?
Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp a…
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation