SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search
Explorar
Noticias de IA
30329 elementos — filtrados, clasificados y sin duplicados
Do Methods Support the Claims? Intra-Paper Verification for Peer Review
EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks
ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language …
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Econo…
The Art of Not Forgetting A Local Learning Architecture for Continual Learning
Living-Harness Is an Interactive-Agent Evolver
Entity Resolution in Practice: Lessons from a Self-Serve Pipeline
High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumpti…
Weight and Height Estimation from a Single Human Image Captured in the Wild
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM …
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research
The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolu…
Evidence-Ledger Adjudication for Claim-Evidence Traceability
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adap…
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection un…
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating …
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM
Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks
PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories
Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
Scientific Knowledge Discovery in the Age of Large Language Models
(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigati…
Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabiliti…
TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions