RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
Explorar
Noticias de IA
21863 elementos — filtrados, clasificados y sin duplicados
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain St…
Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intri…
LAWFUL: Law-Aligned Witness for Faithful Use of Latents
PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learnin…
TerraNova: A Foundation Model for the Anthropocene
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Rie…
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizo…
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Sys…
Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Par…
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Gated Q-learning: Add Off-Policy Bias to Taste
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Exper…
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Ac…
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported …
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Creative Integration: A Decidable Criterion of Creativity
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Rememb…
Beyond Component Testing: Validating Agentic AI Systems
Beyond Retrieval: Analytic Memory for Multimodal Agents
AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
COntExt: Towards Context-Aware Ontology Extension from Operational Metrics