Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritag…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under n…
Why Does Robustness Reduce Superposition?
What LLMs explain is not what they believe: Evaluating explanation sufficiency under mode…
RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Si…
CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajecto…
Towards Comprehensive Basketball Understanding
CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents
What's the Catch? Evaluating Temporal Consistency in Vision-Language Models
Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Langu…
From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Pl…
Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Tr…
Machine Learning Assisted Inverse Design of Pixelated mmWave Patch Antennas
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workpl…
From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers
Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reco…
Runtime Action Interference for AI Control of AlphaStar in StarCraft II
AdaR: A Framework for Equipping LLMs with Adaptive Reasoning
LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platfo…
ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Ga…
Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-…
NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus
Concepts for Securing Agentic AI Coding and the Terok Environment
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM
GIM: Evaluating models via tasks that integrate multiple cognitive domains
Learning from the Test: Self-Referential Differential Testing for Deep RL Agents
Proxy reliance in large language model decisions is uncalibrated to predictive evidence
CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents
On the Role of Citations in Preference Data