Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish…
Explorar
Noticias de IA
30329 elementos — filtrados, clasificados y sin duplicados
TCP-MCP: Landscape-Guided Co-Evolution of Prompts and Communication Topologies for Multi-…
Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs
A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Str…
Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Polic…
Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
A Query Engine for the Agents
Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resoluti…
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Un…
Learning after COVID-19 and the ICT career aspirations: Are students entering the AI era …
PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience …
Anomaly as Non-Conformity via Training-Free Graph Laplacian Energy Minimization
DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Esc…
Architecture-driven Shift: towards a lightweight selector for capturing the trends of log…
Human-AI Collaboration for Estimating Scientific Replicability
Hybrid Neural World Models
Case-Aware Medical Image Classification with Multimodal Knowledge Graphs and Reliability-…
Do Clinical Models Change Treatment Decisions?
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Syst…
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation…
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy…
Reasoning and Planning with Dynamically Changing Norms
Laguna M.1/XS.2 Technical Report
SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping
Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framewo…
Agyn: An Open-Source Platform for AI Agents with Scalable On-Demand Execution, Agent Defi…
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law