C3-Bench: A Context-Aware Change Captioning Benchmark
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
HG-Bench: A Benchmark for Multi-Page Handwritten Answer-Region Grounding in Automated Hom…
The impact of artificial intelligence on enterprise software user roles
Evaluating LLMs on Real-World Software Performance Optimization
An Approach for a Supporting Multi-LLM System for Automated Certification Based on the Ge…
Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automate…
TL++: Accuracy and Privacy Preserving Traversal Learning for Distributed Intelligent Syst…
Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents
Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation
Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimization
Steering Vision-Language Models with Joint Sparse Autoencoders
Gradient-based inverse lithography for EUV masks via the waveguide method and a physics-i…
Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Mo…
Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM…
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources
Edges Before Embeddings: A Confidence-Aware Blur Gate for Vision-Language Pipelines
Semantic Consistency Policy Optimization for Reinforcement Learning of LLM Agents
AutoRelAnnotator: Calibrated Model Cascades for Cost-Efficient Relevance Evaluation in Sp…
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are…
Enhancing Brain MRI Anomaly Detection and Reasoning with ROI Rethink and Synthetic Data
Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts
Multi-Agent Goal Recognition with Team- and Goal-Conditioned Reinforcement Learning and F…
Weave of Formal Thought
SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversati…
FORCE: Efficient VLA Reinforcement Fine-Tuning via Value-Calibrated Warm-up and Self-Dist…
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agen…
Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining
Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity