LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
MulFeRL: Enhancing Reinforcement Learning with Verbal Feedback in a Multi-turn Loop
Safety Alignment of LMs via Non-cooperative Games
MobiBench: Multi-Branch, Modular Benchmark for Mobile GUI Agents
A Unified Evaluation-Instructed Framework for Query-Dependent Prompt Optimization
ACON: Optimizing Context Compression for Long-horizon LLM Agents
Taming System Complexity: Demystifying Software Engineering Agents in Diagnosing Linux Ke…
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Miti…
Learning to Reduce Search Space for Generalizable Neural Routing Solver
Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbat…
Permissive Safety Through Trusted Inference: Verifiable Belief-Space Neural Safety Filter…
Modeling Depth Ambiguity: A Mixture-Density Representation for Flying-Point-Free Depth Es…
Monitoring Agentic Systems Before They're Reliable
Learning When to Translate for Multilingual Reasoning
Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Po…
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Diverge…
FOAM: Frequency and Operator Error-Based Adaptive Damping Method for Reducing Staleness-O…
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018…
Consistency Training while Mitigating Obfuscation via Rate Matching
Multilingual Idioms in Sentences and Conversations Across High-, Medium-, and Low-Resourc…
LALE: Lightweight-Transformer Architecture for Land-Cover Estimation
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
Why Do Time Series Models Need Long Context Windows?
The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue
RadioMaster: Multi-Agent System for Autonomous Radio Signal Generation
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise
Multilinguality of Large Language Models From a Structural Perspective
Construction of Historical Knowledge Graphs Based on BERT and Graph Neural Networks
Two-Fidelity Best-Action Identification for Stochastic Minimax Tree
Latent Collaboration in Multi-Agent Systems