Voluntary Collusion with Secret Tools in Competing LLM Agents
Explorar
Noticias de IA
30308 elementos — filtrados, clasificados y sin duplicados
When prompt perturbations break your A/B test: A valid statistical test for generative su…
LaneRoPE: Positional Encoding for Collaborative Parallel Reasoning and Generation
Soro: A Lightweight Foundation Model and Chatbot for Tajik
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
BIRDS: Characterizing and Understanding Biodiversity Impact of Large Language Model Servi…
Detect by Yourself: Self-Designing Agentic Workflows for Few-Shot Graph Anomaly Detection
DataClawBench: An Agent Benchmark for Exploratory Real-World Financial Data Analysis
Optimal LTLf Synthesis
Democratizing Large-Scale Re-Optimization with LLM-Guided Model Patches
Generalized Holographic Reduced Representations
Improving Requirements Classification with SMOTE-Tomek Preprocessing
Voice "Cloning" is Style Transfer
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Moni…
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptib…
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
Structured Agent Distillation for Large Language Model
Regression Language Models for Code
Object-Centric Vision Token Pruning for Vision Language Models
Detecting and Mitigating the Correct-Answer Extinction Window in Test-Time Reinforcement …
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Ja…
Teaching and Evaluating LLMs to Reason About Polymer Design Related Tasks
You Are in Control of Your State: Why Human Outcomes Are Controllable Through Causal Stat…
Hierarchical Prompt-Domain Control and Learning for Resource-Constrained Agentic Language…
The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntacti…
A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention M…
SkillGrad: Optimizing Agent Skills Like Gradient Descent
ECHO: Entropy-Confidence Hybrid Optimization for Test-Time Reinforcement Learning
Capture Timing-Attention of Events in Clinical Time Series