FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Plan Before Search: Search Agents Need Plan
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Conte…
CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoni…
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective
Measuring Progress Toward AGI: A Cognitive Framework
Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning
From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constr…
Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection
Integrated and Cross-Architecture Interpretation of LLM Reasoning
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language M…
Entropy-aware Masking for Masked Language Modeling
Improving Evaluation of Recombination-based Cartesian Genetic Programming
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Act…
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Optimal LTLf Synthesis
Capture Timing-Attention of Events in Clinical Time Series
MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
Continual Model Routing in Evolving Model Hubs
The Ethics of LLM Sandbox and Persona Dynamics
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verific…
TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-…
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor