Explorar

Noticias de IA

37834 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Research & Papers
FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models
arXiv cs.AI Research & Papers
Plan Before Search: Search Agents Need Plan
arXiv cs.AI Research & Papers
Risk-Controlled Lean-as-Judge for Natural-Language Mathematical Reasoning
arXiv cs.AI Security & Safety
SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents
arXiv cs.AI Security & Safety
MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Conte…
arXiv cs.AI Research & Papers
CyberJurors: A Multi-Agent Simulation Task for E-Commerce Disputes Verdict
arXiv cs.AI Research & Papers
HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoni…
arXiv cs.AI Research & Papers
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
arXiv cs.AI Research & Papers
On the Learnability of Test-Time Adaptation: A Recovery Complexity Perspective
arXiv cs.AI Research & Papers
Measuring Progress Toward AGI: A Cognitive Framework
arXiv cs.AI Research & Papers
Efficient Post-training of LLMs for Code Generation With Offline Reinforcement Learning
arXiv cs.AI Research & Papers
From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constr…
arXiv cs.AI Research & Papers
Let Relations Speak: An End-to-End LLM-GNN Soft Prompt Framework for Fraud Detection
arXiv cs.AI Research & Papers
Integrated and Cross-Architecture Interpretation of LLM Reasoning
arXiv cs.AI Research & Papers
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language M…
arXiv cs.AI Research & Papers
Entropy-aware Masking for Masked Language Modeling
arXiv cs.AI Research & Papers
Improving Evaluation of Recombination-based Cartesian Genetic Programming
arXiv cs.AI Research & Papers
Refusal Before Decoding: Detecting and Exploiting Refusal Signals in Intermediate LLM Act…
arXiv cs.AI Research & Papers
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
arXiv cs.AI Research & Papers
Optimal LTLf Synthesis
arXiv cs.AI Research & Papers
Capture Timing-Attention of Events in Clinical Time Series
arXiv cs.AI Research & Papers
MUSE: Benchmarking Manufacturable, Functional, and Assemblable Text-to-CAD Generation
arXiv cs.AI Research & Papers
Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability
arXiv cs.AI Research & Papers
Continual Model Routing in Evolving Model Hubs
arXiv cs.AI Research & Papers
The Ethics of LLM Sandbox and Persona Dynamics
arXiv cs.AI Research & Papers
DREAM-R: Multimodal Speculative Reasoning with RL-Based Refined Drafting, Precise Verific…
arXiv cs.AI Research & Papers
TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-…
arXiv cs.AI Research & Papers
Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
arXiv cs.AI Research & Papers
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
arXiv cs.AI Research & Papers
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor