Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, a…
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Secu…
SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow
PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
Xetrieval: Mechanistically Explaining Dense Retrieval
The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions vi…
OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster…
ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for …
When Does Persona Prompting Actually Help? A Retrieval and Metric Analysis of Expert Role…
TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evalu…
Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories
It`s All About Speed: AI`s Impact on Workflow in Music Production
Meta-Programming for Linear-time Temporal Answer Set Programming
Compass: Navigating Global Marine Lead Data Integration through Expert-Guided LLM Agent
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
From GPS Points to Travel Patterns: Flexible and Semantic Trajectory Generation with LLMs
Finding DoRI: Discovery of Retained Images in Diffusion Models
Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents
Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers
Reliable Reasoning with Large Language Models via Preference-Based Maximum Satisfiability
Conformal Certification of Reasoning Trace Prefixes
PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers
Enhancing Multi-Agent Communication through Attention Steering with Context Relevance
AgentSchool: An LLM-Powered Multi-Agent Simulation for Education
Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A P…
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientifi…
A comparative study of transformer-based embeddings for topic coherence
NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs