Explorar

Noticias de IA

1057 elementos — filtrados, clasificados y sin duplicados

arXiv cs.AI Security & Safety
Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs
arXiv cs.AI Security & Safety
RiskBridge: Turning CVEs into Business-Aligned Patch Priorities
arXiv cs.AI Security & Safety
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
arXiv cs.AI Security & Safety
Explainable Attention-Guided Stacked Graph Neural Networks for Malware Detection
arXiv cs.AI Security & Safety
Membership Inference Attacks on Tokenizers of Large Language Models
arXiv cs.AI Security & Safety
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
arXiv cs.AI Security & Safety
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimi…
arXiv cs.AI Security & Safety
Hidden-State Privacy Has an Empty Middle
arXiv cs.AI Security & Safety
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-…
arXiv cs.AI Security & Safety
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
Mashable Security & Safety
Megalodon cyberattack infects 5,500 GitHub open-source repositories with malware, researc…
Wired AI Security & Safety
The AI Era Is Creating a Bug Hunting Arms Race
arXiv cs.AI Security & Safety
BarrierSteer: LLM Safety via Learning Barrier Steering
arXiv cs.AI Security & Safety
TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-…
arXiv cs.AI Security & Safety
AI Security Research Should Better Incentivize Defense Research
arXiv cs.AI Security & Safety
PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs
arXiv cs.AI Security & Safety
Security of LLM-generated Code: A Comparative Analysis
arXiv cs.AI Security & Safety
MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structura…
arXiv cs.AI Security & Safety
RAG-Pull: Turning Retrieval into a Code-Injection Channel via Invisible Unicode Perturbat…
arXiv cs.AI Security & Safety
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
arXiv cs.AI Security & Safety
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
arXiv cs.AI Security & Safety
Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android M…
arXiv cs.AI Security & Safety
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
arXiv cs.AI Security & Safety
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from D…
arXiv cs.AI Security & Safety
GenAI-Driven Threat Detection with Microsoft Security Copilot
arXiv cs.AI Security & Safety
Codec-Robust Attacks on Audio LLMs
arXiv cs.AI Security & Safety
The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Syst…
arXiv cs.AI Security & Safety
Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Feat…
TechCrunch AI Security & Safety
Everyone is navigating AI security in real time — even Google
The Verge AI Security & Safety
Hackers are learning to exploit chatbot ‘personalities’