Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs
Explorar
Noticias de IA
1057 elementos — filtrados, clasificados y sin duplicados
RiskBridge: Turning CVEs into Business-Aligned Patch Priorities
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
Explainable Attention-Guided Stacked Graph Neural Networks for Malware Detection
Membership Inference Attacks on Tokenizers of Large Language Models
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimi…
Hidden-State Privacy Has an Empty Middle
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-…
Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence
Megalodon cyberattack infects 5,500 GitHub open-source repositories with malware, researc…
The AI Era Is Creating a Bug Hunting Arms Race
BarrierSteer: LLM Safety via Learning Barrier Steering
TCAP: Tri-Component Attention Profiling for Unsupervised Backdoor Detection in MLLM Fine-…
AI Security Research Should Better Incentivize Defense Research
PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs
Security of LLM-generated Code: A Comparative Analysis
MemAudit: Post-hoc Auditing of Poisoned Agent Memory via Causal Attribution and Structura…
RAG-Pull: Turning Retrieval into a Code-Injection Channel via Invisible Unicode Perturbat…
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG
Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android M…
GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from D…
GenAI-Driven Threat Detection with Microsoft Security Copilot
Codec-Robust Attacks on Audio LLMs
The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Syst…
Content-Aware Attack Detection in LLM Agent Tool-Call Traffic: An Empirical Study of Feat…
Everyone is navigating AI security in real time — even Google
Hackers are learning to exploit chatbot ‘personalities’