Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
Explorar
Noticias de IA
30934 elementos — filtrados, clasificados y sin duplicados
Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe U…
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Are Heterogeneous Graph Neural Networks Truly Effective for Node Classification? A Causal…
Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering
SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pr…
Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD
A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Lear…
Unsupervised Deep Learning for Inverse Problems in Computed Tomography
A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe L…
Verbalizable Representations Form a Global Workspace in Language Models
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream
BrainPilot: Automating Brain Discovery with Agentic Research
Decoupled Alignment for Robust Plug-and-Play Adaptation
Rethinking Quantum Continual Learning with Quantum Fisher Information
Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime St…
From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-…
A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance
MirrorCode: AI can rebuild entire programs from behavior alone
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing…
Recursive Harness Self-Improvement
Can We Trust Item Response Theory for AI Evaluation?
Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personal…
Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resou…
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Mul…
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthe…