Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation
Explorar
Noticias de IA
22115 elementos — filtrados, clasificados y sin duplicados
Reconstructing Synthetic SDO/AIA 193 A EUV Images from He I 10830 A Observations with Dif…
TeamHerald@CHIPSAL 2026: Hate Speech Detection and Sentiment Analysis of Nepali Memes usi…
Not Just After One: Sleep-Inspired Replay Prevents Catastrophic Forgetting After Sequenti…
Segment-level Tree Search for Long Meeting Document Summarization
CoVEBench: Can Video Editing Models Handle Complex Instructions?
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Que…
Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering …
Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without …
HARBOR: A Harness Framework for Agentic Robot Reinforcement Learning
Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Meth…
"I understand your perspective": LLM Persuasion and Sycophancy through the Lens of Commun…
What's the Point? Spatial Grammar & Index Resolution for Sign Language Processing
SafeECGMatch: Calibration-Aware Joint Frequency and Time Space Semi-Supervised Learning f…
GIScholarBench: Benchmarking LLM Overconfidence in GIS Research
Voting Protocols as Coordination Mechanisms for Role-Constrained Multi-Agent Tutoring Sys…
An Information-Theoretic Definition for Open-Ended Learning
Self-Supervised Vision Transformers for CBCT-Based Detection of Temporomandibular Joint O…
Chiaroscuro Attention: Spending Compute in the Dark
"So There's a Catch-22 Here": How Early Adopters Who Build Multi-Agent LLM Systems Concep…
Causal Agent Replay: Counterfactual Attribution for LLM-Agent Failures
AgriGov: A Structured Multilingual Dataset Curation for Indian Government Schemes for Far…
Larch: Learned Query Optimization for Semantic Predicates
3D Oral Modelling with Improved Vertex Distribution Using Matching-Based Learning
The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models
Model Multiplicity for Adversarial Detection in Small Language Model Training on Edge Dev…
Beyond English benchmarks: clinical llm evaluation in Brazilian Portuguese
Does Persona Make LLMs K-pop Fans? A Pilot Study of LLM-Based Online Concert Audience Age…
Jas: AI-Paired Engineering as a Revival of N-Version Programming
DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment