Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Subliminal Learning is a LoRA Artifact
DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for D…
Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics
BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesi…
Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue
pcbGPT: Automatic PCB Schematic Synthesis from Natural Language Requirements
When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry
MViewRouter: Internalizing Geometric Equivariance via Multi-view Alternating Attention fo…
Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation …
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
Dive into Waves: Morlet Spectral Transformer for Cross-Subject Emotion Decoding from EEG
Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and…
Extending Causal Metamodeling to a non-Markovian Queue
Behavior-Invariant Task Representation Learning with Transformer-based World Models for O…
SORA: Free Second-Order Attacks in Fast Adversarial Training
Information-Theoretic Lower Bounds for Bit-Constrained Stochastic Optimization via a Redu…
The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Sh…
Demystifying the Optimal Fair Classifier in Multi-Class Classification
LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psycho…
SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answe…
Improving Visual Representation Alignment Generation with GRPO
Interpretable Policy Distillation for Power Grid Topology Control
CodeCytos: AI-assisted spatial molecular imaging analysis via code-augmented agent action…
PR2: Predictive Routing Replay for MoE-Based LLM Reinforcement Learning
Detector-Evasive LLM Paraphrasing via Constrained Policy Optimization
(HB-ARFM) History-Bootstrapped Flow Matching for Inverse Boiling Reconstruction
Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Lear…