Conf-Gen: Conformal Uncertainty Quantification for Generative Models
Explorar
Noticias de IA
30308 elementos — filtrados, clasificados y sin duplicados
EvA: An Evidence-First Audio Understanding Paradigm for LALMs
FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchma…
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Mode…
Reducing Political Manipulation with Consistency Training
Label-Free Reinforcement Learning via Cross-Model Entropy
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
OISD: On-Policy Internal Self-Distillation of Language Models
unix-ctf: Procedural Environments for Unix-Competence Reinforcement Learning
Multi-Resolution End-to-End Deep Neural Network for Optimizing Latency-Accuracy Tradeoff …
SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medicat…
Parallax: Parameterized Local Linear Attention for Language Modeling
Finding DoRI: Discovery of Retained Images in Diffusion Models
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control
Domain-Informed Representation for Evolutionary Sieving in Integral and Module Lattices
Beyond MSE: Improving Precipitation Nowcasting with Multi-Quantile Regression
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustaina…
Behavior-Aware Auxiliary Corrections for Off-Policy Temporal-Difference Prediction
Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild
The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Ad…
GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents
Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured …
From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaborati…
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
VikingMem: A Memory Base Management System for Stateful LLM-based Applications
Beyond Attack Success Rate: Temporal Logit Observability for LLM Safety Failures
Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Ar…
ReasonOps: Operator Segmentation for LLM Reasoning Traces
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training