CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error…
Fewer yet critical: Reducing Redundant Token Dependencies for Transformer-based Time Seri…
DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aw…
LOBERT: Generative AI Foundation Model for Limit Order Book Messages
CriticGen: Generation-Aware Evaluation as Actionable Feedback
VDiff-Bench: A Challenging Benchmark for Fine-Grained Image Difference Identification
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agen…
Damage-Aware Bandit Pruning for Vision and Language Transformers
Compiling VGDL into Causal Models
PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectori…
Reading Decoder Trajectories: Training-Free Counterfactual Query-Trajectory Reliability f…
The Normalization of Deviance in AI Development
STAR-Pro: Stage-Wise Token Adaptive Reduction with Progressive Refinement for Efficient L…
Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models
Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling
Earth System World Model for What-If Simulations: A Case Study for Terrestrial Ecosystems
The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persi…
SQLMorph: Query Mutation and Fine-Grained Metrics for Text-to-SQL Evaluation
Attention-Weighted Value Projection for KV-Cache Compression
Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscil…
Memory in Deep Time-Series Models
Hyperparameter Scaling Laws Across MoE Sparsity
Silent Revision: Measuring Undisclosed Change in the Safety Frameworks of Frontier AI Dev…
RAPID: Reliability-Aware Pair Importance Distillation
Parallelism Strategy Chaining for Fast Training Convergence
Omni Interaction Agent Technical Report
From the Fluency Fallacy to the Micro-to-Macro Validity Gap: Opportunities and Pitfalls o…
Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence…
TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection