A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
What a battery breakthrough reveals about AI and scientific memory
OpenAI’s sly mathematical breakthrough sends a chill through academia
Google's AI genome system evaluates every possible one-base change
A Stealth Startup Thinks It Just Hacked the Memory Shortage
Students who use AI generally score worse at school
GPT-6 Astra, Looped Transformers, and Hidden Reasoning
Google DeepMind Uses AI to Predict 9 Billion DNA Changes
How An AI math breakthrough ignited a controversy
[AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roug…
GIFT: Reconciling Post-Training Objectives via Variational Finite-Temperature Gibbs Initi…
RAPID: Reliability-Aware Pair Importance Distillation
PAGR: Proof-Carrying Algebraic-Geometric Retrieval: A Quiver-, Provenance-, and Sheaf-The…
PGP-Clinical-TimeKAN: Prior-Guided Joint Probabilistic Forecasting of Clinical Trajectori…
From Where to How: Continuous 4D Interaction Forecasting from Egocentric Video
Neptune: An AI model for Global Ocean Subseasonal Prediction
SynthRCT: Scalable Conditional Deformation Synthesis for Synthetic Repeat CT Generation
Leveraging Cardiac Imaging to Improve ECG-Based Detection of Chagas Disease in Resource-C…
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agen…
Damage-Aware Bandit Pruning for Vision and Language Transformers
Compiling VGDL into Causal Models
SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
Reducing Hallucinations in LLM-based Scientific Literature Analysis Using Peer Context Ou…
Deep belief networks are exact
Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors
Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Met…
Reasoning-Aware Compression: Identifying and Protecting Vulnerable Reasoning Circuits for…
What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record f…
Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses
The Normalization of Deviance in AI Development