Gathering human feedback
Browse
AI News
26304 items — filtered, classified, deduplicated
Better exploration with parameter noise
Proximal Policy Optimization
Robust adversarial inputs
Hindsight Experience Replay
Teacher–student curriculum learning
Faster physics in Python
Learning from human preferences
Learning to cooperate, compete, and communicate
UCB exploration via Q-ensembles
OpenAI Baselines: DQN
Robots that learn
Roboschool
Equivalence between policy gradients and soft Q-learning
Stochastic Neural Networks for hierarchical reinforcement learning
Unsupervised sentiment neuron
Spam detection in the physical world
Evolution strategies as a scalable alternative to reinforcement learning
One-shot imitation learning
Distill
Learning to communicate
Emergence of grounded compositional language in multi-agent populations
Prediction and control with temporal segment models
Third-person imitation learning
Attacking machine learning with adversarial examples
Adversarial attacks on neural network policies
Team update
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other…
Faulty reward functions in the wild
Universe