OpenAI Baselines: ACKTR & A2C
Browse
AI News
25887 items — filtered, classified, deduplicated
More on Dota 2
Dota 2
Gathering human feedback
Better exploration with parameter noise
Proximal Policy Optimization
Robust adversarial inputs
Hindsight Experience Replay
Teacher–student curriculum learning
Faster physics in Python
Learning from human preferences
Learning to cooperate, compete, and communicate
UCB exploration via Q-ensembles
OpenAI Baselines: DQN
Robots that learn
Roboschool
Equivalence between policy gradients and soft Q-learning
Stochastic Neural Networks for hierarchical reinforcement learning
Unsupervised sentiment neuron
Spam detection in the physical world
Evolution strategies as a scalable alternative to reinforcement learning
One-shot imitation learning
Distill
Learning to communicate
Emergence of grounded compositional language in multi-agent populations
Prediction and control with temporal segment models
Third-person imitation learning
Attacking machine learning with adversarial examples
Adversarial attacks on neural network policies
Team update