Learning from human preferences
Explorar
Noticias de IA
29327 elementos — filtrados, clasificados y sin duplicados
Learning to cooperate, compete, and communicate
UCB exploration via Q-ensembles
OpenAI Baselines: DQN
Robots that learn
Roboschool
Equivalence between policy gradients and soft Q-learning
Stochastic Neural Networks for hierarchical reinforcement learning
Unsupervised sentiment neuron
Spam detection in the physical world
Evolution strategies as a scalable alternative to reinforcement learning
One-shot imitation learning
Distill
Learning to communicate
Emergence of grounded compositional language in multi-agent populations
Prediction and control with temporal segment models
Third-person imitation learning
Attacking machine learning with adversarial examples
Adversarial attacks on neural network policies
Team update
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other…
Faulty reward functions in the wild
Universe
#Exploration: A study of count-based exploration for deep reinforcement learning
OpenAI and Microsoft
On the quantitative analysis of decoder-based generative models
A connection between generative adversarial networks, inverse reinforcement learning, and…
RL²: Fast reinforcement learning via slow reinforcement learning
Variational lossy autoencoder
Extensions and limitations of the neural GPU