← All news

Variance reduction for policy gradient with action-dependent factorized baselines

Open the original source for the full article.

Read original at OpenAI News →