chore: import upstream snapshot with attribution

This commit is contained in:
wehub-resource-sync
2026-07-13 12:19:01 +08:00
commit 3b90d1192f
2172 changed files with 594509 additions and 0 deletions
+4
View File
@@ -0,0 +1,4 @@
{
"<h1><a href=\"https://nn.labml.ai/rl/ppo/index.html\">Proximal Policy Optimization - PPO</a></h1>\n<p>This is a <a href=\"https://pytorch.org\">PyTorch</a> implementation of <a href=\"https://arxiv.org/abs/1707.06347\">Proximal Policy Optimization - PPO</a>.</p>\n<p>PPO is a policy gradient method for reinforcement learning. Simple policy gradient methods one do a single gradient update per sample (or a set of samples). Doing multiple gradient steps for a singe sample causes problems because the policy deviates too much producing a bad policy. PPO lets us do multiple gradient updates per sample by trying to keep the policy close to the policy that was used to sample data. It does so by clipping gradient flow if the updated policy is not close to the policy used to sample the data.</p>\n<p>You can find an experiment that uses it <a href=\"https://nn.labml.ai/rl/ppo/experiment.html\">here</a>. The experiment uses <a href=\"https://nn.labml.ai/rl/ppo/gae.html\">Generalized Advantage Estimation</a>.</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/rl/ppo/experiment.ipynb\"><span translate=no>_^_0_^_</span></a> </p>\n": "<h1><a href=\"https://nn.labml.ai/rl/ppo/index.html\">\u8fd1\u63a5\u30dd\u30ea\u30b7\u30fc\u6700\u9069\u5316-PPO</a></h1>\n<p><a href=\"https://arxiv.org/abs/1707.06347\">\u3053\u308c\u306f\u8fd1\u63a5\u30dd\u30ea\u30b7\u30fc\u6700\u9069\u5316</a>\uff08PPO\uff09<a href=\"https://pytorch.org\">\u306ePyTorch\u5b9f\u88c5\u3067\u3059</a>\u3002</p>\n<p>PPO\u306f\u5f37\u5316\u5b66\u7fd2\u306e\u30dd\u30ea\u30b7\u30fc\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u6cd5\u3067\u3059\u3002\u5358\u7d14\u306a\u30dd\u30ea\u30b7\u30fc\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u30e1\u30bd\u30c3\u30c9\u3067\u306f\u3001\u30b5\u30f3\u30d7\u30eb\uff08\u307e\u305f\u306f\u30b5\u30f3\u30d7\u30eb\u306e\u30bb\u30c3\u30c8\uff09\u3054\u3068\u306b1\u3064\u306e\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u66f4\u65b0\u3092\u884c\u3044\u307e\u3059\u30021\u3064\u306e\u30b5\u30f3\u30d7\u30eb\u306b\u5bfe\u3057\u3066\u8907\u6570\u306e\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u30b9\u30c6\u30c3\u30d7\u3092\u5b9f\u884c\u3059\u308b\u3068\u3001\u30dd\u30ea\u30b7\u30fc\u306e\u504f\u5dee\u304c\u5927\u304d\u3059\u304e\u3066\u4e0d\u9069\u5207\u306a\u30dd\u30ea\u30b7\u30fc\u304c\u751f\u6210\u3055\u308c\u308b\u305f\u3081\u3001\u554f\u984c\u304c\u767a\u751f\u3057\u307e\u3059\u3002PPO \u3067\u306f\u3001\u30dd\u30ea\u30b7\u30fc\u3092\u30c7\u30fc\u30bf\u306e\u30b5\u30f3\u30d7\u30ea\u30f3\u30b0\u306b\u4f7f\u7528\u3057\u305f\u30dd\u30ea\u30b7\u30fc\u306b\u8fd1\u3044\u72b6\u614b\u306b\u4fdd\u3064\u3053\u3068\u3067\u3001\u30b5\u30f3\u30d7\u30eb\u3054\u3068\u306b\u8907\u6570\u306e\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u66f4\u65b0\u3092\u884c\u3046\u3053\u3068\u304c\u3067\u304d\u307e\u3059\u3002\u66f4\u65b0\u3055\u308c\u305f\u30dd\u30ea\u30b7\u30fc\u304c\u30c7\u30fc\u30bf\u306e\u30b5\u30f3\u30d7\u30ea\u30f3\u30b0\u306b\u4f7f\u7528\u3055\u308c\u305f\u30dd\u30ea\u30b7\u30fc\u306b\u5408\u308f\u306a\u3044\u5834\u5408\u306f\u3001\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u30d5\u30ed\u30fc\u3092\u30af\u30ea\u30c3\u30d4\u30f3\u30b0\u3057\u3066\u66f4\u65b0\u3057\u307e\u3059</p>\u3002\n<p><a href=\"https://nn.labml.ai/rl/ppo/experiment.html\">\u3053\u308c\u3092\u4f7f\u3063\u305f\u5b9f\u9a13\u306f\u3053\u3061\u3089\u304b\u3089\u3054\u89a7\u3044\u305f\u3060\u3051\u307e\u3059</a>\u3002\u3053\u306e\u5b9f\u9a13\u3067\u306f\u3001<a href=\"https://nn.labml.ai/rl/ppo/gae.html\">\u4e00\u822c\u5316\u30a2\u30c9\u30d0\u30f3\u30c6\u30fc\u30b8\u63a8\u5b9a\u3092\u4f7f\u7528\u3057\u3066\u3044\u307e\u3059</a></p>\u3002\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/rl/ppo/experiment.ipynb\"><span translate=no>_^_0_^_</span></a></p>\n",
"Proximal Policy Optimization - PPO": "\u8fd1\u63a5\u30dd\u30ea\u30b7\u30fc\u6700\u9069\u5316-PPO"
}