chore: import upstream snapshot with attribution

This commit is contained in:
wehub-resource-sync
2026-07-13 12:19:01 +08:00
commit 3b90d1192f
2172 changed files with 594509 additions and 0 deletions
@@ -0,0 +1,29 @@
{
"<h1>Transformer XL</h1>\n<p>This is an implementation of <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n<p>Transformer has a limited attention span, equal to the length of the sequence trained in parallel. All these positions have a fixed positional encoding. Transformer XL increases this attention span by letting each of the positions pay attention to precalculated past embeddings. For instance if the context length is <span translate=no>_^_0_^_</span>, it will keep the embeddings of all layers for previous batch of length <span translate=no>_^_1_^_</span> and feed them to current step. If we use fixed-positional encodings these pre-calculated embeddings will have the same positions as the current context. They introduce relative positional encoding, where the positional encodings are introduced at the attention calculation.</p>\n<p>Annotated implementation of relative multi-headed attention is in <a href=\"relative_mha.html\"><span translate=no>_^_2_^_</span></a>.</p>\n<p>Here&#x27;s <a href=\"experiment.html\">the training code</a> and a notebook for training a transformer XL model on Tiny Shakespeare dataset.</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a></p>\n": "<h1>\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL</h1>\n<p><a href=\"https://pytorch.org\">\u3053\u308c\u306f PyTorch \u306e <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: \u56fa\u5b9a\u9577\u306e\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u3092\u8d85\u3048\u305f\u6ce8\u610f\u6df1\u3044\u8a00\u8a9e\u30e2\u30c7\u30eb\u306e\u5b9f\u88c5\u3067\u3059</a>\u3002</a></p>\n<p>Transformer \u306e\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30b9\u30d1\u30f3\u306f\u3001\u4e26\u884c\u3057\u3066\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3055\u308c\u305f\u30b7\u30fc\u30b1\u30f3\u30b9\u306e\u9577\u3055\u3068\u540c\u3058\u304f\u3089\u3044\u306e\u5236\u9650\u304c\u3042\u308a\u307e\u3059\u3002\u3053\u308c\u3089\u306e\u4f4d\u7f6e\u306f\u3059\u3079\u3066\u56fa\u5b9a\u3055\u308c\u305f\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u306b\u306a\u3063\u3066\u3044\u307e\u3059\u3002Transformer XL\u306f\u3001\u4e8b\u524d\u306b\u8a08\u7b97\u3055\u308c\u305f\u904e\u53bb\u306e\u57cb\u3081\u8fbc\u307f\u306b\u5404\u30dd\u30b8\u30b7\u30e7\u30f3\u306b\u6ce8\u76ee\u3055\u305b\u308b\u3053\u3068\u3067\u3001\u3053\u306e\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30b9\u30d1\u30f3\u3092\u5897\u3084\u3057\u307e\u3059\u3002\u305f\u3068\u3048\u3070\u3001\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u306e\u9577\u3055\u304c\u306e\u5834\u5408<span translate=no>_^_0_^_</span>\u3001<span translate=no>_^_1_^_</span>\u524d\u306e\u30d0\u30c3\u30c1\u306e\u9577\u3055\u306e\u3059\u3079\u3066\u306e\u30ec\u30a4\u30e4\u30fc\u306e\u57cb\u3081\u8fbc\u307f\u3092\u4fdd\u6301\u3057\u3001\u305d\u308c\u3089\u3092\u73fe\u5728\u306e\u30b9\u30c6\u30c3\u30d7\u306b\u9001\u308a\u307e\u3059\u3002\u56fa\u5b9a\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u3092\u4f7f\u7528\u3059\u308b\u3068\u3001\u3053\u308c\u3089\u306e\u4e8b\u524d\u306b\u8a08\u7b97\u3055\u308c\u305f\u57cb\u3081\u8fbc\u307f\u306f\u73fe\u5728\u306e\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u3068\u540c\u3058\u4f4d\u7f6e\u306b\u306a\u308a\u307e\u3059\u3002\u76f8\u5bfe\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u304c\u5c0e\u5165\u3055\u308c\u3001\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u8a08\u7b97\u6642\u306b\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u304c\u5c0e\u5165\u3055\u308c\u307e\u3059</p>\u3002\n<p>\u76f8\u5bfe\u7684\u591a\u9762\u7684\u6ce8\u610f\u306e\u6ce8\u91c8\u4ed8\u304d\u5b9f\u88c5\u304c\u5c0e\u5165\u3055\u308c\u307e\u3057\u305f\u3002<a href=\"relative_mha.html\"><span translate=no>_^_2_^_</span></a></p>\n<p>Tiny <a href=\"experiment.html\">Shakespeare\u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u3067\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fcXL\u30e2\u30c7\u30eb\u3092\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3059\u308b\u305f\u3081\u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u30b3\u30fc\u30c9\u3068\u30ce\u30fc\u30c8\u30d6\u30c3\u30af\u3067\u3059</a>\u3002</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a></p>\n",
"<h2>Transformer XL Layer</h2>\n<p>The transformer XL model comprises of a number of these layers.</p>\n": "<h2>\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL \u30ec\u30a4\u30e4\u30fc</h2>\n<p>\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fcXL\u30e2\u30c7\u30eb\u306f\u3001\u3053\u308c\u3089\u306e\u30ec\u30a4\u30e4\u30fc\u3092\u591a\u6570\u5099\u3048\u3066\u3044\u307e\u3059\u3002</p>\n",
"<h2>Transformer XL Model</h2>\n<p>This consists of multiple transformer XL layers</p>\n": "<h2>\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL \u30e2\u30c7\u30eb</h2>\n<p>\u3053\u308c\u306f\u8907\u6570\u306e\u30c8\u30e9\u30f3\u30b9XL\u5c64\u3067\u69cb\u6210\u3055\u308c\u3066\u3044\u307e\u3059</p>\n",
"<p> </p>\n": "<p></p>\n",
"<p>Add the attention results </p>\n": "<p>\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u7d50\u679c\u3092\u8ffd\u52a0</p>\n",
"<p>Add the feed-forward results back </p>\n": "<p>\u30d5\u30a3\u30fc\u30c9\u30d5\u30a9\u30ef\u30fc\u30c9\u306e\u7d50\u679c\u3092\u8ffd\u52a0\u3057\u76f4\u3059</p>\n",
"<p>Add to the list of feature vectors </p>\n": "<p>\u7279\u5fb4\u30d9\u30af\u30c8\u30eb\u306e\u30ea\u30b9\u30c8\u306b\u8ffd\u52a0</p>\n",
"<p>Attention </p>\n": "<p>\u6ce8\u610f</p>\n",
"<p>Concatenate with <span translate=no>_^_0_^_</span> </p>\n": "<p>\u3068\u9023\u7d50 <span translate=no>_^_0_^_</span></p>\n",
"<p>Final normalization layer </p>\n": "<p>\u6700\u7d42\u6b63\u898f\u5316\u30ec\u30a4\u30e4\u30fc</p>\n",
"<p>Finally, normalize the vectors </p>\n": "<p>\u6700\u5f8c\u306b\u3001\u30d9\u30af\u30c8\u30eb\u3092\u6b63\u898f\u5316\u3057\u307e\u3059\u3002</p>\n",
"<p>If there is memory </p>\n": "<p>\u30e1\u30e2\u30ea\u304c\u3042\u308c\u3070</p>\n",
"<p>Ignore if there is no memory </p>\n": "<p>\u30e1\u30e2\u30ea\u304c\u306a\u3044\u5834\u5408\u306f\u7121\u8996\u3057\u3066\u304f\u3060\u3055\u3044</p>\n",
"<p>List to store token level feature vectors, which will become the memories for the next sequential batch. </p>\n": "<p>\u6b21\u306e\u30b7\u30fc\u30b1\u30f3\u30b7\u30e3\u30eb\u30d0\u30c3\u30c1\u306e\u30e1\u30e2\u30ea\u3068\u306a\u308b\u30c8\u30fc\u30af\u30f3\u30ec\u30d9\u30eb\u306e\u7279\u5fb4\u30d9\u30af\u30c8\u30eb\u3092\u683c\u7d0d\u3059\u308b\u30ea\u30b9\u30c8\u3002</p>\n",
"<p>Make copies of the transformer layer </p>\n": "<p>\u30c8\u30e9\u30f3\u30b9\u30ec\u30a4\u30e4\u30fc\u306e\u30b3\u30d4\u30fc\u3092\u4f5c\u6210</p>\n",
"<p>Memory </p>\n": "<p>\u30e1\u30e2\u30ea\u30fc</p>\n",
"<p>Normalize for feed-forward </p>\n": "<p>\u30d5\u30a3\u30fc\u30c9\u30d5\u30a9\u30ef\u30fc\u30c9\u7528\u306b\u6b63\u898f\u5316</p>\n",
"<p>Normalize it </p>\n": "<p>\u6b63\u898f\u5316\u3057\u3066\u304f\u3060\u3055\u3044</p>\n",
"<p>Normalize the vectors before doing self attention </p>\n": "<p>\u30bb\u30eb\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u3092\u884c\u3046\u524d\u306b\u30d9\u30af\u30c8\u30eb\u3092\u6b63\u898f\u5316\u3057\u3066\u304f\u3060\u3055\u3044</p>\n",
"<p>Pass through the feed-forward network </p>\n": "<p>\u30d5\u30a3\u30fc\u30c9\u30d5\u30a9\u30ef\u30fc\u30c9\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u3092\u901a\u904e</p>\n",
"<p>Run through each transformer layer </p>\n": "<p>\u5404\u5909\u5727\u5668\u5c64\u306b\u901a\u3059</p>\n",
"<p>Run through the transformer XL layer </p>\n": "<p>\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fcXL\u30ec\u30a4\u30e4\u30fc\u3092\u901a\u3059</p>\n",
"<ul><li><span translate=no>_^_0_^_</span> is a tensor of the token embeddings vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is a list of tensors of the past token level feature vectors of shape <span translate=no>_^_3_^_</span> for each layer </li>\n<li><span translate=no>_^_4_^_</span> is the masking matrix</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u30c8\u30fc\u30af\u30f3\u57cb\u3081\u8fbc\u307f\u306e\u5f62\u72b6\u30d9\u30af\u30c8\u30eb\u306e\u30c6\u30f3\u30bd\u30eb\u3067\u3059 <span translate=no>_^_1_^_</span></li>\n<li><span translate=no>_^_2_^_</span>\u904e\u53bb\u306e\u30c8\u30fc\u30af\u30f3\u30ec\u30d9\u30eb\u306e\u30c6\u30f3\u30bd\u30eb\u3001<span translate=no>_^_3_^_</span>\u5404\u30ec\u30a4\u30e4\u30fc\u306e\u5f62\u72b6\u30d9\u30af\u30c8\u30eb\u306e\u30ea\u30b9\u30c8\u3067\u3059</li>\n<li><span translate=no>_^_4_^_</span>\u306f\u30de\u30b9\u30ad\u30f3\u30b0\u30de\u30c8\u30ea\u30c3\u30af\u30b9\u3067\u3059</li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is a tensor of the token level feature vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is a tensor of the past token level feature vectors of shape <span translate=no>_^_3_^_</span> </li>\n<li><span translate=no>_^_4_^_</span> is a matrix of shape <span translate=no>_^_5_^_</span> or <span translate=no>_^_6_^_</span>. <span translate=no>_^_7_^_</span> is true if token at <span translate=no>_^_8_^_</span> can see token at <span translate=no>_^_9_^_</span>.</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u30c8\u30fc\u30af\u30f3\u30ec\u30d9\u30eb\u306e\u5f62\u72b6\u30d9\u30af\u30c8\u30eb\u306e\u30c6\u30f3\u30bd\u30eb\u3067\u3059 <span translate=no>_^_1_^_</span></li>\n<li><span translate=no>_^_2_^_</span>\u904e\u53bb\u306e\u30c8\u30fc\u30af\u30f3\u30ec\u30d9\u30eb\u306e\u5f62\u72b6\u30d9\u30af\u30c8\u30eb\u306e\u30c6\u30f3\u30bd\u30eb\u3067\u3059 <span translate=no>_^_3_^_</span></li>\n<li><span translate=no>_^_4_^_</span><span translate=no>_^_5_^_</span><span translate=no>_^_6_^_</span>\u306f\u5f62\u72b6\u306e\u30de\u30c8\u30ea\u30c3\u30af\u30b9\u304b<span translate=no>_^_7_^_</span>\u30c8\u30fc\u30af\u30f3 at \u304c at <span translate=no>_^_8_^_</span> \u306e\u30c8\u30fc\u30af\u30f3\u3092\u53c2\u7167\u3067\u304d\u308b\u5834\u5408\u306f true <span translate=no>_^_9_^_</span> \u306b\u306a\u308a\u307e\u3059\u3002</li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the token embedding size </li>\n<li><span translate=no>_^_1_^_</span> is the <a href=\"relative_mha.html\">self attention module</a> </li>\n<li><span translate=no>_^_2_^_</span> is the feed forward module </li>\n<li><span translate=no>_^_3_^_</span> is the probability of dropping out after self attention and FFN</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u30c8\u30fc\u30af\u30f3\u306e\u57cb\u3081\u8fbc\u307f\u30b5\u30a4\u30ba\u3067\u3059</li>\n<li><span translate=no>_^_1_^_</span><a href=\"relative_mha.html\">\u30bb\u30eb\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30e2\u30b8\u30e5\u30fc\u30eb\u3067\u3059</a></li>\n<li><span translate=no>_^_2_^_</span>\u30d5\u30a3\u30fc\u30c9\u30d5\u30a9\u30ef\u30fc\u30c9\u30e2\u30b8\u30e5\u30fc\u30eb\u3067\u3059</li>\n<li><span translate=no>_^_3_^_</span>\u30bb\u30eb\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u3068FFN\u306e\u5f8c\u306b\u8131\u843d\u3059\u308b\u78ba\u7387\u3067\u3059</li></ul>\n",
"Documented implementation with explanations of a Transformer-XL model.": "Transformer-XL \u30e2\u30c7\u30eb\u306e\u8aac\u660e\u3092\u542b\u3080\u5b9f\u88c5\u304c\u6587\u66f8\u5316\u3055\u308c\u3066\u3044\u307e\u3059\u3002",
"Transformer XL": "\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL"
}
File diff suppressed because one or more lines are too long
@@ -0,0 +1,29 @@
{
"<h1>Transformer XL</h1>\n<p>This is an implementation of <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n<p>Transformer has a limited attention span, equal to the length of the sequence trained in parallel. All these positions have a fixed positional encoding. Transformer XL increases this attention span by letting each of the positions pay attention to precalculated past embeddings. For instance if the context length is <span translate=no>_^_0_^_</span>, it will keep the embeddings of all layers for previous batch of length <span translate=no>_^_1_^_</span> and feed them to current step. If we use fixed-positional encodings these pre-calculated embeddings will have the same positions as the current context. They introduce relative positional encoding, where the positional encodings are introduced at the attention calculation.</p>\n<p>Annotated implementation of relative multi-headed attention is in <a href=\"relative_mha.html\"><span translate=no>_^_2_^_</span></a>.</p>\n<p>Here&#x27;s <a href=\"experiment.html\">the training code</a> and a notebook for training a transformer XL model on Tiny Shakespeare dataset.</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a></p>\n": "<h1>\u53d8\u538b\u5668 XL</h1>\n<p>\u8fd9\u662f <a href=\"https://pytorch.org\">PyTorch \u4e2d Transfor</a> <a href=\"https://arxiv.org/abs/1901.02860\">mer-XL\uff1a\u8d85\u8d8a\u56fa\u5b9a\u957f\u5ea6\u4e0a\u4e0b\u6587\u7684\u4e13\u5fc3\u8bed\u8a00\u6a21\u578b</a>\u7684\u5b9e\u73b0\u3002</p>\n<p>Transformer \u7684\u6ce8\u610f\u529b\u8de8\u5ea6\u6709\u9650\uff0c\u7b49\u4e8e\u5e76\u884c\u8bad\u7ec3\u5e8f\u5217\u7684\u957f\u5ea6\u3002\u6240\u6709\u8fd9\u4e9b\u4f4d\u7f6e\u90fd\u6709\u56fa\u5b9a\u7684\u4f4d\u7f6e\u7f16\u7801\u3002Transformer XL \u901a\u8fc7\u8ba9\u6bcf\u4e2a\u4f4d\u7f6e\u5173\u6ce8\u8fc7\u53bb\u9884\u5148\u8ba1\u7b97\u7684\u5d4c\u5165\u6b21\u6570\uff0c\u4ece\u800c\u5ef6\u957f\u4e86\u8fd9\u79cd\u6ce8\u610f\u529b\u8de8\u5ea6\u3002\u4f8b\u5982\uff0c\u5982\u679c\u4e0a\u4e0b\u6587\u957f\u5ea6\u4e3a<span translate=no>_^_0_^_</span>\uff0c\u5b83\u5c06\u4fdd\u7559\u524d\u4e00\u6279\u957f\u5ea6\u7684\u6240\u6709\u5c42\u7684\u5d4c\u5165<span translate=no>_^_1_^_</span>\u5e76\u5c06\u5176\u9988\u9001\u5230\u5f53\u524d\u6b65\u9aa4\u3002\u5982\u679c\u6211\u4eec\u4f7f\u7528\u56fa\u5b9a\u4f4d\u7f6e\u7f16\u7801\uff0c\u8fd9\u4e9b\u9884\u5148\u8ba1\u7b97\u7684\u5d4c\u5165\u5c06\u4e0e\u5f53\u524d\u4e0a\u4e0b\u6587\u5177\u6709\u76f8\u540c\u7684\u4f4d\u7f6e\u3002\u5b83\u4eec\u5f15\u5165\u4e86\u76f8\u5bf9\u4f4d\u7f6e\u7f16\u7801\uff0c\u5176\u4e2d\u4f4d\u7f6e\u7f16\u7801\u662f\u5728\u6ce8\u610f\u529b\u8ba1\u7b97\u65f6\u5f15\u5165\u7684\u3002</p>\n<p>\u76f8\u5bf9\u591a\u5934\u6ce8\u610f\u529b\u7684\u5e26\u6ce8\u91ca\u7684\u5b9e\u73b0\u5df2\u7ecf\u5f00\u59cb<a href=\"relative_mha.html\"><span translate=no>_^_2_^_</span></a>\u4e86\u3002</p>\n<p>\u8fd9\u662f\u7528\u4e8e<a href=\"experiment.html\">\u5728 Tiny Shakespeare \u6570\u636e\u96c6\u4e0a\u8bad\u7ec3 transformer XL \u6a21\u578b\u7684\u8bad\u7ec3\u4ee3\u7801</a>\u548c\u7b14\u8bb0\u672c\u3002</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a></p>\n",
"<h2>Transformer XL Layer</h2>\n<p>The transformer XL model comprises of a number of these layers.</p>\n": "<h2>\u53d8\u538b\u5668 XL \u5c42</h2>\n<p>\u53d8\u538b\u5668 XL \u6a21\u578b\u7531\u8bb8\u591a\u8fd9\u6837\u7684\u5c42\u7ec4\u6210\u3002</p>\n",
"<h2>Transformer XL Model</h2>\n<p>This consists of multiple transformer XL layers</p>\n": "<h2>\u53d8\u538b\u5668 XL \u578b\u53f7</h2>\n<p>\u5b83\u7531\u591a\u4e2a\u53d8\u538b\u5668 XL \u5c42\u7ec4\u6210</p>\n",
"<p> </p>\n": "<p></p>\n",
"<p>Add the attention results </p>\n": "<p>\u6dfb\u52a0\u5173\u6ce8\u7ed3\u679c</p>\n",
"<p>Add the feed-forward results back </p>\n": "<p>\u5c06\u524d\u9988\u7ed3\u679c\u6dfb\u52a0\u56de\u6765</p>\n",
"<p>Add to the list of feature vectors </p>\n": "<p>\u6dfb\u52a0\u5230\u7279\u5f81\u5411\u91cf\u5217\u8868\u4e2d</p>\n",
"<p>Attention </p>\n": "<p>\u6ce8\u610f</p>\n",
"<p>Concatenate with <span translate=no>_^_0_^_</span> </p>\n": "<p>\u8fde\u63a5\u4e0e<span translate=no>_^_0_^_</span></p>\n",
"<p>Final normalization layer </p>\n": "<p>\u6700\u7ec8\u5f52\u4e00\u5316\u5c42</p>\n",
"<p>Finally, normalize the vectors </p>\n": "<p>\u6700\u540e\uff0c\u5bf9\u5411\u91cf\u8fdb\u884c\u5f52\u4e00\u5316</p>\n",
"<p>If there is memory </p>\n": "<p>\u5982\u679c\u6709\u8bb0\u5fc6</p>\n",
"<p>Ignore if there is no memory </p>\n": "<p>\u5982\u679c\u6ca1\u6709\u5185\u5b58\uff0c\u5219\u5ffd\u7565</p>\n",
"<p>List to store token level feature vectors, which will become the memories for the next sequential batch. </p>\n": "<p>\u7528\u4e8e\u5b58\u50a8\u4ee4\u724c\u7ea7\u7279\u5f81\u5411\u91cf\u7684\u5217\u8868\uff0c\u8fd9\u4e9b\u5411\u91cf\u5c06\u6210\u4e3a\u4e0b\u4e00\u4e2a\u8fde\u7eed\u6279\u6b21\u7684\u8bb0\u5fc6\u3002</p>\n",
"<p>Make copies of the transformer layer </p>\n": "<p>\u5236\u4f5c\u53d8\u538b\u5668\u5c42\u7684\u526f\u672c</p>\n",
"<p>Memory </p>\n": "<p>\u8bb0\u5fc6</p>\n",
"<p>Normalize for feed-forward </p>\n": "<p>\u6807\u51c6\u5316\u4ee5\u8fdb\u884c\u524d\u9988</p>\n",
"<p>Normalize it </p>\n": "<p>\u89c4\u8303\u5316\u5b83</p>\n",
"<p>Normalize the vectors before doing self attention </p>\n": "<p>\u5728\u8fdb\u884c\u81ea\u6211\u6ce8\u610f\u4e4b\u524d\u5bf9\u5411\u91cf\u8fdb\u884c\u5f52\u4e00\u5316</p>\n",
"<p>Pass through the feed-forward network </p>\n": "<p>\u901a\u8fc7\u524d\u9988\u7f51\u7edc</p>\n",
"<p>Run through each transformer layer </p>\n": "<p>\u7a7f\u8fc7\u6bcf\u4e2a\u53d8\u538b\u5668\u5c42</p>\n",
"<p>Run through the transformer XL layer </p>\n": "<p>\u7a7f\u8fc7\u53d8\u538b\u5668 XL \u5c42</p>\n",
"<ul><li><span translate=no>_^_0_^_</span> is a tensor of the token embeddings vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is a list of tensors of the past token level feature vectors of shape <span translate=no>_^_3_^_</span> for each layer </li>\n<li><span translate=no>_^_4_^_</span> is the masking matrix</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u662f\u5d4c\u5165\u5f62\u72b6\u5411\u91cf\u7684\u4ee4\u724c\u7684\u5f20\u91cf<span translate=no>_^_1_^_</span></li>\n<li><span translate=no>_^_2_^_</span>\u662f\u8fc7\u53bb\u4ee4\u724c\u7ea7\u522b\u7684\u5f20\u91cf\u5217\u8868\uff0c\u6bcf\u4e2a\u5c42\u7684\u5f62\u72b6<span translate=no>_^_3_^_</span>\u5411\u91cf\u7279\u5f81</li>\n<li><span translate=no>_^_4_^_</span>\u662f\u63a9\u7801\u77e9\u9635</li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is a tensor of the token level feature vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is a tensor of the past token level feature vectors of shape <span translate=no>_^_3_^_</span> </li>\n<li><span translate=no>_^_4_^_</span> is a matrix of shape <span translate=no>_^_5_^_</span> or <span translate=no>_^_6_^_</span>. <span translate=no>_^_7_^_</span> is true if token at <span translate=no>_^_8_^_</span> can see token at <span translate=no>_^_9_^_</span>.</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u662f\u4ee4\u724c\u7ea7\u7279\u5f81\u5f62\u72b6\u5411\u91cf\u7684\u5f20\u91cf<span translate=no>_^_1_^_</span></li>\n<li><span translate=no>_^_2_^_</span>\u662f\u8fc7\u53bb\u4ee4\u724c\u7ea7\u522b\u7279\u5f81\u5f62\u72b6\u5411\u91cf\u7684\u5f20\u91cf<span translate=no>_^_3_^_</span></li>\n<li><span translate=no>_^_4_^_</span>\u662f\u5f62\u72b6\u7684\u77e9\u9635<span translate=no>_^_5_^_</span>\u6216<span translate=no>_^_6_^_</span>\u3002<span translate=no>_^_7_^_</span>\u5982\u679c token<span translate=no>_^_8_^_</span> \u53ef\u4ee5\u5728\u5904\u770b\u5230\u4ee4\u724c\uff0c\u5219\u4e3a true<span translate=no>_^_9_^_</span>\u3002</li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the token embedding size </li>\n<li><span translate=no>_^_1_^_</span> is the <a href=\"relative_mha.html\">self attention module</a> </li>\n<li><span translate=no>_^_2_^_</span> is the feed forward module </li>\n<li><span translate=no>_^_3_^_</span> is the probability of dropping out after self attention and FFN</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u662f\u4ee4\u724c\u5d4c\u5165\u7684\u5927\u5c0f</li>\n<li><span translate=no>_^_1_^_</span>\u662f<a href=\"relative_mha.html\">\u81ea\u6211\u5173\u6ce8\u6a21\u5757</a></li>\n<li><span translate=no>_^_2_^_</span>\u662f\u524d\u9988\u6a21\u5757</li>\n<li><span translate=no>_^_3_^_</span>\u662f\u81ea\u6211\u5173\u6ce8\u548c FFN \u540e\u9000\u5b66\u7684\u6982\u7387</li></ul>\n",
"Documented implementation with explanations of a Transformer-XL model.": "\u8bb0\u5f55\u4e86\u5b9e\u73b0\uff0c\u5e76\u89e3\u91ca\u4e86 Transformer-XL \u6a21\u578b\u3002",
"Transformer XL": "\u53d8\u538b\u5668 XL"
}
@@ -0,0 +1,74 @@
{
"<h1>Transformer XL Experiment</h1>\n<p>This is an annotated PyTorch experiment to train a transformer xl model.</p>\n": "<h1>\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL \u5b9f\u9a13</h1>\n<p>\u3053\u308c\u306f\u3001\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc xl \u30e2\u30c7\u30eb\u3092\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3059\u308b\u305f\u3081\u306e\u6ce8\u91c8\u4ed8\u304d\u306e PyTorch \u5b9f\u9a13\u3067\u3059\u3002</p>\n",
"<h2>Auto regressive model</h2>\n": "<h2>\u81ea\u52d5\u56de\u5e30\u30e2\u30c7\u30eb</h2>\n",
"<h2>Configurations</h2>\n<p>The default configs can and will be over-ridden when we start the experiment</p>\n": "<h2>\u30b3\u30f3\u30d5\u30a3\u30ae\u30e5\u30ec\u30fc\u30b7\u30e7\u30f3</h2>\n<p>\u30c7\u30d5\u30a9\u30eb\u30c8\u306e\u8a2d\u5b9a\u306f\u3001\u5b9f\u9a13\u3092\u958b\u59cb\u3057\u305f\u3068\u304d\u306b\u4e0a\u66f8\u304d\u3067\u304d\u3001\u307e\u305f\u4e0a\u66f8\u304d\u3055\u308c\u307e\u3059\u3002</p>\n",
"<h3>Initialize the auto-regressive model</h3>\n": "<h3>\u81ea\u5df1\u56de\u5e30\u30e2\u30c7\u30eb\u3092\u521d\u671f\u5316</h3>\n",
"<h3>Run the experiment</h3>\n": "<h3>\u5b9f\u9a13\u3092\u5b9f\u884c\u3059\u308b</h3>\n",
"<h3>Sampling function to generate samples periodically while training</h3>\n": "<h3>\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u4e2d\u306b\u5b9a\u671f\u7684\u306b\u30b5\u30f3\u30d7\u30eb\u3092\u751f\u6210\u3059\u308b\u30b5\u30f3\u30d7\u30ea\u30f3\u30b0\u6a5f\u80fd</h3>\n",
"<h3>Training/validation step</h3>\n": "<h3>\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0/\u691c\u8a3c\u30b9\u30c6\u30c3\u30d7</h3>\n",
"<p> </p>\n": "<p></p>\n",
"<p> Concatenate memories and remove old memories to keep a maximum of <span translate=no>_^_0_^_</span> memories.</p>\n": "<p>\u8a18\u61b6\u3092\u9023\u7d50\u3057\u3001\u53e4\u3044\u8a18\u61b6\u3092\u524a\u9664\u3057\u3066\u3001\u8a18\u61b6\u3092\u6700\u5927\u9650\u306b\u6d3b\u7528\u3057\u3066\u304f\u3060\u3055\u3044\u3002<span translate=no>_^_0_^_</span></p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span></p>\n",
"<p>A dictionary of configurations to override </p>\n": "<p>\u30aa\u30fc\u30d0\u30fc\u30e9\u30a4\u30c9\u3059\u308b\u8a2d\u5b9a\u306e\u8f9e\u66f8</p>\n",
"<p>Add a hook to log module outputs </p>\n": "<p>\u30e2\u30b8\u30e5\u30fc\u30eb\u51fa\u529b\u3092\u30ed\u30b0\u306b\u8a18\u9332\u3059\u308b\u30d5\u30c3\u30af\u3092\u8ffd\u52a0</p>\n",
"<p>Add the prediction for logging </p>\n": "<p>\u30ed\u30ae\u30f3\u30b0\u7528\u306e\u4e88\u6e2c\u3092\u8ffd\u52a0</p>\n",
"<p>Add the prediction to prompt </p>\n": "<p>\u4e88\u6e2c\u3092\u30d7\u30ed\u30f3\u30d7\u30c8\u306b\u8ffd\u52a0</p>\n",
"<p>Calculate and log accuracy </p>\n": "<p>\u7cbe\u5ea6\u306e\u8a08\u7b97\u3068\u8a18\u9332</p>\n",
"<p>Calculate and log cross entropy loss </p>\n": "<p>\u30af\u30ed\u30b9\u30a8\u30f3\u30c8\u30ed\u30d4\u30fc\u640d\u5931\u306e\u8a08\u7b97\u3068\u8a18\u9332</p>\n",
"<p>Calculate gradients </p>\n": "<p>\u52fe\u914d\u306e\u8a08\u7b97</p>\n",
"<p>Clear the gradients </p>\n": "<p>\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3\u3092\u30af\u30ea\u30a2</p>\n",
"<p>Clip gradients </p>\n": "<p>\u30af\u30ea\u30c3\u30d7\u30b0\u30e9\u30c7\u30fc\u30b7\u30e7\u30f3</p>\n",
"<p>Collect output for printing </p>\n": "<p>\u5370\u5237\u7528\u306e\u51fa\u529b\u3092\u53ce\u96c6</p>\n",
"<p>Concatenate the masks if there is memory </p>\n": "<p>\u30e1\u30e2\u30ea\u304c\u3042\u308b\u5834\u5408\u306f\u30de\u30b9\u30af\u3092\u9023\u7d50\u3057\u3066\u304f\u3060\u3055\u3044</p>\n",
"<p>Concatenate with old memory </p>\n": "<p>\u53e4\u3044\u30e1\u30e2\u30ea\u3068\u9023\u7d50</p>\n",
"<p>Create a subsequent mask for tokens </p>\n": "<p>\u30c8\u30fc\u30af\u30f3\u306e\u30de\u30b9\u30af\u3092\u5f8c\u304b\u3089\u4f5c\u6210</p>\n",
"<p>Create an all ones (full visibility) mask for memory </p>\n": "<p>\u30e1\u30e2\u30ea\u7528\u306e\u30aa\u30fc\u30eb\u30ef\u30f3 (\u30d5\u30eb\u30d3\u30b8\u30d3\u30ea\u30c6\u30a3) \u30de\u30b9\u30af\u3092\u4f5c\u6210</p>\n",
"<p>Create configs </p>\n": "<p>\u30b3\u30f3\u30d5\u30a3\u30b0\u306e\u4f5c\u6210</p>\n",
"<p>Create experiment </p>\n": "<p>\u5b9f\u9a13\u3092\u4f5c\u6210</p>\n",
"<p>Dropout probability </p>\n": "<p>\u8131\u843d\u78ba\u7387</p>\n",
"<p>Final layer </p>\n": "<p>\u6700\u7d42\u30ec\u30a4\u30e4\u30fc</p>\n",
"<p>Generate logits of the next token </p>\n": "<p>\u6b21\u306e\u30c8\u30fc\u30af\u30f3\u306e\u30ed\u30b8\u30c3\u30c8\u3092\u751f\u6210</p>\n",
"<p>Get memories </p>\n": "<p>\u601d\u3044\u51fa\u3092\u30b2\u30c3\u30c8</p>\n",
"<p>Get the model output </p>\n": "<p>\u30e2\u30c7\u30eb\u51fa\u529b\u3092\u53d6\u5f97</p>\n",
"<p>Get the model prediction (greedy) </p>\n": "<p>\u30e2\u30c7\u30eb\u4e88\u6e2c\u3092\u53d6\u5f97 (\u6b32\u5f35\u308a)</p>\n",
"<p>If it&#x27;s configured not to use memory </p>\n": "<p>\u30e1\u30e2\u30ea\u3092\u4f7f\u7528\u3057\u306a\u3044\u3088\u3046\u306b\u8a2d\u5b9a\u3055\u308c\u3066\u3044\u308b\u5834\u5408</p>\n",
"<p>Length of the memory </p>\n": "<p>\u30e1\u30e2\u30ea\u306e\u9577\u3055</p>\n",
"<p>Load configurations </p>\n": "<p>\u69cb\u6210\u3092\u30ed\u30fc\u30c9</p>\n",
"<p>Log the model parameters and gradients on last batch of every epoch </p>\n": "<p>\u5404\u30a8\u30dd\u30c3\u30af\u306e\u6700\u5f8c\u306e\u30d0\u30c3\u30c1\u3067\u30e2\u30c7\u30eb\u30d1\u30e9\u30e1\u30fc\u30bf\u3068\u52fe\u914d\u3092\u8a18\u9332\u3057\u307e\u3059</p>\n",
"<p>Masks </p>\n": "<p>\u30de\u30b9\u30af</p>\n",
"<p>Merge memory </p>\n": "<p>\u30de\u30fc\u30b8\u30e1\u30e2\u30ea</p>\n",
"<p>Move data to the device </p>\n": "<p>\u30c7\u30fc\u30bf\u3092\u30c7\u30d0\u30a4\u30b9\u306b\u79fb\u52d5</p>\n",
"<p>Move to device </p>\n": "<p>\u30c7\u30d0\u30a4\u30b9\u306b\u79fb\u52d5</p>\n",
"<p>Number of attention heads </p>\n": "<p>\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30d8\u30c3\u30c9\u306e\u6570</p>\n",
"<p>Number of features in FFN hidden layer </p>\n": "<p>FFN \u96a0\u308c\u30ec\u30a4\u30e4\u30fc\u306e\u30d5\u30a3\u30fc\u30c1\u30e3\u6570</p>\n",
"<p>Number of memories to keep </p>\n": "<p>\u4fdd\u5b58\u3059\u308b\u30e1\u30e2\u30ea\u306e\u6570</p>\n",
"<p>Number of transformer layers </p>\n": "<p>\u5909\u5727\u5668\u5c64\u306e\u6570</p>\n",
"<p>Only feed the last character to model in next iteration, rest will go in as memories </p>\n": "<p>\u6b21\u306e\u30a4\u30c6\u30ec\u30fc\u30b7\u30e7\u30f3\u3067\u306f\u6700\u5f8c\u306e\u6587\u5b57\u3060\u3051\u3092\u30e2\u30c7\u30eb\u306b\u30d5\u30a3\u30fc\u30c9\u3057\u3001\u6b8b\u308a\u306f\u30e1\u30e2\u30ea\u3068\u3057\u3066\u6b8b\u308a\u307e\u3059</p>\n",
"<p>Print the sampled output </p>\n": "<p>\u30b5\u30f3\u30d7\u30eb\u51fa\u529b\u3092\u5370\u5237\u3059\u308b</p>\n",
"<p>Run it through the transformer </p>\n": "<p>\u5909\u5727\u5668\u306b\u901a\u3057\u3066\u304f\u3060\u3055\u3044</p>\n",
"<p>Run the model </p>\n": "<p>\u30e2\u30c7\u30eb\u3092\u5b9f\u884c</p>\n",
"<p>Sample 25 tokens </p>\n": "<p>25\u30c8\u30fc\u30af\u30f3\u306e\u30b5\u30f3\u30d7\u30eb</p>\n",
"<p>Save the tracked metrics </p>\n": "<p>\u8ffd\u8de1\u3057\u305f\u30e1\u30c8\u30ea\u30af\u30b9\u3092\u4fdd\u5b58\u3059\u308b</p>\n",
"<p>Set models for saving and loading </p>\n": "<p>\u4fdd\u5b58\u304a\u3088\u3073\u8aad\u307f\u8fbc\u307f\u7528\u306e\u30e2\u30c7\u30eb\u3092\u8a2d\u5b9a\u3059\u308b</p>\n",
"<p>Set tracker configurations </p>\n": "<p>\u30c8\u30e9\u30c3\u30ab\u30fc\u69cb\u6210\u3092\u8a2d\u5b9a</p>\n",
"<p>Start the experiment </p>\n": "<p>\u5b9f\u9a13\u3092\u59cb\u3081\u308b</p>\n",
"<p>Starting prompt </p>\n": "<p>\u8d77\u52d5\u30d7\u30ed\u30f3\u30d7\u30c8</p>\n",
"<p>State module to maintain memories when switching between training and validation </p>\n": "<p>\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3068\u691c\u8a3c\u3092\u5207\u308a\u66ff\u3048\u308b\u3068\u304d\u306b\u30e1\u30e2\u30ea\u3092\u7dad\u6301\u3059\u308b\u30b9\u30c6\u30fc\u30c8\u30e2\u30b8\u30e5\u30fc\u30eb</p>\n",
"<p>Take optimizer step </p>\n": "<p>\u6700\u9069\u5316\u306e\u4e00\u6b69\u3092\u8e0f\u307f\u51fa\u3059</p>\n",
"<p>This will keep the accuracy metric stats and memories separate for training and validation. </p>\n": "<p>\u3053\u308c\u306b\u3088\u308a\u3001\u7cbe\u5ea6\u30e1\u30c8\u30ea\u30c3\u30af\u306e\u7d71\u8a08\u60c5\u5831\u3068\u30e1\u30e2\u30ea\u304c\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3068\u691c\u8a3c\u7528\u306b\u5225\u3005\u306b\u4fdd\u6301\u3055\u308c\u307e\u3059\u3002</p>\n",
"<p>Token embedding module </p>\n": "<p>\u30c8\u30fc\u30af\u30f3\u57cb\u3081\u8fbc\u307f\u30e2\u30b8\u30e5\u30fc\u30eb</p>\n",
"<p>Token embedding size </p>\n": "<p>\u30c8\u30fc\u30af\u30f3\u306e\u57cb\u3081\u8fbc\u307f\u30b5\u30a4\u30ba</p>\n",
"<p>Token embeddings </p>\n": "<p>\u30c8\u30fc\u30af\u30f3\u306e\u57cb\u3081\u8fbc\u307f</p>\n",
"<p>Tokenize the prompt </p>\n": "<p>\u30d7\u30ed\u30f3\u30d7\u30c8\u3092\u30c8\u30fc\u30af\u30f3\u5316</p>\n",
"<p>Train the model </p>\n": "<p>\u30e2\u30c7\u30eb\u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0</p>\n",
"<p>Transformer </p>\n": "<p>\u5909\u5727\u5668</p>\n",
"<p>Truncate old memories </p>\n": "<p>\u53e4\u3044\u601d\u3044\u51fa\u3092\u5207\u308a\u6368\u3066\u308b</p>\n",
"<p>Update global step (number of tokens processed) when in training mode </p>\n": "<p>\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u30e2\u30fc\u30c9\u6642\u306b\u30b0\u30ed\u30fc\u30d0\u30eb\u30b9\u30c6\u30c3\u30d7 (\u51e6\u7406\u3055\u308c\u305f\u30c8\u30fc\u30af\u30f3\u306e\u6570) \u3092\u66f4\u65b0</p>\n",
"<p>Update memories </p>\n": "<p>\u30e1\u30e2\u30ea\u30fc\u3092\u66f4\u65b0</p>\n",
"<p>Update memory </p>\n": "<p>\u30e1\u30e2\u30ea\u3092\u66f4\u65b0</p>\n",
"<p>Use the subsequent mask otherwise </p>\n": "<p>\u305d\u308c\u4ee5\u5916\u306e\u5834\u5408\u306f\u3001\u5f8c\u7d9a\u306e\u30de\u30b9\u30af\u3092\u4f7f\u7528\u3057\u3066\u304f\u3060\u3055\u3044\u3002</p>\n",
"<p>Whether to capture model outputs </p>\n": "<p>\u30e2\u30c7\u30eb\u51fa\u529b\u3092\u30ad\u30e3\u30d7\u30c1\u30e3\u3059\u308b\u304b\u3069\u3046\u304b</p>\n",
"<p>memory </p>\n": "<p>\u8a18\u61b6</p>\n",
"This experiment trains a transformer XL model on tiny Shakespeare dataset.": "\u3053\u306e\u5b9f\u9a13\u3067\u306f\u3001\u5c0f\u3055\u306a\u30b7\u30a7\u30a4\u30af\u30b9\u30d4\u30a2\u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u3067\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL \u30e2\u30c7\u30eb\u3092\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3057\u307e\u3059\u3002",
"Transformer XL Experiment": "\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL \u5b9f\u9a13"
}
@@ -0,0 +1,74 @@
{
"<h1>Transformer XL Experiment</h1>\n<p>This is an annotated PyTorch experiment to train a transformer xl model.</p>\n": "<h1>\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca\u0d91\u0d9a\u0dca\u0dc3\u0dca\u0d91\u0dbd\u0dca \u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf \u0db6\u0dd0\u0dbd\u0dd3\u0db8</h1>\n<p>\u0db8\u0dd9\u0dba\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca xl \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0d9a\u0dca \u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0d9a\u0dbb\u0db1 \u0dbd\u0daf \u0db4\u0dba\u0dd2\u0da7\u0ddd\u0dbb\u0dca\u0da0\u0dca \u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf \u0db6\u0dd0\u0dbd\u0dd3\u0db8\u0d9a\u0dd2. </p>\n",
"<h2>Auto regressive model</h2>\n": "<h2>\u0dc3\u0dca\u0dc0\u0dba\u0d82\u0d9a\u0dca\u0dbb\u0dd3\u0dba\u0db4\u0dca\u0dbb\u0dad\u0dd2\u0d9c\u0dcf\u0db8\u0dd3 \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba</h2>\n",
"<h2>Configurations</h2>\n<p>The default configs can and will be over-ridden when we start the experiment</p>\n": "<h2>\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dca</h2>\n<p>\u0d85\u0db4\u0dd2\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf \u0db6\u0dd0\u0dbd\u0dd3\u0db8 \u0d86\u0dbb\u0db8\u0dca\u0db7 \u0d9a\u0dbb\u0db1 \u0dc0\u0dd2\u0da7 \u0db4\u0dd9\u0dbb\u0db1\u0dd2\u0db8\u0dd2 \u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3 \u0d9a\u0dc5 \u0dc4\u0dd0\u0d9a\u0dd2 \u0d85\u0dad\u0dbb \u0d91\u0dba \u0d85\u0db0\u0dd2\u0d9a \u0dbd\u0dd9\u0dc3 \u0db0\u0dcf\u0dc0\u0db1\u0dba \u0dc0\u0db1\u0dd4 \u0d87\u0dad</p>\n",
"<h3>Initialize the auto-regressive model</h3>\n": "<h3>\u0dc3\u0dca\u0dc0\u0dba\u0d82\u0d9a\u0dca\u0dbb\u0dd3\u0dba\u0db4\u0dca\u0dbb\u0dad\u0dd2\u0d9c\u0dcf\u0db8\u0dd3 \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba \u0d86\u0dbb\u0db8\u0dca\u0db7 \u0d9a\u0dbb\u0db1\u0dca\u0db1</h3>\n",
"<h3>Run the experiment</h3>\n": "<h3>\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf\u0db6\u0dd0\u0dbd\u0dd3\u0db8 \u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0dad\u0dca\u0db8\u0d9a \u0d9a\u0dbb\u0db1\u0dca\u0db1</h3>\n",
"<h3>Sampling function to generate samples periodically while training</h3>\n": "<h3>\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4\u0dc0\u0d85\u0dad\u0dbb\u0dad\u0dd4\u0dbb \u0dc0\u0dbb\u0dd2\u0db1\u0dca \u0dc0\u0dbb \u0dc3\u0dcf\u0db8\u0dca\u0db4\u0dbd \u0da2\u0db1\u0db1\u0dba \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0db1\u0dd2\u0dba\u0dd0\u0daf\u0dd2 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda \u0d9a\u0dcf\u0dbb\u0dca\u0dba\u0dba</h3>\n",
"<h3>Training/validation step</h3>\n": "<h3>\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4\u0dc0/\u0dc0\u0dbd\u0d82\u0d9c\u0dd4\u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda \u0db4\u0dd2\u0dba\u0dc0\u0dbb</h3>\n",
"<p> </p>\n": "<p> </p>\n",
"<p> Concatenate memories and remove old memories to keep a maximum of <span translate=no>_^_0_^_</span> memories.</p>\n": "<p> \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca\u0dc3\u0d82\u0d9a\u0ddd\u0da0\u0db1\u0dba \u0d9a\u0dbb \u0d8b\u0db4\u0dbb\u0dd2\u0db8 \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0dad\u0db6\u0dcf \u0d9c\u0dd0\u0db1\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0db4\u0dd0\u0dbb\u0dab\u0dd2 <span translate=no>_^_0_^_</span> \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0d89\u0dc0\u0dad\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1. </p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span> </p>\n",
"<p>A dictionary of configurations to override </p>\n": "<p>\u0d85\u0db7\u0dd2\u0db6\u0dc0\u0dcf\u0dba\u0dcf\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0dba\u0db1\u0dca \u0db4\u0dd2\u0dc5\u0dd2\u0db6\u0db3 \u0dc1\u0db6\u0dca\u0daf\u0d9a\u0ddd\u0dc2\u0dba\u0d9a\u0dca </p>\n",
"<p>Add a hook to log module outputs </p>\n": "<p>\u0db8\u0ddc\u0da9\u0dd2\u0dba\u0dd4\u0dbd\u0db4\u0dca\u0dbb\u0dad\u0dd2\u0daf\u0dcf\u0db1\u0dba\u0db1\u0dca \u0dbd\u0ddc\u0d9c\u0dca \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0da7 \u0d9a\u0ddc\u0d9a\u0dca\u0d9a\u0d9a\u0dca \u0d91\u0d9a\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Add the prediction for logging </p>\n": "<p>\u0dbd\u0ddc\u0d9c\u0dca\u0dc0\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0d85\u0db1\u0dcf\u0dc0\u0dd0\u0d9a\u0dd2\u0dba \u0d91\u0d9a\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Add the prediction to prompt </p>\n": "<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0da7\u0d85\u0db1\u0dcf\u0dc0\u0dd0\u0d9a\u0dd2\u0dba \u0d91\u0d9a\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Calculate and log accuracy </p>\n": "<p>\u0d9c\u0dab\u0db1\u0dba\u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0dc4 \u0dbd\u0ddc\u0d9c\u0dca \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda \u0db1\u0dd2\u0dbb\u0dc0\u0daf\u0dca\u0dba\u0dad\u0dcf\u0dc0\u0dba </p>\n",
"<p>Calculate and log cross entropy loss </p>\n": "<p>\u0dc4\u0dbb\u0dc3\u0dca\u0d91\u0db1\u0dca\u0da7\u0dca\u0dbb\u0ddc\u0db4\u0dd2 \u0d85\u0dbd\u0dcf\u0db7\u0dba \u0d9c\u0dab\u0db1\u0dba \u0d9a\u0dbb \u0dbd\u0ddc\u0d9c\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Calculate gradients </p>\n": "<p>\u0d85\u0db1\u0dd4\u0d9a\u0dca\u0dbb\u0db8\u0dd2\u0d9a\u0d9c\u0dab\u0db1\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Clear the gradients </p>\n": "<p>\u0d85\u0db1\u0dd4\u0d9a\u0dca\u0dbb\u0db8\u0dd2\u0d9a\u0d89\u0dc0\u0dad\u0dca </p>\n",
"<p>Clip gradients </p>\n": "<p>\u0d9a\u0dca\u0dbd\u0dd2\u0db4\u0dca\u0d85\u0db1\u0dd4\u0d9a\u0dca\u0dbb\u0db8\u0dd2\u0d9a </p>\n",
"<p>Collect output for printing </p>\n": "<p>\u0db8\u0dd4\u0daf\u0dca\u0dbb\u0dab\u0dba\u0dc3\u0db3\u0dc4\u0dcf \u0db4\u0dca\u0dbb\u0dad\u0dd2\u0daf\u0dcf\u0db1\u0dba \u0d91\u0d9a\u0dad\u0dd4 \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Concatenate the masks if there is memory </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0d9a\u0dca\u0dad\u0dd2\u0db6\u0dda \u0db1\u0db8\u0dca \u0dc0\u0dd9\u0dc3\u0dca \u0db8\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 \u0dc3\u0d82\u0dba\u0dd4\u0d9a\u0dca\u0dad \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Concatenate with old memory </p>\n": "<p>\u0db4\u0dd0\u0dbb\u0dab\u0dd2\u0db8\u0dad\u0d9a\u0dba \u0dc3\u0db8\u0d9f \u0dc3\u0d82\u0dba\u0dd4\u0d9a\u0dca\u0dad \u0dc0\u0db1\u0dca\u0db1 </p>\n",
"<p>Create a subsequent mask for tokens </p>\n": "<p>\u0da7\u0ddd\u0d9a\u0db1\u0dc3\u0db3\u0dc4\u0dcf \u0db4\u0dc3\u0dd4\u0d9a\u0dcf\u0dbd\u0dd3\u0db1 \u0dc0\u0dd9\u0dc3\u0dca \u0db8\u0dd4\u0dc4\u0dd4\u0dab\u0d9a\u0dca \u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Create an all ones (full visibility) mask for memory </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0dc3\u0db3\u0dc4\u0dcf \u0dc3\u0dd2\u0dba\u0dbd\u0dd4 (\u0dc3\u0db8\u0dca\u0db4\u0dd6\u0dbb\u0dca\u0dab \u0daf\u0dd8\u0dc1\u0dca\u0dba\u0dad\u0dcf\u0dc0) \u0dc0\u0dd9\u0dc3\u0dca\u0db8\u0dd4\u0dc4\u0dd4\u0dab\u0d9a\u0dca \u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Create configs </p>\n": "<p>\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Create experiment </p>\n": "<p>\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf\u0db6\u0dd0\u0dbd\u0dd3\u0db8 \u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Dropout probability </p>\n": "<p>\u0d85\u0dad\u0dc4\u0dd0\u0dbb\u0daf\u0dd0\u0db8\u0dd3\u0db8\u0dda \u0dc3\u0db8\u0dca\u0db7\u0dcf\u0dc0\u0dd2\u0dad\u0dcf\u0dc0 </p>\n",
"<p>Final layer </p>\n": "<p>\u0d85\u0dc0\u0dc3\u0db1\u0dca\u0dc3\u0dca\u0dae\u0dbb\u0dba </p>\n",
"<p>Generate logits of the next token </p>\n": "<p>\u0d8a\u0dc5\u0d9f\u0da7\u0ddd\u0d9a\u0db1\u0dba\u0dda \u0db4\u0dd2\u0dc0\u0dd2\u0dc3\u0dd4\u0db8\u0dca \u0da2\u0db1\u0db1\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Get memories </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca\u0dbd\u0db6\u0dcf \u0d9c\u0db1\u0dca\u0db1 </p>\n",
"<p>Get the model output </p>\n": "<p>\u0d86\u0daf\u0dbb\u0dca\u0dc1\u0db4\u0dca\u0dbb\u0dad\u0dd2\u0daf\u0dcf\u0db1\u0dba \u0dbd\u0db6\u0dcf \u0d9c\u0db1\u0dca\u0db1 </p>\n",
"<p>Get the model prediction (greedy) </p>\n": "<p>\u0d86\u0daf\u0dbb\u0dca\u0dc1\u0d85\u0db1\u0dcf\u0dc0\u0dd0\u0d9a\u0dd2\u0dba \u0dbd\u0db6\u0dcf \u0d9c\u0db1\u0dca\u0db1 (\u0d9a\u0dd1\u0daf\u0dbb) </p>\n",
"<p>If it&#x27;s configured not to use memory </p>\n": "<p>\u0d91\u0dba\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3 \u0d9a\u0dbb \u0d87\u0dad\u0dca\u0db1\u0db8\u0dca \u0db8\u0dad\u0d9a\u0dba \u0db7\u0dcf\u0dc0\u0dd2\u0dad\u0dcf \u0db1\u0ddc\u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0da7 </p>\n",
"<p>Length of the memory </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0dda\u0daf\u0dd2\u0d9c </p>\n",
"<p>Load configurations </p>\n": "<p>\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0dba\u0db1\u0dca\u0db4\u0dd6\u0dbb\u0dab\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Log the model parameters and gradients on last batch of every epoch </p>\n": "<p>\u0dc3\u0dd1\u0db8\u0dba\u0dd4\u0d9c\u0dbd\u0dba\u0d9a\u0db8 \u0d85\u0dc0\u0dc3\u0dcf\u0db1 \u0d9a\u0dab\u0dca\u0da9\u0dcf\u0dba\u0db8\u0dda \u0d86\u0daf\u0dbb\u0dca\u0dc1 \u0db4\u0dbb\u0dcf\u0db8\u0dd2\u0dad\u0dd3\u0db1\u0dca \u0dc3\u0dc4 \u0d85\u0db1\u0dd4\u0d9a\u0dca\u0dbb\u0db8\u0dd2\u0d9a \u0dbd\u0ddc\u0d9c\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Masks </p>\n": "<p>\u0dc0\u0dd9\u0dc3\u0dca\u0db8\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 </p>\n",
"<p>Merge memory </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0d92\u0d9a\u0dcf\u0db6\u0daf\u0dca\u0db0 \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Move data to the device </p>\n": "<p>\u0d8b\u0db4\u0dcf\u0d82\u0d9c\u0dba\u0dc0\u0dd9\u0dad \u0daf\u0dad\u0dca\u0dad \u0d9c\u0dd9\u0db1\u0dba\u0db1\u0dca\u0db1 </p>\n",
"<p>Move to device </p>\n": "<p>\u0d8b\u0db4\u0dcf\u0d82\u0d9c\u0dba\u0dc0\u0dd9\u0dad \u0d9c\u0dd9\u0db1 \u0dba\u0db1\u0dca\u0db1 </p>\n",
"<p>Number of attention heads </p>\n": "<p>\u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba\u0dba\u0ddc\u0db8\u0dd4 \u0db4\u0dca\u0dbb\u0db0\u0dcf\u0db1\u0dd3\u0db1\u0dca \u0d9c\u0dab\u0db1 </p>\n",
"<p>Number of features in FFN hidden layer </p>\n": "<p>FFN\u0dc3\u0dd0\u0d9f\u0dc0\u0dd4\u0dab\u0dd4 \u0dc3\u0dca\u0dae\u0dbb\u0dba\u0dda \u0dc0\u0dd2\u0dc1\u0dda\u0dc2\u0dcf\u0d82\u0d9c \u0d9c\u0dab\u0db1 </p>\n",
"<p>Number of memories to keep </p>\n": "<p>\u0dad\u0db6\u0dcf\u0d9c\u0dad \u0dba\u0dd4\u0dad\u0dd4 \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0d9c\u0dab\u0db1 </p>\n",
"<p>Number of transformer layers </p>\n": "<p>\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca\u0dc3\u0dca\u0dae\u0dbb \u0d9c\u0dab\u0db1 </p>\n",
"<p>Only feed the last character to model in next iteration, rest will go in as memories </p>\n": "<p>\u0d8a\u0dc5\u0d9f\u0db4\u0dd4\u0db1\u0dbb\u0dcf\u0dc0\u0dbb\u0dca\u0dad\u0db1\u0dba\u0dda\u0daf\u0dd3 \u0d85\u0dc0\u0dc3\u0dcf\u0db1 \u0da0\u0dbb\u0dd2\u0dad\u0dba \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0da7 \u0db4\u0db8\u0dab\u0d9a\u0dca \u0db4\u0ddd\u0dc2\u0dab\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1, \u0dc0\u0dd2\u0dc0\u0dda\u0d9a\u0dba \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0dbd\u0dd9\u0dc3 \u0d89\u0daf\u0dd2\u0dbb\u0dd2\u0dba\u0da7 \u0dba\u0db1\u0dd4 \u0d87\u0dad </p>\n",
"<p>Print the sampled output </p>\n": "<p>\u0db1\u0dd2\u0dba\u0dd0\u0daf\u0dd2\u0db4\u0dca\u0dbb\u0dad\u0dd2\u0daf\u0dcf\u0db1\u0dba \u0db8\u0dd4\u0daf\u0dca\u0dbb\u0dab\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Run it through the transformer </p>\n": "<p>\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dba\u0dc4\u0dbb\u0dc4\u0dcf \u0d91\u0dba \u0db0\u0dcf\u0dc0\u0db1\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Run the model </p>\n": "<p>\u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0db0\u0dcf\u0dc0\u0db1\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Sample 25 tokens </p>\n": "<p>\u0dc3\u0dcf\u0db8\u0dca\u0db4\u0dbd25 \u0da7\u0ddd\u0d9a\u0db1 </p>\n",
"<p>Save the tracked metrics </p>\n": "<p>\u0dbd\u0dd4\u0dc4\u0dd4\u0db6\u0dd0\u0db3\u0d87\u0dad\u0dd2 \u0db4\u0dca\u0dbb\u0db8\u0dd2\u0dad\u0dd2\u0d9a \u0dc3\u0dd4\u0dbb\u0d9a\u0dd2\u0db1\u0dca\u0db1 </p>\n",
"<p>Set models for saving and loading </p>\n": "<p>\u0d89\u0dad\u0dd2\u0dbb\u0dd2\u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0dc4 \u0db4\u0dd0\u0da7\u0dc0\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0d86\u0d9a\u0dd8\u0dad\u0dd2 \u0dc3\u0d9a\u0dc3\u0db1\u0dca\u0db1 </p>\n",
"<p>Set tracker configurations </p>\n": "<p>\u0da7\u0dca\u0dbb\u0dd0\u0d9a\u0dbb\u0dca\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0dba\u0db1\u0dca \u0dc3\u0d9a\u0dc3\u0db1\u0dca\u0db1 </p>\n",
"<p>Start the experiment </p>\n": "<p>\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf\u0db6\u0dd0\u0dbd\u0dd3\u0db8 \u0d86\u0dbb\u0db8\u0dca\u0db7 \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Starting prompt </p>\n": "<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0d9a\u0dca\u0d86\u0dbb\u0db8\u0dca\u0db7 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 </p>\n",
"<p>State module to maintain memories when switching between training and validation </p>\n": "<p>\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4\u0dc0\u0dc3\u0dc4 \u0dc0\u0dbd\u0d82\u0d9c\u0dd4 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0d85\u0dad\u0dbb \u0db8\u0dcf\u0dbb\u0dd4\u0dc0\u0dd3\u0db8\u0dda\u0daf\u0dd3 \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0db4\u0dc0\u0dad\u0dca\u0dc0\u0dcf \u0d9c\u0dd0\u0db1\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0dbb\u0dcf\u0da2\u0dca\u0dba \u0db8\u0ddc\u0da9\u0dd2\u0dba\u0dd4\u0dbd\u0dba </p>\n",
"<p>Take optimizer step </p>\n": "<p>\u0db4\u0dca\u0dbb\u0dc1\u0dc3\u0dca\u0dad\u0dd2\u0d9a\u0dbb\u0dab\u0db4\u0dd2\u0dba\u0dc0\u0dbb \u0d9c\u0db1\u0dca\u0db1 </p>\n",
"<p>This will keep the accuracy metric stats and memories separate for training and validation. </p>\n": "<p>\u0db8\u0dd9\u0dba\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4\u0dc0 \u0dc3\u0dc4 \u0dc0\u0dbd\u0d82\u0d9c\u0dd4 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0db1\u0dd2\u0dbb\u0dc0\u0daf\u0dca\u0dba\u0dad\u0dcf \u0db8\u0dd9\u0da7\u0dca\u0dbb\u0dd2\u0d9a\u0dca \u0dc3\u0d82\u0d9b\u0dca\u0dba\u0dcf\u0db1 \u0dc3\u0dc4 \u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0dc0\u0dd9\u0db1\u0db8 \u0dad\u0db6\u0dcf \u0d9c\u0db1\u0dd3. </p>\n",
"<p>Token embedding module </p>\n": "<p>\u0da7\u0ddd\u0d9a\u0db1\u0dca\u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8 \u0db8\u0ddc\u0da9\u0dd2\u0dba\u0dd4\u0dbd\u0dba </p>\n",
"<p>Token embedding size </p>\n": "<p>\u0da7\u0ddd\u0d9a\u0db1\u0dca\u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dda \u0db4\u0dca\u0dbb\u0db8\u0dcf\u0dab\u0dba </p>\n",
"<p>Token embeddings </p>\n": "<p>\u0da7\u0ddd\u0d9a\u0db1\u0dca\u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dca </p>\n",
"<p>Tokenize the prompt </p>\n": "<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0da7\u0ddd\u0d9a\u0dd9\u0db1\u0dca\u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Train the model </p>\n": "<p>\u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Transformer </p>\n": "<p>\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca </p>\n",
"<p>Truncate old memories </p>\n": "<p>\u0db4\u0dd0\u0dbb\u0dab\u0dd2\u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca \u0d89\u0dc0\u0dad\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Update global step (number of tokens processed) when in training mode </p>\n": "<p>\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4\u0db4\u0dca\u0dbb\u0d9a\u0dcf\u0dbb\u0dba\u0dda\u0daf\u0dd3 \u0d9c\u0ddd\u0dbd\u0dd3\u0dba \u0db4\u0dd2\u0dba\u0dc0\u0dbb \u0dba\u0dcf\u0dc0\u0dad\u0dca\u0d9a\u0dcf\u0dbd\u0dd3\u0db1 \u0d9a\u0dbb\u0db1\u0dca\u0db1 (\u0dc3\u0dd0\u0d9a\u0dc3\u0dd6 \u0da7\u0ddd\u0d9a\u0db1 \u0d9c\u0dab\u0db1) </p>\n",
"<p>Update memories </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0db1\u0dca\u0dba\u0dcf\u0dc0\u0dad\u0dca\u0d9a\u0dcf\u0dbd\u0dd3\u0db1 \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Update memory </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba\u0dba\u0dcf\u0dc0\u0dad\u0dca\u0d9a\u0dcf\u0dbd\u0dd3\u0db1 \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Use the subsequent mask otherwise </p>\n": "<p>\u0db4\u0dc3\u0dd4\u0d9a\u0dcf\u0dbd\u0dd3\u0db1\u0d86\u0dc0\u0dbb\u0dab \u0dc0\u0dd9\u0db1\u0dad\u0dca \u0d86\u0d9a\u0dcf\u0dbb\u0dba\u0d9a\u0dd2\u0db1\u0dca \u0db7\u0dcf\u0dc0\u0dd2\u0dad\u0dcf \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Whether to capture model outputs </p>\n": "<p>\u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0db4\u0dca\u0dbb\u0dad\u0dd2\u0daf\u0dcf\u0db1\u0dba\u0db1\u0dca \u0d9c\u0dca\u0dbb\u0dc4\u0dab\u0dba \u0d9a\u0dbb \u0d9c\u0dad \u0dba\u0dd4\u0dad\u0dd4\u0daf \u0dba\u0db1\u0dca\u0db1 </p>\n",
"<p>memory </p>\n": "<p>\u0db8\u0dad\u0d9a\u0dba </p>\n",
"This experiment trains a transformer XL model on tiny Shakespeare dataset.": "\u0db8\u0dd9\u0db8 \u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf \u0db6\u0dd0\u0dbd\u0dd3\u0db8 \u0d9a\u0dd4\u0da9\u0dcf \u0dc2\u0dda\u0d9a\u0dca\u0dc3\u0dca\u0db4\u0dd2\u0dba\u0dbb\u0dca \u0daf\u0dad\u0dca\u0dad \u0d9a\u0da7\u0dca\u0da7\u0dbd\u0dba\u0dda \u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca \u0d91\u0d9a\u0dca\u0dc3\u0dca\u0d91\u0dbd\u0dca \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0d9a\u0dca \u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 \u0d9a\u0dbb\u0dba\u0dd2.",
"Transformer XL Experiment": "\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca \u0d91\u0d9a\u0dca\u0dc3\u0dca\u0d91\u0dbd\u0dca \u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf \u0db6\u0dd0\u0dbd\u0dd3\u0db8"
}
@@ -0,0 +1,74 @@
{
"<h1>Transformer XL Experiment</h1>\n<p>This is an annotated PyTorch experiment to train a transformer xl model.</p>\n": "<h1>\u53d8\u538b\u5668 XL \u5b9e\u9a8c</h1>\n<p>\u8fd9\u662f\u4e00\u4e2a\u5e26\u6ce8\u91ca\u7684 PyTorch \u5b9e\u9a8c\uff0c\u7528\u4e8e\u8bad\u7ec3\u53d8\u538b\u5668 xl \u6a21\u578b\u3002</p>\n",
"<h2>Auto regressive model</h2>\n": "<h2>\u81ea\u52a8\u56de\u5f52\u6a21\u578b</h2>\n",
"<h2>Configurations</h2>\n<p>The default configs can and will be over-ridden when we start the experiment</p>\n": "<h2>\u914d\u7f6e</h2>\n<p>\u5f53\u6211\u4eec\u5f00\u59cb\u5b9e\u9a8c\u65f6\uff0c\u9ed8\u8ba4\u914d\u7f6e\u53ef\u4ee5\u800c\u4e14\u5c06\u4f1a\u88ab\u8986\u76d6</p>\n",
"<h3>Initialize the auto-regressive model</h3>\n": "<h3>\u521d\u59cb\u5316\u81ea\u56de\u5f52\u6a21\u578b</h3>\n",
"<h3>Run the experiment</h3>\n": "<h3>\u8fd0\u884c\u5b9e\u9a8c</h3>\n",
"<h3>Sampling function to generate samples periodically while training</h3>\n": "<h3>\u91c7\u6837\u529f\u80fd\u53ef\u5728\u8bad\u7ec3\u65f6\u5b9a\u671f\u751f\u6210\u6837\u672c</h3>\n",
"<h3>Training/validation step</h3>\n": "<h3>\u57f9\u8bad/\u9a8c\u8bc1\u6b65\u9aa4</h3>\n",
"<p> </p>\n": "<p></p>\n",
"<p> Concatenate memories and remove old memories to keep a maximum of <span translate=no>_^_0_^_</span> memories.</p>\n": "<p>\u8fde\u63a5\u8bb0\u5fc6\u5e76\u5220\u9664\u65e7\u8bb0\u5fc6\u4ee5\u6700\u5927\u9650\u5ea6\u5730\u4fdd\u7559\u5185<span translate=no>_^_0_^_</span>\u5b58\u3002</p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span></p>\n",
"<p>A dictionary of configurations to override </p>\n": "<p>\u8981\u8986\u76d6\u7684\u914d\u7f6e\u5b57\u5178</p>\n",
"<p>Add a hook to log module outputs </p>\n": "<p>\u5411\u65e5\u5fd7\u6a21\u5757\u8f93\u51fa\u6dfb\u52a0\u94a9\u5b50</p>\n",
"<p>Add the prediction for logging </p>\n": "<p>\u6dfb\u52a0\u65e5\u5fd7\u8bb0\u5f55\u7684\u9884\u6d4b</p>\n",
"<p>Add the prediction to prompt </p>\n": "<p>\u5c06\u9884\u6d4b\u6dfb\u52a0\u5230\u63d0\u793a\u7b26\u4e2d</p>\n",
"<p>Calculate and log accuracy </p>\n": "<p>\u8ba1\u7b97\u548c\u8bb0\u5f55\u7cbe\u5ea6</p>\n",
"<p>Calculate and log cross entropy loss </p>\n": "<p>\u8ba1\u7b97\u548c\u8bb0\u5f55\u4ea4\u53c9\u71b5\u635f\u5931</p>\n",
"<p>Calculate gradients </p>\n": "<p>\u8ba1\u7b97\u68af\u5ea6</p>\n",
"<p>Clear the gradients </p>\n": "<p>\u6e05\u9664\u6e10\u53d8</p>\n",
"<p>Clip gradients </p>\n": "<p>\u526a\u8f91\u6e10\u53d8</p>\n",
"<p>Collect output for printing </p>\n": "<p>\u6536\u96c6\u8f93\u51fa\u4ee5\u8fdb\u884c\u6253\u5370</p>\n",
"<p>Concatenate the masks if there is memory </p>\n": "<p>\u5982\u679c\u6709\u5185\u5b58\uff0c\u5219\u8fde\u63a5\u63a9\u7801</p>\n",
"<p>Concatenate with old memory </p>\n": "<p>\u4e0e\u65e7\u5185\u5b58\u4e32\u8054</p>\n",
"<p>Create a subsequent mask for tokens </p>\n": "<p>\u4e3a\u4ee4\u724c\u521b\u5efa\u540e\u7eed\u63a9\u7801</p>\n",
"<p>Create an all ones (full visibility) mask for memory </p>\n": "<p>\u4e3a\u5185\u5b58\u521b\u5efa\u4e00\u4e2a\u5168\u4e00\uff08\u5b8c\u5168\u53ef\u89c1\u6027\uff09\u63a9\u7801</p>\n",
"<p>Create configs </p>\n": "<p>\u521b\u5efa\u914d\u7f6e</p>\n",
"<p>Create experiment </p>\n": "<p>\u521b\u5efa\u5b9e\u9a8c</p>\n",
"<p>Dropout probability </p>\n": "<p>\u8f8d\u5b66\u6982\u7387</p>\n",
"<p>Final layer </p>\n": "<p>\u6700\u540e\u4e00\u5c42</p>\n",
"<p>Generate logits of the next token </p>\n": "<p>\u751f\u6210\u4e0b\u4e00\u4e2a\u4ee4\u724c\u7684\u65e5\u5fd7</p>\n",
"<p>Get memories </p>\n": "<p>\u83b7\u5f97\u56de\u5fc6</p>\n",
"<p>Get the model output </p>\n": "<p>\u83b7\u53d6\u6a21\u578b\u8f93\u51fa</p>\n",
"<p>Get the model prediction (greedy) </p>\n": "<p>\u83b7\u53d6\u6a21\u578b\u9884\u6d4b\uff08\u8d2a\u5a6a\uff09</p>\n",
"<p>If it&#x27;s configured not to use memory </p>\n": "<p>\u5982\u679c\u914d\u7f6e\u4e3a\u4e0d\u4f7f\u7528\u5185\u5b58</p>\n",
"<p>Length of the memory </p>\n": "<p>\u5185\u5b58\u7684\u957f\u5ea6</p>\n",
"<p>Load configurations </p>\n": "<p>\u88c5\u8f7d\u914d\u7f6e</p>\n",
"<p>Log the model parameters and gradients on last batch of every epoch </p>\n": "<p>\u8bb0\u5f55\u6bcf\u4e2a\u7eaa\u5143\u6700\u540e\u4e00\u6279\u7684\u6a21\u578b\u53c2\u6570\u548c\u68af\u5ea6</p>\n",
"<p>Masks </p>\n": "<p>\u53e3\u7f69</p>\n",
"<p>Merge memory </p>\n": "<p>\u5408\u5e76\u5185\u5b58</p>\n",
"<p>Move data to the device </p>\n": "<p>\u5c06\u6570\u636e\u79fb\u52a8\u5230\u8bbe\u5907</p>\n",
"<p>Move to device </p>\n": "<p>\u79fb\u81f3\u8bbe\u5907</p>\n",
"<p>Number of attention heads </p>\n": "<p>\u6ce8\u610f\u5934\u6570\u91cf</p>\n",
"<p>Number of features in FFN hidden layer </p>\n": "<p>FFN \u9690\u85cf\u5c42\u4e2d\u7684\u8981\u7d20\u6570\u91cf</p>\n",
"<p>Number of memories to keep </p>\n": "<p>\u8981\u4fdd\u7559\u7684\u8bb0\u5fc6\u6570\u91cf</p>\n",
"<p>Number of transformer layers </p>\n": "<p>\u53d8\u538b\u5668\u5c42\u6570</p>\n",
"<p>Only feed the last character to model in next iteration, rest will go in as memories </p>\n": "<p>\u5728\u4e0b\u4e00\u6b21\u8fed\u4ee3\u4e2d\u53ea\u5582\u6700\u540e\u4e00\u4e2a\u89d2\u8272\u8fdb\u884c\u5efa\u6a21\uff0c\u5176\u4f59\u90e8\u5206\u5c06\u4f5c\u4e3a\u8bb0\u5fc6\u8fdb\u53bb</p>\n",
"<p>Print the sampled output </p>\n": "<p>\u6253\u5370\u91c7\u6837\u8f93\u51fa</p>\n",
"<p>Run it through the transformer </p>\n": "<p>\u7528\u5b83\u7a7f\u8fc7\u53d8\u538b\u5668</p>\n",
"<p>Run the model </p>\n": "<p>\u8fd0\u884c\u6a21\u578b</p>\n",
"<p>Sample 25 tokens </p>\n": "<p>\u6837\u672c 25 \u4e2a\u4ee3\u5e01</p>\n",
"<p>Save the tracked metrics </p>\n": "<p>\u4fdd\u5b58\u8ddf\u8e2a\u7684\u6307\u6807</p>\n",
"<p>Set models for saving and loading </p>\n": "<p>\u8bbe\u7f6e\u7528\u4e8e\u4fdd\u5b58\u548c\u52a0\u8f7d\u7684\u6a21\u578b</p>\n",
"<p>Set tracker configurations </p>\n": "<p>\u8bbe\u7f6e\u8ddf\u8e2a\u5668\u914d\u7f6e</p>\n",
"<p>Start the experiment </p>\n": "<p>\u5f00\u59cb\u5b9e\u9a8c</p>\n",
"<p>Starting prompt </p>\n": "<p>\u542f\u52a8\u63d0\u793a</p>\n",
"<p>State module to maintain memories when switching between training and validation </p>\n": "<p>\u72b6\u6001\u6a21\u5757\u7528\u4e8e\u5728\u8bad\u7ec3\u548c\u9a8c\u8bc1\u4e4b\u95f4\u5207\u6362\u65f6\u4fdd\u6301\u8bb0\u5fc6</p>\n",
"<p>Take optimizer step </p>\n": "<p>\u91c7\u53d6\u4f18\u5316\u5668\u6b65\u9aa4</p>\n",
"<p>This will keep the accuracy metric stats and memories separate for training and validation. </p>\n": "<p>\u8fd9\u5c06\u4f7f\u7cbe\u5ea6\u6307\u6807\u7edf\u8ba1\u6570\u636e\u548c\u8bb0\u5fc6\u5206\u5f00\uff0c\u4ee5\u4fbf\u8bad\u7ec3\u548c\u9a8c\u8bc1\u3002</p>\n",
"<p>Token embedding module </p>\n": "<p>\u4ee4\u724c\u5d4c\u5165\u6a21\u5757</p>\n",
"<p>Token embedding size </p>\n": "<p>\u4ee4\u724c\u5d4c\u5165\u5927\u5c0f</p>\n",
"<p>Token embeddings </p>\n": "<p>\u4ee4\u724c\u5d4c\u5165</p>\n",
"<p>Tokenize the prompt </p>\n": "<p>\u5c06\u63d0\u793a\u7b26\u53f7\u5316</p>\n",
"<p>Train the model </p>\n": "<p>\u8bad\u7ec3\u6a21\u578b</p>\n",
"<p>Transformer </p>\n": "<p>\u53d8\u538b\u5668</p>\n",
"<p>Truncate old memories </p>\n": "<p>\u622a\u65ad\u65e7\u7684\u8bb0\u5fc6</p>\n",
"<p>Update global step (number of tokens processed) when in training mode </p>\n": "<p>\u5728\u8bad\u7ec3\u6a21\u5f0f\u4e0b\u66f4\u65b0\u5168\u5c40\u6b65\u957f\uff08\u5904\u7406\u7684\u4ee4\u724c\u6570\uff09</p>\n",
"<p>Update memories </p>\n": "<p>\u66f4\u65b0\u8bb0\u5fc6</p>\n",
"<p>Update memory </p>\n": "<p>\u66f4\u65b0\u5185\u5b58</p>\n",
"<p>Use the subsequent mask otherwise </p>\n": "<p>\u5426\u5219\uff0c\u8bf7\u4f7f\u7528\u540e\u7eed\u7684\u63a9\u7801</p>\n",
"<p>Whether to capture model outputs </p>\n": "<p>\u662f\u5426\u6355\u83b7\u6a21\u578b\u8f93\u51fa</p>\n",
"<p>memory </p>\n": "<p>\u8bb0\u5fc6</p>\n",
"This experiment trains a transformer XL model on tiny Shakespeare dataset.": "\u8fd9\u4e2a\u5b9e\u9a8c\u5728\u5fae\u5c0f\u7684\u838e\u58eb\u6bd4\u4e9a\u6570\u636e\u96c6\u4e0a\u8bad\u7ec3\u4e00\u4e2a\u53d8\u538b\u5668 XL \u6a21\u578b\u3002",
"Transformer XL Experiment": "\u53d8\u538b\u5668 XL \u5b9e\u9a8c"
}
@@ -0,0 +1,4 @@
{
"<h1><a href=\"https://nn.labml.ai/transformers/xl/index.html\">Transformer XL</a></h1>\n<p>This is an implementation of <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n<p>Transformer has a limited attention span, equal to the length of the sequence trained in parallel. All these positions have a fixed positional encoding. Transformer XL increases this attention span by letting each of the positions pay attention to precalculated past embeddings. For instance if the context length is <span translate=no>_^_0_^_</span>, it will keep the embeddings of all layers for previous batch of length <span translate=no>_^_1_^_</span> and feed them to current step. If we use fixed-positional encodings these pre-calculated embeddings will have the same positions as the current context. They introduce relative positional encoding, where the positional encodings are introduced at the attention calculation.</p>\n<p>Annotated implementation of relative multi-headed attention is in <a href=\"https://nn.labml.ai/transformers/xl/relative_mha.html\"><span translate=no>_^_2_^_</span></a>.</p>\n<p>Here&#x27;s <a href=\"https://nn.labml.ai/transformers/xl/experiment.html\">the training code</a> and a notebook for training a transformer XL model on Tiny Shakespeare dataset.</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a> </p>\n": "<h1><a href=\"https://nn.labml.ai/transformers/xl/index.html\">\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL</a></h1>\n<p><a href=\"https://pytorch.org\">\u3053\u308c\u306f PyTorch \u306e <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: \u56fa\u5b9a\u9577\u306e\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u3092\u8d85\u3048\u305f\u6ce8\u610f\u6df1\u3044\u8a00\u8a9e\u30e2\u30c7\u30eb\u306e\u5b9f\u88c5\u3067\u3059</a>\u3002</a></p>\n<p>Transformer \u306e\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30b9\u30d1\u30f3\u306f\u3001\u4e26\u884c\u3057\u3066\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3055\u308c\u305f\u30b7\u30fc\u30b1\u30f3\u30b9\u306e\u9577\u3055\u3068\u540c\u3058\u304f\u3089\u3044\u306e\u5236\u9650\u304c\u3042\u308a\u307e\u3059\u3002\u3053\u308c\u3089\u306e\u4f4d\u7f6e\u306f\u3059\u3079\u3066\u56fa\u5b9a\u3055\u308c\u305f\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u306b\u306a\u3063\u3066\u3044\u307e\u3059\u3002Transformer XL\u306f\u3001\u4e8b\u524d\u306b\u8a08\u7b97\u3055\u308c\u305f\u904e\u53bb\u306e\u57cb\u3081\u8fbc\u307f\u306b\u5404\u30dd\u30b8\u30b7\u30e7\u30f3\u306b\u6ce8\u76ee\u3055\u305b\u308b\u3053\u3068\u3067\u3001\u3053\u306e\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30b9\u30d1\u30f3\u3092\u5897\u3084\u3057\u307e\u3059\u3002\u305f\u3068\u3048\u3070\u3001\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u306e\u9577\u3055\u304c\u306e\u5834\u5408<span translate=no>_^_0_^_</span>\u3001<span translate=no>_^_1_^_</span>\u524d\u306e\u30d0\u30c3\u30c1\u306e\u9577\u3055\u306e\u3059\u3079\u3066\u306e\u30ec\u30a4\u30e4\u30fc\u306e\u57cb\u3081\u8fbc\u307f\u3092\u4fdd\u6301\u3057\u3001\u305d\u308c\u3089\u3092\u73fe\u5728\u306e\u30b9\u30c6\u30c3\u30d7\u306b\u9001\u308a\u307e\u3059\u3002\u56fa\u5b9a\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u3092\u4f7f\u7528\u3059\u308b\u3068\u3001\u3053\u308c\u3089\u306e\u4e8b\u524d\u306b\u8a08\u7b97\u3055\u308c\u305f\u57cb\u3081\u8fbc\u307f\u306f\u73fe\u5728\u306e\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u3068\u540c\u3058\u4f4d\u7f6e\u306b\u306a\u308a\u307e\u3059\u3002\u76f8\u5bfe\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u304c\u5c0e\u5165\u3055\u308c\u3001\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u8a08\u7b97\u6642\u306b\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u304c\u5c0e\u5165\u3055\u308c\u307e\u3059</p>\u3002\n<p>\u76f8\u5bfe\u7684\u591a\u9762\u7684\u6ce8\u610f\u306e\u6ce8\u91c8\u4ed8\u304d\u5b9f\u88c5\u304c\u5c0e\u5165\u3055\u308c\u307e\u3057\u305f\u3002<a href=\"https://nn.labml.ai/transformers/xl/relative_mha.html\"><span translate=no>_^_2_^_</span></a></p>\n<p>Tiny <a href=\"https://nn.labml.ai/transformers/xl/experiment.html\">Shakespeare\u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u3067\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fcXL\u30e2\u30c7\u30eb\u3092\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3059\u308b\u305f\u3081\u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u30b3\u30fc\u30c9\u3068\u30ce\u30fc\u30c8\u30d6\u30c3\u30af\u3067\u3059</a>\u3002</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a></p>\n",
"Transformer XL": "\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fc XL"
}
File diff suppressed because one or more lines are too long
@@ -0,0 +1,4 @@
{
"<h1><a href=\"https://nn.labml.ai/transformers/xl/index.html\">Transformer XL</a></h1>\n<p>This is an implementation of <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n<p>Transformer has a limited attention span, equal to the length of the sequence trained in parallel. All these positions have a fixed positional encoding. Transformer XL increases this attention span by letting each of the positions pay attention to precalculated past embeddings. For instance if the context length is <span translate=no>_^_0_^_</span>, it will keep the embeddings of all layers for previous batch of length <span translate=no>_^_1_^_</span> and feed them to current step. If we use fixed-positional encodings these pre-calculated embeddings will have the same positions as the current context. They introduce relative positional encoding, where the positional encodings are introduced at the attention calculation.</p>\n<p>Annotated implementation of relative multi-headed attention is in <a href=\"https://nn.labml.ai/transformers/xl/relative_mha.html\"><span translate=no>_^_2_^_</span></a>.</p>\n<p>Here&#x27;s <a href=\"https://nn.labml.ai/transformers/xl/experiment.html\">the training code</a> and a notebook for training a transformer XL model on Tiny Shakespeare dataset.</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a> </p>\n": "<h1><a href=\"https://nn.labml.ai/transformers/xl/index.html\">\u53d8\u538b\u5668 XL</a></h1>\n<p>\u8fd9\u662f <a href=\"https://pytorch.org\">PyTorch \u4e2d Transfor</a> <a href=\"https://arxiv.org/abs/1901.02860\">mer-XL\uff1a\u8d85\u8d8a\u56fa\u5b9a\u957f\u5ea6\u4e0a\u4e0b\u6587\u7684\u4e13\u5fc3\u8bed\u8a00\u6a21\u578b</a>\u7684\u5b9e\u73b0\u3002</p>\n<p>Transformer \u7684\u6ce8\u610f\u529b\u8de8\u5ea6\u6709\u9650\uff0c\u7b49\u4e8e\u5e76\u884c\u8bad\u7ec3\u5e8f\u5217\u7684\u957f\u5ea6\u3002\u6240\u6709\u8fd9\u4e9b\u4f4d\u7f6e\u90fd\u6709\u56fa\u5b9a\u7684\u4f4d\u7f6e\u7f16\u7801\u3002Transformer XL \u901a\u8fc7\u8ba9\u6bcf\u4e2a\u4f4d\u7f6e\u5173\u6ce8\u8fc7\u53bb\u9884\u5148\u8ba1\u7b97\u7684\u5d4c\u5165\u6b21\u6570\uff0c\u4ece\u800c\u5ef6\u957f\u4e86\u8fd9\u79cd\u6ce8\u610f\u529b\u8de8\u5ea6\u3002\u4f8b\u5982\uff0c\u5982\u679c\u4e0a\u4e0b\u6587\u957f\u5ea6\u4e3a<span translate=no>_^_0_^_</span>\uff0c\u5b83\u5c06\u4fdd\u7559\u524d\u4e00\u6279\u957f\u5ea6\u7684\u6240\u6709\u5c42\u7684\u5d4c\u5165<span translate=no>_^_1_^_</span>\u5e76\u5c06\u5176\u9988\u9001\u5230\u5f53\u524d\u6b65\u9aa4\u3002\u5982\u679c\u6211\u4eec\u4f7f\u7528\u56fa\u5b9a\u4f4d\u7f6e\u7f16\u7801\uff0c\u8fd9\u4e9b\u9884\u5148\u8ba1\u7b97\u7684\u5d4c\u5165\u5c06\u4e0e\u5f53\u524d\u4e0a\u4e0b\u6587\u5177\u6709\u76f8\u540c\u7684\u4f4d\u7f6e\u3002\u5b83\u4eec\u5f15\u5165\u4e86\u76f8\u5bf9\u4f4d\u7f6e\u7f16\u7801\uff0c\u5176\u4e2d\u4f4d\u7f6e\u7f16\u7801\u662f\u5728\u6ce8\u610f\u529b\u8ba1\u7b97\u65f6\u5f15\u5165\u7684\u3002</p>\n<p>\u76f8\u5bf9\u591a\u5934\u6ce8\u610f\u529b\u7684\u5e26\u6ce8\u91ca\u7684\u5b9e\u73b0\u5df2\u7ecf\u5f00\u59cb<a href=\"https://nn.labml.ai/transformers/xl/relative_mha.html\"><span translate=no>_^_2_^_</span></a>\u4e86\u3002</p>\n<p>\u8fd9\u662f\u7528\u4e8e<a href=\"https://nn.labml.ai/transformers/xl/experiment.html\">\u5728 Tiny Shakespeare \u6570\u636e\u96c6\u4e0a\u8bad\u7ec3 transformer XL \u6a21\u578b\u7684\u8bad\u7ec3\u4ee3\u7801</a>\u548c\u7b14\u8bb0\u672c\u3002</p>\n<p><a href=\"https://colab.research.google.com/github/labmlai/annotated_deep_learning_paper_implementations/blob/master/labml_nn/transformers/xl/experiment.ipynb\"><span translate=no>_^_3_^_</span></a></p>\n",
"Transformer XL": "\u53d8\u538b\u5668 XL"
}
@@ -0,0 +1,20 @@
{
"<h1>Relative Multi-Headed Attention</h1>\n<p>This is an implementation of relative multi-headed attention from paper <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n": "<h1>\u76f8\u5bfe\u7684\u591a\u9762\u7684\u6ce8\u610f</h1>\n<p><a href=\"https://pytorch.org\">\u3053\u308c\u306f\u3001\u8ad6\u6587\u300c<a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: PyTorch \u306b\u304a\u3051\u308b\u56fa\u5b9a\u9577\u306e\u30b3\u30f3\u30c6\u30ad\u30b9\u30c8\u3092\u8d85\u3048\u305f\u6ce8\u610f\u306e\u884c\u304d\u5c4a\u3044\u305f\u8a00\u8a9e\u30e2\u30c7\u30eb\u300d\u306e\u6bd4\u8f03\u7684\u591a\u9762\u7684\u306a\u6ce8\u610f\u306e\u5b9f\u88c5\u3067\u3059</a>\u3002</a></p>\n",
"<h2>Relative Multi-Head Attention Module</h2>\n<p>We override <a href=\"mha.html\">Multi-Head Attention</a> module so we only need to write the <span translate=no>_^_0_^_</span> method.</p>\n": "<h2>\u76f8\u5bfe\u30de\u30eb\u30c1\u30d8\u30c3\u30c9\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30e2\u30b8\u30e5\u30fc\u30eb</h2>\n<p><a href=\"mha.html\">\u30de\u30eb\u30c1\u30d8\u30c3\u30c9\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30e2\u30b8\u30e5\u30fc\u30eb\u3092\u30aa\u30fc\u30d0\u30fc\u30e9\u30a4\u30c9\u3059\u308b\u306e\u3067</a>\u3001<span translate=no>_^_0_^_</span>\u30e1\u30bd\u30c3\u30c9\u3092\u8a18\u8ff0\u3059\u308b\u3060\u3051\u3067\u6e08\u307f\u307e\u3059\u3002</p>\n",
"<h3>Get relative attention scores</h3>\n<p>With absolute attention</p>\n<span translate=no>_^_0_^_</span><p>where <span translate=no>_^_1_^_</span>, are linear transformations of original embeddings <span translate=no>_^_2_^_</span> and <span translate=no>_^_3_^_</span> are linear transformations of absolute positional encodings <span translate=no>_^_4_^_</span>.</p>\n<p>They reason out that the attention to a given key should be the same regardless of the position of query. Hence replace <span translate=no>_^_5_^_</span> with a constant <span translate=no>_^_6_^_</span>.</p>\n<p>For the second and third terms relative positional encodings are introduced. So <span translate=no>_^_7_^_</span> is replaced with <span translate=no>_^_8_^_</span> and <span translate=no>_^_9_^_</span> with <span translate=no>_^_10_^_</span>.</p>\n<span translate=no>_^_11_^_</span>": "<h3>\u76f8\u5bfe\u7684\u306a\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30b9\u30b3\u30a2\u3092\u53d6\u5f97</h3>\n<p>\u7d76\u5bfe\u7684\u306a\u6ce8\u610f\u3092\u6255\u3063\u3066</p>\n<span translate=no>_^_0_^_</span><p>\u3053\u3053\u3067<span translate=no>_^_1_^_</span>\u3001<span translate=no>_^_2_^_</span>\u306f\u5143\u306e\u57cb\u3081\u8fbc\u307f\u306e\u7dda\u5f62\u5909\u63db\u3067\u3001<span translate=no>_^_3_^_</span>\u306f\u7d76\u5bfe\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u306e\u7dda\u5f62\u5909\u63db\u3067\u3059\u3002<span translate=no>_^_4_^_</span></p>\n<p>\u5f7c\u3089\u306f\u3001\u7279\u5b9a\u306e\u30ad\u30fc\u3078\u306e\u6ce8\u610f\u306f\u3001\u30af\u30a8\u30ea\u306e\u4f4d\u7f6e\u306b\u95a2\u4fc2\u306a\u304f\u540c\u3058\u3067\u3042\u308b\u3079\u304d\u3060\u3068\u63a8\u8ad6\u3057\u3066\u3044\u307e\u3059\u3002\u3057\u305f\u304c\u3063\u3066\u3001<span translate=no>_^_5_^_</span>\u5b9a\u6570\u306b\u7f6e\u304d\u63db\u3048\u3066\u304f\u3060\u3055\u3044<span translate=no>_^_6_^_</span>\u3002</p>\n<p>\u7b2c2\u7528\u8a9e\u3068\u7b2c3\u7528\u8a9e\u3067\u306f\u3001\u76f8\u5bfe\u4f4d\u7f6e\u30a8\u30f3\u30b3\u30fc\u30c7\u30a3\u30f3\u30b0\u304c\u5c0e\u5165\u3055\u308c\u3066\u3044\u307e\u3059\u3002<span translate=no>_^_7_^_</span>So \u306f\u3001<span translate=no>_^_8_^_</span>\u3068\u3001<span translate=no>_^_9_^_</span>\u306b\u7f6e\u304d\u63db\u3048\u3089\u308c\u307e\u3059<span translate=no>_^_10_^_</span>\u3002</p>\n<span translate=no>_^_11_^_</span>",
"<p> </p>\n": "<p></p>\n",
"<p> This method shifts <span translate=no>_^_0_^_</span> row of a matrix by <span translate=no>_^_1_^_</span> columns.</p>\n<p>If the input is <span translate=no>_^_2_^_</span>, the shifted result would be <span translate=no>_^_3_^_</span>. <em>Ideally we should mask out the lower triangle but it&#x27;s ok for our purpose</em>.</p>\n": "<p>\u3053\u306e\u30e1\u30bd\u30c3\u30c9\u306f\u3001<span translate=no>_^_0_^_</span><span translate=no>_^_1_^_</span>\u884c\u5217\u306e\u884c\u3092\u5217\u3054\u3068\u306b\u30b7\u30d5\u30c8\u3057\u307e\u3059\u3002</p>\n<p>\u5165\u529b\u304c\u306e\u5834\u5408<span translate=no>_^_2_^_</span>\u3001\u30b7\u30d5\u30c8\u3055\u308c\u305f\u7d50\u679c\u306f\u6b21\u306e\u3088\u3046\u306b\u306a\u308a\u307e\u3059\u3002<span translate=no>_^_3_^_</span><em>\u4e0b\u306e\u4e09\u89d2\u5f62\u3092\u30de\u30b9\u30af\u3059\u308b\u306e\u304c\u7406\u60f3\u7684\u3067\u3059\u304c\u3001\u3053\u306e\u76ee\u7684\u306b\u306f\u554f\u984c\u3042\u308a\u307e\u305b\u3093</em>\u3002</p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span></p>\n",
"<p>Concatenate a column of zeros </p>\n": "<p>0 \u306e\u5217\u3092\u9023\u7d50\u3059\u308b</p>\n",
"<p>Number of relative positions </p>\n": "<p>\u76f8\u5bfe\u4f4d\u7f6e\u306e\u6570</p>\n",
"<p>Positional embeddings for the query is independent of the position of the query </p>\n": "<p>\u30af\u30a8\u30ea\u306e\u4f4d\u7f6e\u57cb\u3081\u8fbc\u307f\u306f\u30af\u30a8\u30ea\u306e\u4f4d\u7f6e\u3068\u306f\u7121\u95a2\u4fc2\u3067\u3059</p>\n",
"<p>Relative positional embedding bias for key relative to the query. </p>\n": "<p>\u30af\u30a8\u30ea\u306b\u5bfe\u3059\u308b\u30ad\u30fc\u306e\u76f8\u5bfe\u7684\u306a\u4f4d\u7f6e\u57cb\u3081\u8fbc\u307f\u30d0\u30a4\u30a2\u30b9\u3002</p>\n",
"<p>Relative positional embeddings for key relative to the query. We need <span translate=no>_^_0_^_</span> embeddings because the keys can be before or after the query. </p>\n": "<p>\u30af\u30a8\u30ea\u3092\u57fa\u6e96\u3068\u3057\u305f\u30ad\u30fc\u306e\u76f8\u5bfe\u4f4d\u7f6e\u57cb\u3081\u8fbc\u307f\u3002\u30ad\u30fc\u306f\u30af\u30a8\u30ea\u306e\u524d\u3067\u3082\u5f8c\u3067\u3082\u69cb\u308f\u306a\u3044\u306e\u3067\u3001<span translate=no>_^_0_^_</span>\u57cb\u3081\u8fbc\u307f\u304c\u5fc5\u8981\u3067\u3059</p>\u3002\n",
"<p>Remove extra positions </p>\n": "<p>\u4f59\u5206\u306a\u30dd\u30b8\u30b7\u30e7\u30f3\u3092\u524a\u9664</p>\n",
"<p>Reshape and remove excess elements from the end </p>\n": "<p>\u5f62\u3092\u5909\u3048\u3066\u7aef\u304b\u3089\u4f59\u5206\u306a\u8981\u7d20\u3092\u53d6\u308a\u9664\u304f</p>\n",
"<p>Return the sum <span translate=no>_^_0_^_</span> </p>\n": "<p>\u5408\u8a08\u3092\u8fd4\u3059 <span translate=no>_^_0_^_</span></p>\n",
"<p>Shift the rows of <span translate=no>_^_0_^_</span> to get <span translate=no>_^_1_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span>\u884c\u3092\u30b7\u30d5\u30c8\u3059\u308b\u3068 <span translate=no>_^_1_^_</span></p>\n",
"<p>The linear transformations do not need a bias since we explicitly include it when calculating scores. However having a bias for <span translate=no>_^_0_^_</span> might make sense. </p>\n": "<p>\u7dda\u5f62\u5909\u63db\u306f\u30b9\u30b3\u30a2\u306e\u8a08\u7b97\u6642\u306b\u660e\u793a\u7684\u306b\u542b\u3081\u308b\u306e\u3067\u3001\u30d0\u30a4\u30a2\u30b9\u306f\u5fc5\u8981\u3042\u308a\u307e\u305b\u3093\u3002\u305f\u3060\u3057\u3001<span translate=no>_^_0_^_</span>\u504f\u898b\u3092\u6301\u3064\u3053\u3068\u306f\u7406\u306b\u304b\u306a\u3063\u3066\u3044\u308b\u304b\u3082\u3057\u308c\u307e\u305b\u3093.</p>\n",
"Documented implementation with explanations of Relative Multi-Headed Attention from paper Transformer-XL.": "\u30da\u30fc\u30d1\u30fc\u30fb\u30c8\u30e9\u30f3\u30b9\u30d5\u30a9\u30fc\u30de\u30fcXL\u306e\u76f8\u5bfe\u7684\u30de\u30eb\u30c1\u30d8\u30c3\u30c9\u30fb\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u306e\u8aac\u660e\u3092\u542b\u3080\u5b9f\u88c5\u304c\u6587\u66f8\u5316\u3055\u308c\u3066\u3044\u307e\u3059\u3002",
"Relative Multi-Headed Attention": "\u76f8\u5bfe\u7684\u591a\u9762\u7684\u6ce8\u610f"
}
@@ -0,0 +1,20 @@
{
"<h1>Relative Multi-Headed Attention</h1>\n<p>This is an implementation of relative multi-headed attention from paper <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n": "<h1>\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0db6\u0dc4\u0dd4-\u0dc1\u0dd3\u0dbb\u0dca\u0dc2 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba</h1>\n<p>\u0db8\u0dd9\u0dba\u0d9a\u0da9\u0daf\u0dcf\u0dc3\u0dd2 <a href=\"https://arxiv.org/abs/1901.02860\">\u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca-\u0d91\u0d9a\u0dca\u0dc3\u0dca\u0d91\u0dbd\u0dca \u0dc0\u0dd9\u0dad\u0dd2\u0db1\u0dca \u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0db6\u0dc4\u0dd4-\u0dc1\u0dd3\u0dbb\u0dca\u0dc2 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0dad\u0dca\u0db8\u0d9a \u0d9a\u0dd2\u0dbb\u0dd3\u0db8: <a href=\"https://pytorch.org\">\u0db4\u0dba\u0dd2\u0da7\u0ddd\u0dbb\u0dca\u0da0\u0dca</a> \u0dc4\u0dd2 \u0dc3\u0dca\u0dae\u0dcf\u0dc0\u0dbb \u0daf\u0dd2\u0d9c \u0dc3\u0db1\u0dca\u0daf\u0dbb\u0dca\u0db7\u0dba\u0d9a\u0dd2\u0db1\u0dca \u0d94\u0db6\u0dca\u0db6\u0da7 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dba\u0ddc\u0db8\u0dd4 \u0d9a\u0dbb\u0db1 \u0db7\u0dcf\u0dc2\u0dcf \u0d86\u0d9a\u0dd8\u0dad\u0dd2</a> . </p>\n",
"<h2>Relative Multi-Head Attention Module</h2>\n<p>We override <a href=\"mha.html\">Multi-Head Attention</a> module so we only need to write the <span translate=no>_^_0_^_</span> method.</p>\n": "<h2>\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0db6\u0dc4\u0dd4-\u0dc4\u0dd2\u0dc3 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0db8\u0ddc\u0da9\u0dd2\u0dba\u0dd4\u0dbd\u0dba</h2>\n<p>\u0d85\u0db4\u0dd2 <a href=\"mha.html\">\u0db6\u0dc4\u0dd4-\u0db4\u0dca\u0dbb\u0db0\u0dcf\u0db1 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba</a> \u0db8\u0ddc\u0da9\u0dd2\u0dba\u0dd4\u0dbd\u0dba \u0d85\u0db7\u0dd2\u0db6\u0dc0\u0dcf \u0dba\u0db1 \u0db6\u0dd0\u0dc0\u0dd2\u0db1\u0dca \u0d85\u0db4\u0da7 \u0d85\u0dc0\u0dc1\u0dca\u0dba \u0dc0\u0db1\u0dca\u0db1\u0dda <span translate=no>_^_0_^_</span> \u0d9a\u0dca\u0dbb\u0db8\u0dba \u0dbd\u0dd2\u0dc0\u0dd3\u0db8\u0da7 \u0db4\u0db8\u0dab\u0dd2. </p>\n",
"<h3>Get relative attention scores</h3>\n<p>With absolute attention</p>\n<span translate=no>_^_0_^_</span><p>where <span translate=no>_^_1_^_</span>, are linear transformations of original embeddings <span translate=no>_^_2_^_</span> and <span translate=no>_^_3_^_</span> are linear transformations of absolute positional encodings <span translate=no>_^_4_^_</span>.</p>\n<p>They reason out that the attention to a given key should be the same regardless of the position of query. Hence replace <span translate=no>_^_5_^_</span> with a constant <span translate=no>_^_6_^_</span>.</p>\n<p>For the second and third terms relative positional encodings are introduced. So <span translate=no>_^_7_^_</span> is replaced with <span translate=no>_^_8_^_</span> and <span translate=no>_^_9_^_</span> with <span translate=no>_^_10_^_</span>.</p>\n<span translate=no>_^_11_^_</span>": "<h3>\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dbd\u0d9a\u0dd4\u0dab\u0dd4 \u0dbd\u0db6\u0dcf \u0d9c\u0db1\u0dca\u0db1</h3>\n<p>\u0db1\u0dd2\u0dbb\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba\u0dd9\u0db1\u0dca</p>\n<span translate=no>_^_0_^_</span><p>\u0db8\u0dd4\u0dbd\u0dca\u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dca \u0dc0\u0dbd \u0dbb\u0dda\u0d9b\u0dd3\u0dba \u0db4\u0dbb\u0dd2\u0dc0\u0dbb\u0dca\u0dad\u0db1\u0dba\u0db1\u0dca <span translate=no>_^_3_^_</span> \u0dc0\u0db1 <span translate=no>_^_2_^_</span> \u0d85\u0dad\u0dbb \u0db1\u0dd2\u0dbb\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0dc3\u0dca\u0dae\u0dcf\u0db1\u0dd3\u0dba \u0d9a\u0dda\u0dad\u0dd3\u0d9a\u0dbb\u0dab\u0dba\u0dda \u0dbb\u0dda\u0d9b\u0dd3\u0dba \u0db4\u0dbb\u0dd2\u0dc0\u0dbb\u0dca\u0dad\u0db1\u0dba\u0db1\u0dca \u0dc0\u0dda <span translate=no>_^_1_^_</span> <span translate=no>_^_4_^_</span>. </p>\n<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0dca\u0dad\u0dad\u0dca\u0dad\u0dca\u0dc0\u0dba \u0db1\u0ddc\u0dc3\u0dbd\u0d9a\u0dcf \u0daf\u0dd3 \u0d87\u0dad\u0dd2 \u0dba\u0dad\u0dd4\u0dbb\u0d9a\u0dca \u0d9a\u0dd9\u0dbb\u0dd9\u0dc4\u0dd2 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dba\u0ddc\u0db8\u0dd4 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0db8\u0dcf\u0db1 \u0dc0\u0dd2\u0dba \u0dba\u0dd4\u0dad\u0dd4 \u0db6\u0dc0 \u0d94\u0dc0\u0dd4\u0dc4\u0dd4 \u0db4\u0dd9\u0db1\u0dca\u0dc0\u0dcf \u0daf\u0dd9\u0dad\u0dd2. \u0d91\u0db6\u0dd0\u0dc0\u0dd2\u0db1\u0dca \u0db1\u0dd2\u0dba\u0dad\u0dba\u0d9a\u0dca <span translate=no>_^_5_^_</span> \u0dc3\u0db8\u0d9f \u0db4\u0dca\u0dbb\u0dad\u0dd2\u0dc3\u0dca\u0dae\u0dcf\u0db4\u0db1\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1 <span translate=no>_^_6_^_</span>. </p>\n<p>\u0daf\u0dd9\u0dc0\u0db1\u0dc4\u0dcf \u0dad\u0dd9\u0dc0\u0db1 \u0d9a\u0ddc\u0db1\u0dca\u0daf\u0dda\u0dc3\u0dd2 \u0dc3\u0db3\u0dc4\u0dcf \u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0dc3\u0dca\u0dae\u0dcf\u0db1\u0dd3\u0dba \u0d9a\u0dda\u0dad\u0db1 \u0dc4\u0db3\u0dd4\u0db1\u0dca\u0dc0\u0dcf \u0daf\u0dd9\u0db1\u0dd4 \u0dbd\u0dd0\u0db6\u0dda. \u0d92 <span translate=no>_^_7_^_</span> \u0db1\u0dd2\u0dc3\u0dcf <span translate=no>_^_8_^_</span> \u0dc4\u0dcf <span translate=no>_^_9_^_</span> \u0dc3\u0db8\u0d9f \u0db4\u0dca\u0dbb\u0dad\u0dd2\u0dc3\u0dca\u0dae\u0dcf\u0db4\u0db1\u0dba <span translate=no>_^_10_^_</span>\u0dc0\u0dda. </p>\n<span translate=no>_^_11_^_</span>",
"<p> </p>\n": "<p> </p>\n",
"<p> This method shifts <span translate=no>_^_0_^_</span> row of a matrix by <span translate=no>_^_1_^_</span> columns.</p>\n<p>If the input is <span translate=no>_^_2_^_</span>, the shifted result would be <span translate=no>_^_3_^_</span>. <em>Ideally we should mask out the lower triangle but it&#x27;s ok for our purpose</em>.</p>\n": "<p> \u0db8\u0dd9\u0db8\u0d9a\u0dca\u0dbb\u0db8\u0dba <span translate=no>_^_1_^_</span> \u0dad\u0dd3\u0dbb\u0dd4 \u0db8\u0d9c\u0dd2\u0db1\u0dca \u0d85\u0db1\u0dd4\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0d9a <span translate=no>_^_0_^_</span> \u0db4\u0dda\u0dc5\u0dd2\u0dba \u0db8\u0dcf\u0dbb\u0dd4 \u0d9a\u0dbb\u0dba\u0dd2. </p>\n<p>\u0d86\u0daf\u0dcf\u0db1\u0dba\u0db1\u0db8\u0dca <span translate=no>_^_2_^_</span>, \u0db8\u0dcf\u0dbb\u0dd4 \u0d9a\u0dc5 \u0db4\u0dca\u0dbb\u0dad\u0dd2 result \u0dbd\u0dba \u0dc0\u0db1\u0dd4 <span translate=no>_^_3_^_</span>\u0d87\u0dad. <em>\u0d89\u0dad\u0dcf\u0db8\u0dd0\u0db1\u0dc0\u0dd2\u0db1\u0dca \u0d85\u0db4\u0dd2 \u0db4\u0dc4\u0dc5 \u0dad\u0dca\u0dbb\u0dd2\u0d9a\u0ddd\u0dab\u0dba \u0dc0\u0dc3\u0d82 \u0d9a\u0dc5 \u0dba\u0dd4\u0dad\u0dd4 \u0db1\u0db8\u0dd4\u0dad\u0dca \u0d91\u0dba \u0d85\u0db4\u0d9c\u0dda \u0d85\u0dbb\u0db8\u0dd4\u0dab \u0dc3\u0db3\u0dc4\u0dcf \u0dc4\u0dbb\u0dd2</em>. </p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span> </p>\n",
"<p>Concatenate a column of zeros </p>\n": "<p>\u0dc1\u0dd4\u0db1\u0dca\u0dba\u0dad\u0dd3\u0dbb\u0dd4\u0dc0\u0d9a\u0dca \u0dc3\u0d82\u0dba\u0dd4\u0d9a\u0dca\u0dad \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Number of relative positions </p>\n": "<p>\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0dad\u0db1\u0dad\u0dd4\u0dbb\u0dd4 \u0d9c\u0dab\u0db1 </p>\n",
"<p>Positional embeddings for the query is independent of the position of the query </p>\n": "<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0dc3\u0db3\u0dc4\u0dcf \u0dc3\u0dca\u0dae\u0dcf\u0db1\u0dd3\u0dba \u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dca \u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0dda \u0db4\u0dd2\u0dc4\u0dd2\u0da7\u0dd4\u0db8\u0dd9\u0db1\u0dca \u0dc3\u0dca\u0dc0\u0dcf\u0db0\u0dd3\u0db1 \u0dc0\u0dda </p>\n",
"<p>Relative positional embedding bias for key relative to the query. </p>\n": "<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0da7\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0dc0 \u0dba\u0dad\u0dd4\u0dbb \u0dc3\u0db3\u0dc4\u0dcf \u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0dc3\u0dca\u0dae\u0dcf\u0db1\u0dd3\u0dba \u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dda \u0db1\u0dd0\u0db9\u0dd4\u0dbb\u0dd4\u0dc0. </p>\n",
"<p>Relative positional embeddings for key relative to the query. We need <span translate=no>_^_0_^_</span> embeddings because the keys can be before or after the query. </p>\n": "<p>\u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0da7\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2\u0dc0 \u0dba\u0dad\u0dd4\u0dbb \u0dc3\u0db3\u0dc4\u0dcf \u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0dc3\u0dca\u0dae\u0dcf\u0db1\u0dd3\u0dba \u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dca. \u0d85\u0db4\u0da7 <span translate=no>_^_0_^_</span> \u0d9a\u0dcf\u0dc0\u0dd0\u0daf\u0dca\u0daf\u0dd3\u0db8\u0dca \u0d85\u0dc0\u0dc1\u0dca\u0dba \u0dc0\u0db1\u0dca\u0db1\u0dda \u0dba\u0dad\u0dd4\u0dbb\u0dd4 \u0dc0\u0dd2\u0db8\u0dc3\u0dd3\u0db8\u0da7 \u0db4\u0dd9\u0dbb \u0dc4\u0ddd \u0db4\u0dc3\u0dd4\u0dc0 \u0dc0\u0dd2\u0dba \u0dc4\u0dd0\u0d9a\u0dd2 \u0db6\u0dd0\u0dc0\u0dd2\u0db1\u0dd2. </p>\n",
"<p>Remove extra positions </p>\n": "<p>\u0d85\u0db8\u0dad\u0dbb\u0dad\u0db1\u0dad\u0dd4\u0dbb\u0dd4 \u0d89\u0dc0\u0dad\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Reshape and remove excess elements from the end </p>\n": "<p>\u0d85\u0dc0\u0dc3\u0dcf\u0db1\u0dba\u0dda\u0dc3\u0dd2\u0da7 \u0d85\u0dad\u0dd2\u0dbb\u0dd2\u0d9a\u0dca\u0dad \u0db8\u0dd6\u0dbd\u0daf\u0dca\u0dbb\u0dc0\u0dca\u0dba \u0db1\u0dd0\u0dc0\u0dad \u0dc3\u0d9a\u0dc3\u0dca \u0d9a\u0dbb \u0d89\u0dc0\u0dad\u0dca \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Return the sum <span translate=no>_^_0_^_</span> </p>\n": "<p>\u0db8\u0dd4\u0daf\u0dbd\u0d86\u0db4\u0dc3\u0dd4 \u0daf\u0dd9\u0db1\u0dca\u0db1 <span translate=no>_^_0_^_</span> </p>\n",
"<p>Shift the rows of <span translate=no>_^_0_^_</span> to get <span translate=no>_^_1_^_</span> </p>\n": "<p>\u0dbd\u0db6\u0dcf\u0d9c\u0dd0\u0db1\u0dd3\u0db8\u0da7 \u0db4\u0dda\u0dc5\u0dd2 \u0db8\u0dcf\u0dbb\u0dd4 <span translate=no>_^_0_^_</span> \u0d9a\u0dbb\u0db1\u0dca\u0db1 <span translate=no>_^_1_^_</span> </p>\n",
"<p>The linear transformations do not need a bias since we explicitly include it when calculating scores. However having a bias for <span translate=no>_^_0_^_</span> might make sense. </p>\n": "<p>\u0dbd\u0d9a\u0dd4\u0dab\u0dd4\u0d9c\u0dab\u0db1\u0dba \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda\u0daf\u0dd3 \u0d85\u0db4\u0dd2 \u0db4\u0dd0\u0dc4\u0dd0\u0daf\u0dd2\u0dbd\u0dd2\u0dc0\u0db8 \u0d91\u0dba \u0d87\u0dad\u0dd4\u0dc5\u0dad\u0dca \u0d9a\u0dbb \u0d87\u0dad\u0dd2 \u0db6\u0dd0\u0dc0\u0dd2\u0db1\u0dca \u0dbb\u0dda\u0d9b\u0dd3\u0dba \u0db4\u0dbb\u0dd2\u0dc0\u0dbb\u0dca\u0dad\u0db1\u0dba\u0db1\u0dca\u0da7 \u0db1\u0dd0\u0db9\u0dd4\u0dbb\u0dd4\u0dc0\u0d9a\u0dca \u0d85\u0dc0\u0dc1\u0dca\u0dba \u0db1\u0ddc\u0dc0\u0dda. \u0d9a\u0dd9\u0dc3\u0dda \u0dc0\u0dd9\u0dad\u0dad\u0dca \u0db4\u0d9a\u0dca\u0dc2\u0d9c\u0dca\u0dbb\u0dcf\u0dc4\u0dd3\u0dc0 \u0dc3\u0dd2\u0da7\u0dd3\u0db8 \u0d85\u0dbb\u0dca\u0dae\u0dc0\u0dad\u0dca <span translate=no>_^_0_^_</span> \u0dc0\u0dd2\u0dba \u0dc4\u0dd0\u0d9a\u0dd2\u0dba. </p>\n",
"Documented implementation with explanations of Relative Multi-Headed Attention from paper Transformer-XL.": "\u0d9a\u0da9\u0daf\u0dcf\u0dc3\u0dd2 \u0da7\u0dca\u0dbb\u0dcf\u0db1\u0dca\u0dc3\u0dca\u0dc6\u0ddd\u0db8\u0dbb\u0dca-\u0d91\u0d9a\u0dca\u0dc3\u0dca\u0d91\u0dbd\u0dca \u0dc0\u0dd9\u0dad\u0dd2\u0db1\u0dca \u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0db6\u0dc4\u0dd4-\u0dc1\u0dd3\u0dbb\u0dca\u0dc2 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0db4\u0dd2\u0dc5\u0dd2\u0db6\u0db3 \u0db4\u0dd0\u0dc4\u0dd0\u0daf\u0dd2\u0dbd\u0dd2 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dca \u0dc3\u0db8\u0d9f \u0dbd\u0dda\u0d9b\u0db1\u0d9c\u0dad \u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0dad\u0dca\u0db8\u0d9a \u0d9a\u0dd2\u0dbb\u0dd3\u0db8.",
"Relative Multi-Headed Attention": "\u0dc3\u0dcf\u0db4\u0dda\u0d9a\u0dca\u0dc2 \u0db6\u0dc4\u0dd4-\u0dc1\u0dd3\u0dbb\u0dca\u0dc2 \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba"
}
@@ -0,0 +1,20 @@
{
"<h1>Relative Multi-Headed Attention</h1>\n<p>This is an implementation of relative multi-headed attention from paper <a href=\"https://arxiv.org/abs/1901.02860\">Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context</a> in <a href=\"https://pytorch.org\">PyTorch</a>.</p>\n": "<h1>\u76f8\u5bf9\u591a\u5934\u6ce8\u610f\u529b</h1>\n<p>\u8fd9\u662f paper Transfor <a href=\"https://arxiv.org/abs/1901.02860\">mer-XL\uff1a<a href=\"https://pytorch.org\">PyTorch</a> \u4e2d\u56fa\u5b9a\u957f\u5ea6\u4e0a\u4e0b\u6587\u4e4b\u5916\u7684\u7ec6\u5fc3\u8bed\u8a00\u6a21\u578b</a>\u4e2d\u76f8\u5bf9\u591a\u5934\u5173\u6ce8\u7684\u5b9e\u73b0\u3002</p>\n",
"<h2>Relative Multi-Head Attention Module</h2>\n<p>We override <a href=\"mha.html\">Multi-Head Attention</a> module so we only need to write the <span translate=no>_^_0_^_</span> method.</p>\n": "<h2>\u76f8\u5bf9\u591a\u5934\u6ce8\u610f\u6a21\u5757</h2>\n<p>\u6211\u4eec\u91cd\u5199\u4e86<a href=\"mha.html\">\u591a\u5934\u6ce8\u610f</a>\u6a21\u5757\uff0c\u56e0\u6b64\u6211\u4eec\u53ea\u9700\u8981\u7f16\u5199\u8be5<span translate=no>_^_0_^_</span>\u65b9\u6cd5\u5373\u53ef\u3002</p>\n",
"<h3>Get relative attention scores</h3>\n<p>With absolute attention</p>\n<span translate=no>_^_0_^_</span><p>where <span translate=no>_^_1_^_</span>, are linear transformations of original embeddings <span translate=no>_^_2_^_</span> and <span translate=no>_^_3_^_</span> are linear transformations of absolute positional encodings <span translate=no>_^_4_^_</span>.</p>\n<p>They reason out that the attention to a given key should be the same regardless of the position of query. Hence replace <span translate=no>_^_5_^_</span> with a constant <span translate=no>_^_6_^_</span>.</p>\n<p>For the second and third terms relative positional encodings are introduced. So <span translate=no>_^_7_^_</span> is replaced with <span translate=no>_^_8_^_</span> and <span translate=no>_^_9_^_</span> with <span translate=no>_^_10_^_</span>.</p>\n<span translate=no>_^_11_^_</span>": "<h3>\u83b7\u53d6\u76f8\u5bf9\u6ce8\u610f\u529b\u5206\u6570</h3>\n<p>\u7edd\u5bf9\u5173\u6ce8</p>\n<span translate=no>_^_0_^_</span><p>\u5176\u4e2d<span translate=no>_^_1_^_</span>\uff0c\u662f\u539f\u59cb\u5d4c\u5165\u7684\u7ebf\u6027\u53d8\u6362<span translate=no>_^_2_^_</span>\uff0c<span translate=no>_^_3_^_</span>\u662f\u7edd\u5bf9\u4f4d\u7f6e\u7f16\u7801\u7684\u7ebf\u6027\u53d8\u6362<span translate=no>_^_4_^_</span>\u3002</p>\n<p>\u4ed6\u4eec\u8ba4\u4e3a\uff0c\u65e0\u8bba\u67e5\u8be2\u7684\u4f4d\u7f6e\u5982\u4f55\uff0c\u5bf9\u7ed9\u5b9a\u952e\u7684\u5173\u6ce8\u90fd\u5e94\u8be5\u76f8\u540c\u3002\u56e0\u6b64\uff0c<span translate=no>_^_5_^_</span>\u7528\u5e38\u91cf\u66ff\u6362<span translate=no>_^_6_^_</span>\u3002</p>\n<p>\u5bf9\u4e8e\u7b2c\u4e8c\u9879\u548c\u7b2c\u4e09\u9879\uff0c\u5f15\u5165\u4e86\u76f8\u5bf9\u4f4d\u7f6e\u7f16\u7801\u3002\u56e0\u6b64<span translate=no>_^_7_^_</span>\uff0c\u66ff\u6362<span translate=no>_^_8_^_</span>\u4e3a<span translate=no>_^_9_^_</span>\u548c<span translate=no>_^_10_^_</span>\u3002</p>\n<span translate=no>_^_11_^_</span>",
"<p> </p>\n": "<p></p>\n",
"<p> This method shifts <span translate=no>_^_0_^_</span> row of a matrix by <span translate=no>_^_1_^_</span> columns.</p>\n<p>If the input is <span translate=no>_^_2_^_</span>, the shifted result would be <span translate=no>_^_3_^_</span>. <em>Ideally we should mask out the lower triangle but it&#x27;s ok for our purpose</em>.</p>\n": "<p>\u6b64\u65b9\u6cd5\u5c06\u77e9\u9635\u7684<span translate=no>_^_0_^_</span>\u884c\u6309<span translate=no>_^_1_^_</span>\u5217\u79fb\u52a8\u3002</p>\n<p>\u5982\u679c\u8f93\u5165\u4e3a<span translate=no>_^_2_^_</span>\uff0c\u5219\u79fb\u4f4d\u7684\u7ed3\u679c\u5c06\u4e3a<span translate=no>_^_3_^_</span>\u3002<em>\u7406\u60f3\u60c5\u51b5\u4e0b\uff0c\u6211\u4eec\u5e94\u8be5\u63a9\u76d6\u4e0b\u4e09\u89d2\u5f62\uff0c\u4f46\u8fd9\u5bf9\u6211\u4eec\u7684\u76ee\u7684\u6765\u8bf4\u662f\u53ef\u4ee5\u7684</em>\u3002</p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span></p>\n",
"<p>Concatenate a column of zeros </p>\n": "<p>\u8fde\u63a5\u4e00\u5217\u96f6</p>\n",
"<p>Number of relative positions </p>\n": "<p>\u76f8\u5bf9\u4f4d\u7f6e\u7684\u6570\u91cf</p>\n",
"<p>Positional embeddings for the query is independent of the position of the query </p>\n": "<p>\u67e5\u8be2\u7684\u4f4d\u7f6e\u5d4c\u5165\u4e0e\u67e5\u8be2\u7684\u4f4d\u7f6e\u65e0\u5173</p>\n",
"<p>Relative positional embedding bias for key relative to the query. </p>\n": "<p>\u952e\u76f8\u5bf9\u4e8e\u67e5\u8be2\u7684\u76f8\u5bf9\u4f4d\u7f6e\u5d4c\u5165\u504f\u5dee\u3002</p>\n",
"<p>Relative positional embeddings for key relative to the query. We need <span translate=no>_^_0_^_</span> embeddings because the keys can be before or after the query. </p>\n": "<p>\u952e\u76f8\u5bf9\u4e8e\u67e5\u8be2\u7684\u76f8\u5bf9\u4f4d\u7f6e\u5d4c\u5165\u3002\u6211\u4eec\u9700\u8981<span translate=no>_^_0_^_</span>\u5d4c\u5165\uff0c\u56e0\u4e3a\u952e\u53ef\u4ee5\u5728\u67e5\u8be2\u4e4b\u524d\u6216\u4e4b\u540e\u3002</p>\n",
"<p>Remove extra positions </p>\n": "<p>\u79fb\u9664\u591a\u4f59\u7684\u5934\u5bf8</p>\n",
"<p>Reshape and remove excess elements from the end </p>\n": "<p>\u91cd\u5851\u5e76\u4ece\u672b\u7aef\u79fb\u9664\u591a\u4f59\u7684\u5143\u7d20</p>\n",
"<p>Return the sum <span translate=no>_^_0_^_</span> </p>\n": "<p>\u8fd4\u56de\u603b\u548c<span translate=no>_^_0_^_</span></p>\n",
"<p>Shift the rows of <span translate=no>_^_0_^_</span> to get <span translate=no>_^_1_^_</span> </p>\n": "<p>\u79fb\u52a8\u884c<span translate=no>_^_0_^_</span>\u4ee5\u83b7\u53d6<span translate=no>_^_1_^_</span></p>\n",
"<p>The linear transformations do not need a bias since we explicitly include it when calculating scores. However having a bias for <span translate=no>_^_0_^_</span> might make sense. </p>\n": "<p>\u7ebf\u6027\u53d8\u6362\u4e0d\u9700\u8981\u504f\u5dee\uff0c\u56e0\u4e3a\u6211\u4eec\u5728\u8ba1\u7b97\u5206\u6570\u65f6\u4f1a\u660e\u786e\u5305\u542b\u504f\u5dee\u3002\u4f46\u662f\uff0c\u6709\u504f\u89c1<span translate=no>_^_0_^_</span>\u53ef\u80fd\u662f\u6709\u9053\u7406\u7684\u3002</p>\n",
"Documented implementation with explanations of Relative Multi-Headed Attention from paper Transformer-XL.": "\u8bb0\u5f55\u4e86\u5b9e\u73b0\uff0c\u5e76\u89e3\u91ca\u4e86\u6765\u81ea paper Transformer-XL \u7684\u76f8\u5bf9\u591a\u5934\u6ce8\u610f\u529b\u3002",
"Relative Multi-Headed Attention": "\u76f8\u5bf9\u591a\u5934\u6ce8\u610f\u529b"
}