chore: import upstream snapshot with attribution

This commit is contained in:
wehub-resource-sync
2026-07-13 12:19:01 +08:00
commit 3b90d1192f
2172 changed files with 594509 additions and 0 deletions
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,34 @@
{
"<h1>Graph Attention Networks v2 (GATv2)</h1>\n<p>This is a <a href=\"https://pytorch.org\">PyTorch</a> implementation of the GATv2 operator from the paper <a href=\"https://arxiv.org/abs/2105.14491\">How Attentive are Graph Attention Networks?</a>.</p>\n<p>GATv2s work on graph data similar to <a href=\"../gat/index.html\">GAT</a>. A graph consists of nodes and edges connecting nodes. For example, in Cora dataset the nodes are research papers and the edges are citations that connect the papers.</p>\n<p>The GATv2 operator fixes the static attention problem of the standard <a href=\"../gat/index.html\">GAT</a>. Static attention is when the attention to the key nodes has the same rank (order) for any query node. <a href=\"../gat/index.html\">GAT</a> computes attention from query node <span translate=no>_^_0_^_</span> to key node <span translate=no>_^_1_^_</span> as,</p>\n<span translate=no>_^_2_^_</span><p>Note that for any query node <span translate=no>_^_3_^_</span>, the attention rank (<span translate=no>_^_4_^_</span>) of keys depends only on <span translate=no>_^_5_^_</span>. Therefore the attention rank of keys remains the same (<em>static</em>) for all queries.</p>\n<p>GATv2 allows dynamic attention by changing the attention mechanism,</p>\n<span translate=no>_^_6_^_</span><p>The paper shows that GATs static attention mechanism fails on some graph problems with a synthetic dictionary lookup dataset. It&#x27;s a fully connected bipartite graph where one set of nodes (query nodes) have a key associated with it and the other set of nodes have both a key and a value associated with it. The goal is to predict the values of query nodes. GAT fails on this task because of its limited static attention.</p>\n<p>Here is <a href=\"experiment.html\">the training code</a> for training a two-layer GATv2 on Cora dataset.</p>\n": "<h1>Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 (Gatv2)</h1>\n<p>\u8fd9\u662f <a href=\"https://pytorch.org\">PyTorch</a> \u5bf9 Gatv2 \u8fd0\u7b97\u7b26\u7684\u5b9e\u73b0\uff0c\u6458\u81ea\u300a<a href=\"https://arxiv.org/abs/2105.14491\">\u56fe\u6ce8\u610f\u529b\u7f51\u7edc\u6709\u591a\u4e13\u5fc3\uff1f</a>\u300b</p>\u3002\n<p>Gatv2 \u5904\u7406\u7684\u56fe\u5f62\u6570\u636e\u4e0e <a href=\"../gat/index.html\">GAT</a> \u7c7b\u4f3c\u3002\u56fe\u7531\u8282\u70b9\u548c\u8fde\u63a5\u8282\u70b9\u7684\u8fb9\u7ec4\u6210\u3002\u4f8b\u5982\uff0c\u5728 Cora \u6570\u636e\u96c6\u4e2d\uff0c\u8282\u70b9\u662f\u7814\u7a76\u8bba\u6587\uff0c\u8fb9\u7f18\u662f\u8fde\u63a5\u8bba\u6587\u7684\u5f15\u6587\u3002</p>\n<p>Gatv2 \u64cd\u4f5c\u5458\u4fee\u590d\u4e86\u6807\u51c6 <a href=\"../gat/index.html\">G</a> AT \u7684\u9759\u6001\u6ce8\u610f\u529b\u95ee\u9898\u3002\u9759\u6001\u6ce8\u610f\u529b\u662f\u6307\u4efb\u4f55\u67e5\u8be2\u8282\u70b9\u5bf9\u5173\u952e\u8282\u70b9\u7684\u5173\u6ce8\u7b49\u7ea7\uff08\u987a\u5e8f\uff09\u76f8\u540c\u3002<a href=\"../gat/index.html\">GAT</a> \u5c06\u4ece\u67e5\u8be2\u8282\u70b9<span translate=no>_^_0_^_</span>\u5230\u5173\u952e\u8282\u70b9\u7684\u6ce8\u610f\u529b\u8ba1\u7b97<span translate=no>_^_1_^_</span>\u4e3a\uff0c</p>\n<span translate=no>_^_2_^_</span><p>\u8bf7\u6ce8\u610f\uff0c\u5bf9\u4e8e\u4efb\u4f55\u67e5\u8be2\u8282\u70b9<span translate=no>_^_3_^_</span>\uff0c\u952e\u7684\u6ce8\u610f\u529b\u7b49\u7ea7 (<span translate=no>_^_4_^_</span>) \u4ec5\u53d6\u51b3\u4e8e<span translate=no>_^_5_^_</span>\u3002\u56e0\u6b64\uff0c\u6240\u6709\u67e5\u8be2\u7684\u952e\u7684\u6ce8\u610f\u529b\u7b49\u7ea7\u4fdd\u6301\u4e0d\u53d8\uff08<em>\u9759\u6001</em>\uff09\u3002</p>\n<p>Gatv2 \u901a\u8fc7\u6539\u53d8\u6ce8\u610f\u529b\u673a\u5236\u6765\u5141\u8bb8\u52a8\u6001\u5173\u6ce8\uff0c</p>\n<span translate=no>_^_6_^_</span><p>\u8be5\u8bba\u6587\u8868\u660e\uff0cGAT\u7684\u9759\u6001\u6ce8\u610f\u529b\u673a\u5236\u5728\u5408\u6210\u5b57\u5178\u67e5\u627e\u6570\u636e\u96c6\u7684\u67d0\u4e9b\u56fe\u5f62\u95ee\u9898\u4e0a\u4f1a\u5931\u8d25\u3002\u8fd9\u662f\u4e00\u4e2a\u5b8c\u5168\u8fde\u63a5\u7684\u4e8c\u5206\u56fe\uff0c\u5176\u4e2d\u4e00\u7ec4\u8282\u70b9\uff08\u67e5\u8be2\u8282\u70b9\uff09\u5177\u6709\u4e0e\u4e4b\u5173\u8054\u7684\u5bc6\u94a5\uff0c\u800c\u53e6\u4e00\u7ec4\u8282\u70b9\u65e2\u6709\u952e\u53c8\u6709\u4e0e\u4e4b\u5173\u8054\u7684\u503c\u3002\u76ee\u6807\u662f\u9884\u6d4b\u67e5\u8be2\u8282\u70b9\u7684\u503c\u3002GAT \u65e0\u6cd5\u5b8c\u6210\u6b64\u4efb\u52a1\uff0c\u56e0\u4e3a\u5176\u9759\u6001\u6ce8\u610f\u529b\u6709\u9650\u3002</p>\n<p>\u4ee5\u4e0b\u662f<a href=\"experiment.html\">\u5728 Cora \u6570\u636e\u96c6\u4e0a\u8bad\u7ec3\u53cc\u5c42 Gatv2 \u7684\u8bad\u7ec3\u4ee3\u7801</a>\u3002</p>\n",
"<h2>Graph attention v2 layer</h2>\n<p>This is a single graph attention v2 layer. A GATv2 is made up of multiple such layers. It takes <span translate=no>_^_0_^_</span>, where <span translate=no>_^_1_^_</span> as input and outputs <span translate=no>_^_2_^_</span>, where <span translate=no>_^_3_^_</span>.</p>\n": "<h2>Graph \u6ce8\u610f\u529b v2 \u5c42</h2>\n<p>\u8fd9\u662f\u5355\u56fe\u5173\u6ce8 v2 \u5c42\u3002GATv2 \u7531\u591a\u4e2a\u8fd9\u6837\u7684\u5c42\u7ec4\u6210\u3002\u5b83\u9700\u8981<span translate=no>_^_0_^_</span>\uff0c\u5176\u4e2d<span translate=no>_^_1_^_</span>\u4f5c\u4e3a\u8f93\u5165\u548c\u8f93\u51fa<span translate=no>_^_2_^_</span>\uff0c\u5728\u54ea\u91cc<span translate=no>_^_3_^_</span>\u3002</p>\n",
"<h4>Calculate attention score</h4>\n<p>We calculate these for each head <span translate=no>_^_0_^_</span>. <em>We have omitted <span translate=no>_^_1_^_</span> for simplicity</em>.</p>\n<p><span translate=no>_^_2_^_</span></p>\n<p><span translate=no>_^_3_^_</span> is the attention score (importance) from node <span translate=no>_^_4_^_</span> to node <span translate=no>_^_5_^_</span>. We calculate this for each head.</p>\n<p><span translate=no>_^_6_^_</span> is the attention mechanism, that calculates the attention score. The paper sums <span translate=no>_^_7_^_</span>, <span translate=no>_^_8_^_</span> followed by a <span translate=no>_^_9_^_</span> and does a linear transformation with a weight vector <span translate=no>_^_10_^_</span></p>\n<p><span translate=no>_^_11_^_</span> Note: The paper desrcibes <span translate=no>_^_12_^_</span> as <span translate=no>_^_13_^_</span> which is equivalent to the definition we use here. </p>\n": "<h4>\u8ba1\u7b97\u6ce8\u610f\u529b\u5206\u6570</h4>\n<p>\u6211\u4eec\u4e3a\u6bcf\u4e2a\u5934\u90e8\u8ba1\u7b97\u8fd9\u4e9b<span translate=no>_^_0_^_</span>\u3002<em><span translate=no>_^_1_^_</span>\u4e3a\u7b80\u5355\u8d77\u89c1\uff0c\u6211\u4eec\u7701\u7565\u4e86</em>\u3002</p>\n<p><span translate=no>_^_2_^_</span></p>\n<p><span translate=no>_^_3_^_</span>\u662f\u4ece\u4e00\u4e2a\u8282\u70b9\u5230\u53e6\u4e00\u4e2a\u8282\u70b9\u7684<span translate=no>_^_4_^_</span>\u6ce8\u610f\u529b\u5206\u6570\uff08\u91cd\u8981\u6027\uff09<span translate=no>_^_5_^_</span>\u3002\u6211\u4eec\u4e3a\u6bcf\u4e2a\u5934\u90e8\u8ba1\u7b97\u8fd9\u4e2a\u503c\u3002</p>\n<p><span translate=no>_^_6_^_</span>\u662f\u8ba1\u7b97\u6ce8\u610f\u529b\u5206\u6570\u7684\u6ce8\u610f\u529b\u673a\u5236\u3002\u672c\u6587\u6c42\u548c<span translate=no>_^_7_^_</span>\uff0c<span translate=no>_^_8_^_</span>\u7136\u540e\u662f a\uff0c<span translate=no>_^_9_^_</span>\u7136\u540e\u4f7f\u7528\u6743\u91cd\u5411\u91cf\u8fdb\u884c\u7ebf\u6027\u53d8\u6362<span translate=no>_^_10_^_</span></p>\n<p><span translate=no>_^_11_^_</span>\u6ce8\u610f\uff1a\u672c\u6587\u63cf\u8ff0\u7684\u5185\u5bb9<span translate=no>_^_12_^_</span>\u7b49\u540c<span translate=no>_^_13_^_</span>\u4e8e\u6211\u4eec\u5728\u6b64\u5904\u4f7f\u7528\u7684\u5b9a\u4e49\u3002</p>\n",
"<p><span translate=no>_^_0_^_</span> </p>\n": "<p><span translate=no>_^_0_^_</span></p>\n",
"<p><span translate=no>_^_0_^_</span> gets <span translate=no>_^_1_^_</span> where each node embedding is repeated <span translate=no>_^_2_^_</span> times. </p>\n": "<p><span translate=no>_^_0_^_</span>\u83b7\u53d6\u6bcf\u4e2a\u8282\u70b9\u5d4c\u5165\u91cd\u590d<span translate=no>_^_2_^_</span>\u6b21\u6570<span translate=no>_^_1_^_</span>\u7684\u4f4d\u7f6e\u3002</p>\n",
"<p>Apply dropout regularization </p>\n": "<p>\u5e94\u7528\u8f8d\u5b66\u6b63\u5219\u5316</p>\n",
"<p>Calculate <span translate=no>_^_0_^_</span> <span translate=no>_^_1_^_</span> is of shape <span translate=no>_^_2_^_</span> </p>\n": "<p>\u8ba1\u7b97<span translate=no>_^_0_^_</span><span translate=no>_^_1_^_</span>\u662f\u5f62\u72b6\u7684<span translate=no>_^_2_^_</span></p>\n",
"<p>Calculate final output for each head <span translate=no>_^_0_^_</span> </p>\n": "<p>\u8ba1\u7b97\u6bcf\u4e2a\u5934\u7684\u6700\u7ec8\u8f93\u51fa<span translate=no>_^_0_^_</span></p>\n",
"<p>Calculate the number of dimensions per head </p>\n": "<p>\u8ba1\u7b97\u6bcf\u5934\u7684\u5c3a\u5bf8\u6570</p>\n",
"<p>Concatenate the heads </p>\n": "<p>\u8fde\u63a5\u5934\u90e8</p>\n",
"<p>Dropout layer to be applied for attention </p>\n": "<p>\u8981\u5e94\u7528\u7684\u6389\u843d\u5c42\u4ee5\u5f15\u8d77\u6ce8\u610f</p>\n",
"<p>First we calculate <span translate=no>_^_0_^_</span> for all pairs of <span translate=no>_^_1_^_</span>.</p>\n<p><span translate=no>_^_2_^_</span> gets <span translate=no>_^_3_^_</span> where each node embedding is repeated <span translate=no>_^_4_^_</span> times. </p>\n": "<p>\u9996\u5148\uff0c\u6211\u4eec\u8ba1\u7b97<span translate=no>_^_0_^_</span>\u6240\u6709\u5bf9<span translate=no>_^_1_^_</span>.</p>\n<p><span translate=no>_^_2_^_</span>\u83b7\u53d6\u6bcf\u4e2a\u8282\u70b9\u5d4c\u5165\u91cd\u590d<span translate=no>_^_4_^_</span>\u6b21\u6570<span translate=no>_^_3_^_</span>\u7684\u4f4d\u7f6e\u3002</p>\n",
"<p>If <span translate=no>_^_0_^_</span> is <span translate=no>_^_1_^_</span> the same linear layer is used for the target nodes </p>\n": "<p>\u5982\u679c<span translate=no>_^_0_^_</span>\u662f<span translate=no>_^_1_^_</span>\uff0c\u5219\u4e3a\u76ee\u6807\u8282\u70b9\u4f7f\u7528\u76f8\u540c\u7684\u7ebf\u6027\u5c42</p>\n",
"<p>If we are averaging the multiple heads </p>\n": "<p>\u5982\u679c\u6211\u4eec\u5e73\u5747\u591a\u5934</p>\n",
"<p>If we are concatenating the multiple heads </p>\n": "<p>\u5982\u679c\u6211\u4eec\u8981\u8fde\u63a5\u591a\u4e2a\u5934</p>\n",
"<p>Linear layer for initial source transformation; i.e. to transform the source node embeddings before self-attention </p>\n": "<p>\u7528\u4e8e\u521d\u59cb\u6e90\u53d8\u6362\u7684\u7ebf\u6027\u5c42\uff1b\u5373\u5728\u81ea\u6211\u5173\u6ce8\u4e4b\u524d\u8f6c\u6362\u6e90\u8282\u70b9\u5d4c\u5165</p>\n",
"<p>Linear layer to compute attention score <span translate=no>_^_0_^_</span> </p>\n": "<p>\u7528\u4e8e\u8ba1\u7b97\u6ce8\u610f\u529b\u5206\u6570\u7684\u7ebf\u6027\u56fe\u5c42<span translate=no>_^_0_^_</span></p>\n",
"<p>Mask <span translate=no>_^_0_^_</span> based on adjacency matrix. <span translate=no>_^_1_^_</span> is set to <span translate=no>_^_2_^_</span> if there is no edge from <span translate=no>_^_3_^_</span> to <span translate=no>_^_4_^_</span>. </p>\n": "<p><span translate=no>_^_0_^_</span>\u57fa\u4e8e\u90bb\u63a5\u77e9\u9635\u7684\u63a9\u7801\u3002<span translate=no>_^_1_^_</span><span translate=no>_^_2_^_</span>\u5982\u679c\u6ca1\u6709\u4ece\u5230\u7684\u8fb9\u7f18\uff0c\u5219\u8bbe\u7f6e<span translate=no>_^_3_^_</span>\u4e3a<span translate=no>_^_4_^_</span>\u3002</p>\n",
"<p>Now we add the two tensors to get <span translate=no>_^_0_^_</span> </p>\n": "<p>\u73b0\u5728\u6211\u4eec\u6dfb\u52a0\u4e24\u4e2a\u5f20\u91cf\u6765\u83b7\u5f97<span translate=no>_^_0_^_</span></p>\n",
"<p>Number of nodes </p>\n": "<p>\u8282\u70b9\u6570\u91cf</p>\n",
"<p>Remove the last dimension of size <span translate=no>_^_0_^_</span> </p>\n": "<p>\u79fb\u9664\u5927\u5c0f\u7684\u6700\u540e\u4e00\u4e2a\u7ef4\u5ea6<span translate=no>_^_0_^_</span></p>\n",
"<p>Reshape so that <span translate=no>_^_0_^_</span> is <span translate=no>_^_1_^_</span> </p>\n": "<p>\u91cd\u5851<span translate=no>_^_0_^_</span>\u5c31\u662f\u8fd9\u6837<span translate=no>_^_1_^_</span></p>\n",
"<p>Softmax to compute attention <span translate=no>_^_0_^_</span> </p>\n": "<p>Softmax \u9700\u8981\u8ba1\u7b97\u6ce8\u610f\u529b<span translate=no>_^_0_^_</span></p>\n",
"<p>Take the mean of the heads </p>\n": "<p>\u4ee5\u5934\u8111\u7684\u610f\u601d\u4e3a\u4f8b</p>\n",
"<p>The activation for attention score <span translate=no>_^_0_^_</span> </p>\n": "<p>\u6fc0\u6d3b\u6ce8\u610f\u529b\u5206\u6570<span translate=no>_^_0_^_</span></p>\n",
"<p>The adjacency matrix should have shape <span translate=no>_^_0_^_</span> or<span translate=no>_^_1_^_</span> </p>\n": "<p>\u90bb\u63a5\u77e9\u9635\u7684\u5f62\u72b6\u5e94<span translate=no>_^_0_^_</span>\u4e3a<span translate=no>_^_1_^_</span></p>\n",
"<p>The initial transformations, <span translate=no>_^_0_^_</span> <span translate=no>_^_1_^_</span> for each head. We do two linear transformations and then split it up for each head. </p>\n": "<p>\u6bcf\u4e2a\u5934\u90e8\u7684\u521d\u59cb\u53d8\u6362\u3002<span translate=no>_^_0_^_</span><span translate=no>_^_1_^_</span>\u6211\u4eec\u505a\u4e86\u4e24\u4e2a\u7ebf\u6027\u53d8\u6362\uff0c\u7136\u540e\u5c06\u5176\u62c6\u5206\u4e3a\u6bcf\u4e2a\u5934\u90e8\u3002</p>\n",
"<p>We then normalize attention scores (or coefficients) <span translate=no>_^_0_^_</span></p>\n<p>where <span translate=no>_^_1_^_</span> is the set of nodes connected to <span translate=no>_^_2_^_</span>.</p>\n<p>We do this by setting unconnected <span translate=no>_^_3_^_</span> to <span translate=no>_^_4_^_</span> which makes <span translate=no>_^_5_^_</span> for unconnected pairs. </p>\n": "<p>\u7136\u540e\uff0c\u6211\u4eec\u5c06\u6ce8\u610f\u529b\u5206\u6570\uff08\u6216\u7cfb\u6570\uff09\u5f52\u4e00\u5316<span translate=no>_^_0_^_</span></p>\n<p>\u5176\u4e2d<span translate=no>_^_1_^_</span>\u662f\u8fde\u63a5\u5230\u7684\u8282\u70b9\u96c6<span translate=no>_^_2_^_</span>\u3002</p>\n<p>\u6211\u4eec\u901a\u8fc7<span translate=no>_^_3_^_</span>\u5c06\u672a\u8fde\u63a5\u7684\u914d\u5bf9\u8bbe\u7f6e<span translate=no>_^_5_^_</span>\u4e3a\u672a\u8fde\u63a5<span translate=no>_^_4_^_</span>\u7684\u914d\u5bf9\u6765\u5b9e\u73b0\u6b64\u76ee\u7684\u3002</p>\n",
"<ul><li><span translate=no>_^_0_^_</span>, <span translate=no>_^_1_^_</span> is the input node embeddings of shape <span translate=no>_^_2_^_</span>. </li>\n<li><span translate=no>_^_3_^_</span> is the adjacency matrix of shape <span translate=no>_^_4_^_</span>. We use shape <span translate=no>_^_5_^_</span> since the adjacency is the same for each head. Adjacency matrix represent the edges (or connections) among nodes. <span translate=no>_^_6_^_</span> is <span translate=no>_^_7_^_</span> if there is an edge from node <span translate=no>_^_8_^_</span> to node <span translate=no>_^_9_^_</span>.</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\uff0c<span translate=no>_^_1_^_</span>\u662f shape \u7684\u8f93\u5165\u8282\u70b9\u5d4c\u5165<span translate=no>_^_2_^_</span>\u3002</li>\n<li><span translate=no>_^_3_^_</span>\u662f\u5f62\u72b6\u7684\u90bb\u63a5\u77e9\u9635<span translate=no>_^_4_^_</span>\u3002\u6211\u4eec\u4f7f\u7528\u5f62\u72b6\uff0c<span translate=no>_^_5_^_</span>\u56e0\u4e3a\u6bcf\u4e2a\u5934\u90e8\u7684\u90bb\u63a5\u662f\u76f8\u540c\u7684\u3002\u90bb\u63a5\u77e9\u9635\u8868\u793a\u8282\u70b9\u4e4b\u95f4\u7684\u8fb9\uff08\u6216\u8fde\u63a5\uff09\u3002<span translate=no>_^_6_^_</span><span translate=no>_^_7_^_</span>\u5982\u679c\u8282\u70b9\u4e0e\u8282<span translate=no>_^_8_^_</span>\u70b9\u4e4b\u95f4\u5b58\u5728\u8fb9\u7f18<span translate=no>_^_9_^_</span>\u3002</li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span>, <span translate=no>_^_1_^_</span>, is the number of input features per node </li>\n<li><span translate=no>_^_2_^_</span>, <span translate=no>_^_3_^_</span>, is the number of output features per node </li>\n<li><span translate=no>_^_4_^_</span>, <span translate=no>_^_5_^_</span>, is the number of attention heads </li>\n<li><span translate=no>_^_6_^_</span> whether the multi-head results should be concatenated or averaged </li>\n<li><span translate=no>_^_7_^_</span> is the dropout probability </li>\n<li><span translate=no>_^_8_^_</span> is the negative slope for leaky relu activation </li>\n<li><span translate=no>_^_9_^_</span> if set to <span translate=no>_^_10_^_</span>, the same matrix will be applied to the source and the target node of every edge</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span><span translate=no>_^_1_^_</span>\uff0c\u662f\u6bcf\u4e2a\u8282\u70b9\u7684\u8f93\u5165\u8981\u7d20\u6570</li>\n<li><span translate=no>_^_2_^_</span><span translate=no>_^_3_^_</span>\uff0c\u662f\u6bcf\u4e2a\u8282\u70b9\u7684\u8f93\u51fa\u8981\u7d20\u6570</li>\n<li><span translate=no>_^_4_^_</span><span translate=no>_^_5_^_</span>\uff0c\u662f\u6ce8\u610f\u5934\u7684\u6570\u91cf</li>\n<li><span translate=no>_^_6_^_</span>\u591a\u5934\u7ed3\u679c\u5e94\u8be5\u662f\u4e32\u8054\u8fd8\u662f\u6c42\u5e73\u5747\u503c</li>\n<li><span translate=no>_^_7_^_</span>\u662f\u8f8d\u5b66\u6982\u7387</li>\n<li><span translate=no>_^_8_^_</span>\u662f\u6cc4\u6f0f\u7684 relu \u6fc0\u6d3b\u7684\u8d1f\u659c\u7387</li>\n<li><span translate=no>_^_9_^_</span>\u5982\u679c\u8bbe\u7f6e\u4e3a<span translate=no>_^_10_^_</span>\uff0c\u5219\u540c\u4e00\u77e9\u9635\u5c06\u5e94\u7528\u4e8e\u6bcf\u6761\u8fb9\u7684\u6e90\u8282\u70b9\u548c\u76ee\u6807\u8282\u70b9</li></ul>\n",
"A PyTorch implementation/tutorial of Graph Attention Networks v2.": "Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 \u7684 PyTorch \u5b9e\u73b0/\u6559\u7a0b\u3002",
"Graph Attention Networks v2 (GATv2)": "Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 (GATv2)"
}
@@ -0,0 +1,27 @@
{
"<h1>Train a Graph Attention Network v2 (GATv2) on Cora dataset</h1>\n": "<h1>Cora \u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u3067\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30cd\u30c3\u30c8\u30ef\u30fc\u30af v2 (GATv2) \u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0</h1>\n",
"<h2>Configurations</h2>\n<p>Since the experiment is same as <a href=\"../gat/experiment.html\">GAT experiment</a> but with <a href=\"index.html\">GATv2 model</a> we extend the same configs and change the model.</p>\n": "<h2>\u30b3\u30f3\u30d5\u30a3\u30ae\u30e5\u30ec\u30fc\u30b7\u30e7\u30f3</h2>\n<p><a href=\"../gat/experiment.html\">\u5b9f\u9a13\u306fGAT\u5b9f\u9a13\u3068\u540c\u3058\u3067\u3059\u304c\u3001<a href=\"index.html\">GATv2\u30e2\u30c7\u30eb\u3067\u306f\u540c\u3058\u69cb\u6210\u3092\u62e1\u5f35\u3057\u3066\u30e2\u30c7\u30eb\u3092\u5909\u66f4\u3057\u307e\u3059</a></a>\u3002</p>\n",
"<h2>Graph Attention Network v2 (GATv2)</h2>\n<p>This graph attention network has two <a href=\"index.html\">graph attention layers</a>.</p>\n": "<h2>\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30cd\u30c3\u30c8\u30ef\u30fc\u30af v2 (GATv2)</h2>\n<p>\u3053\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u306b\u306f 2 <a href=\"index.html\">\u3064\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30ec\u30a4\u30e4\u30fc\u304c\u3042\u308a\u307e\u3059</a>\u3002</p>\n",
"<p> </p>\n": "<p></p>\n",
"<p> Create GATv2 model</p>\n": "<p>GATv2 \u30e2\u30c7\u30eb\u306e\u4f5c\u6210</p>\n",
"<p>Activation function </p>\n": "<p>\u30a2\u30af\u30c6\u30a3\u30d9\u30fc\u30b7\u30e7\u30f3\u6a5f\u80fd</p>\n",
"<p>Activation function after first graph attention layer </p>\n": "<p>\u6700\u521d\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30ec\u30a4\u30e4\u30fc\u5f8c\u306e\u30a2\u30af\u30c6\u30a3\u30d9\u30fc\u30b7\u30e7\u30f3\u6a5f\u80fd</p>\n",
"<p>Adam optimizer </p>\n": "<p>\u30a2\u30c0\u30e0\u30fb\u30aa\u30d7\u30c6\u30a3\u30de\u30a4\u30b6\u30fc</p>\n",
"<p>Apply dropout to the input </p>\n": "<p>\u5165\u529b\u306b\u30c9\u30ed\u30c3\u30d7\u30a2\u30a6\u30c8\u3092\u9069\u7528</p>\n",
"<p>Calculate configurations. </p>\n": "<p>\u69cb\u6210\u3092\u8a08\u7b97\u3057\u307e\u3059\u3002</p>\n",
"<p>Create an experiment </p>\n": "<p>\u30c6\u30b9\u30c8\u3092\u4f5c\u6210</p>\n",
"<p>Create configurations </p>\n": "<p>\u69cb\u6210\u306e\u4f5c\u6210</p>\n",
"<p>Dropout </p>\n": "<p>\u30c9\u30ed\u30c3\u30d7\u30a2\u30a6\u30c8</p>\n",
"<p>Final graph attention layer where we average the heads </p>\n": "<p>\u30d8\u30c3\u30c9\u3092\u5e73\u5747\u5316\u3059\u308b\u6700\u5f8c\u306e\u30b0\u30e9\u30d5\u30fb\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30fb\u30ec\u30a4\u30e4\u30fc</p>\n",
"<p>First graph attention layer </p>\n": "<p>\u6700\u521d\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30ec\u30a4\u30e4\u30fc</p>\n",
"<p>First graph attention layer where we concatenate the heads </p>\n": "<p>\u30d8\u30c3\u30c9\u3092\u9023\u7d50\u3059\u308b\u6700\u521d\u306e\u30b0\u30e9\u30d5\u30fb\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30fb\u30ec\u30a4\u30e4\u30fc</p>\n",
"<p>Output layer (without activation) for logits </p>\n": "<p>\u30ed\u30b8\u30c3\u30c8\u306e\u51fa\u529b\u30ec\u30a4\u30e4\u30fc (\u30a2\u30af\u30c6\u30a3\u30d9\u30fc\u30b7\u30e7\u30f3\u306a\u3057)</p>\n",
"<p>Run the training </p>\n": "<p>\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3092\u5b9f\u884c</p>\n",
"<p>Set the model </p>\n": "<p>\u30e2\u30c7\u30eb\u3092\u8a2d\u5b9a\u3059\u308b</p>\n",
"<p>Start and watch the experiment </p>\n": "<p>\u5b9f\u9a13\u3092\u958b\u59cb\u3057\u3066\u898b\u308b</p>\n",
"<p>Whether to share weights for source and target nodes of edges </p>\n": "<p>\u30a8\u30c3\u30b8\u306e\u30bd\u30fc\u30b9\u30ce\u30fc\u30c9\u3068\u30bf\u30fc\u30b2\u30c3\u30c8\u30ce\u30fc\u30c9\u306e\u30a6\u30a7\u30a4\u30c8\u3092\u5171\u6709\u3059\u308b\u304b\u3069\u3046\u304b</p>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the features vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is the adjacency matrix of the form <span translate=no>_^_3_^_</span> or <span translate=no>_^_4_^_</span></li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u306f\u5f62\u72b6\u306e\u7279\u5fb4\u30d9\u30af\u30c8\u30eb\u3067\u3059 <span translate=no>_^_1_^_</span></li>\n<li><span translate=no>_^_2_^_</span><span translate=no>_^_3_^_</span>\u306f\u6b21\u306e\u5f62\u5f0f\u306e\u96a3\u63a5\u884c\u5217\u3067\u3059 <span translate=no>_^_4_^_</span></li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the number of features per node </li>\n<li><span translate=no>_^_1_^_</span> is the number of features in the first graph attention layer </li>\n<li><span translate=no>_^_2_^_</span> is the number of classes </li>\n<li><span translate=no>_^_3_^_</span> is the number of heads in the graph attention layers </li>\n<li><span translate=no>_^_4_^_</span> is the dropout probability </li>\n<li><span translate=no>_^_5_^_</span> if set to True, the same matrix will be applied to the source and the target node of every edge</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u306f\u30ce\u30fc\u30c9\u3042\u305f\u308a\u306e\u30d5\u30a3\u30fc\u30c1\u30e3\u6570</li>\n<li><span translate=no>_^_1_^_</span>\u306f\u6700\u521d\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30ec\u30a4\u30e4\u30fc\u306b\u542b\u307e\u308c\u308b\u30d5\u30a3\u30fc\u30c1\u30e3\u306e\u6570\u3067\u3059</li>\n<li><span translate=no>_^_2_^_</span>\u306f\u30af\u30e9\u30b9\u306e\u6570</li>\n<li><span translate=no>_^_3_^_</span>\u30b0\u30e9\u30d5\u30fb\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30fb\u30ec\u30a4\u30e4\u30fc\u306e\u30d8\u30c3\u30c9\u6570\u3067\u3059</li>\n<li><span translate=no>_^_4_^_</span>\u306f\u8131\u843d\u78ba\u7387\u3067\u3059</li>\n<li><span translate=no>_^_5_^_</span>True \u306b\u8a2d\u5b9a\u3059\u308b\u3068\u3001\u3059\u3079\u3066\u306e\u30a8\u30c3\u30b8\u306e\u30bd\u30fc\u30b9\u30ce\u30fc\u30c9\u3068\u30bf\u30fc\u30b2\u30c3\u30c8\u30ce\u30fc\u30c9\u306b\u540c\u3058\u30de\u30c8\u30ea\u30c3\u30af\u30b9\u304c\u9069\u7528\u3055\u308c\u307e\u3059</li></ul>\n",
"This trains is a Graph Attention Network v2 (GATv2) on Cora dataset": "\u3053\u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u306f\u3001Cora\u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30cd\u30c3\u30c8\u30ef\u30fc\u30afv2\uff08GATv2\uff09\u3067\u3059\u3002",
"Train a Graph Attention Network v2 (GATv2) on Cora dataset": "Cora \u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u3067\u306e\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30cd\u30c3\u30c8\u30ef\u30fc\u30af v2 (GATv2) \u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0"
}
@@ -0,0 +1,27 @@
{
"<h1>Train a Graph Attention Network v2 (GATv2) on Cora dataset</h1>\n<p><a href=\"https://app.labml.ai/run/34b1e2f6ed6f11ebb860997901a2d1e3\"><span translate=no>_^_0_^_</span></a></p>\n": "<h1>\u0d9a\u0ddd\u0dbb\u0dcf\u0daf\u0dad\u0dca\u0dad \u0d9a\u0da7\u0dca\u0da7\u0dbd\u0dba \u0db8\u0dad \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0da2\u0dcf\u0dbd\u0dba v2 (GATV2) \u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4</h1>\n<p><a href=\"https://app.labml.ai/run/34b1e2f6ed6f11ebb860997901a2d1e3\"><span translate=no>_^_0_^_</span></a></p>\n",
"<h2>Configurations</h2>\n<p>Since the experiment is same as <a href=\"../gat/experiment.html\">GAT experiment</a> but with <a href=\"index.html\">GATv2 model</a> we extend the same configs and change the model.</p>\n": "<h2>\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dca</h2>\n<p>\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf\u0db6\u0dd0\u0dbd\u0dd3\u0db8 <a href=\"../gat/experiment.html\">GAT \u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf \u0db6\u0dd0\u0dbd\u0dd3\u0db8\u0da7</a> \u0dc3\u0db8\u0dcf\u0db1 \u0dc0\u0db1 \u0db1\u0db8\u0dd4\u0dad\u0dca <a href=\"index.html\">GATV2 \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba</a> \u0dc3\u0db8\u0d9f \u0d85\u0db4\u0dd2 \u0d91\u0d9a\u0db8 \u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3 \u0daf\u0dd2\u0d9c\u0dd4 \u0d9a\u0dbb \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba \u0dc0\u0dd9\u0db1\u0dc3\u0dca \u0d9a\u0dbb\u0db8\u0dd4. </p>\n",
"<h2>Graph Attention Network v2 (GATv2)</h2>\n<p>This graph attention network has two <a href=\"index.html\">graph attention layers</a>.</p>\n": "<h2>\u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0da2\u0dcf\u0dbd\u0dba v2 (GATV2)</h2>\n<p>\u0db8\u0dd9\u0db8\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0da2\u0dcf\u0dbd\u0dba <a href=\"index.html\">\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb</a>\u0daf\u0dd9\u0d9a\u0d9a\u0dca \u0d87\u0dad. </p>\n",
"<p> </p>\n": "<p> </p>\n",
"<p> Create GATv2 model</p>\n": "<p> GATV2\u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba \u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1</p>\n",
"<p>Activation function </p>\n": "<p>\u0dc3\u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda \u0d9a\u0dcf\u0dbb\u0dca\u0dba\u0dba </p>\n",
"<p>Activation function after first graph attention layer </p>\n": "<p>\u0db4\u0dc5\u0db8\u0dd4\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dae\u0dcf\u0dbb \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb\u0dba\u0dd9\u0db1\u0dca \u0db4\u0dc3\u0dd4 \u0dc3\u0d9a\u0dca\u0dbb\u0dd2\u0dba \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda \u0d9a\u0dcf\u0dbb\u0dca\u0dba\u0dba </p>\n",
"<p>Adam optimizer </p>\n": "<p>\u0d86\u0daf\u0db8\u0dca\u0db4\u0dca\u0dbb\u0dc1\u0dc3\u0dca\u0dad\u0d9a\u0dbb\u0dab\u0dba </p>\n",
"<p>Apply dropout to the input </p>\n": "<p>\u0d86\u0daf\u0dcf\u0db1\u0dba\u0da7\u0d85\u0dad\u0dc4\u0dd0\u0dbb \u0daf\u0dd0\u0db8\u0dd3\u0db8 \u0dba\u0ddc\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Calculate configurations. </p>\n": "<p>\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0dba\u0db1\u0dca\u0d9c\u0dab\u0db1\u0dba \u0d9a\u0dbb\u0db1\u0dca\u0db1. </p>\n",
"<p>Create an experiment </p>\n": "<p>\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf\u0db6\u0dd0\u0dbd\u0dd3\u0db8\u0d9a\u0dca \u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Create configurations </p>\n": "<p>\u0dc0\u0dd2\u0db1\u0dca\u0dba\u0dcf\u0dc3\u0dba\u0db1\u0dca\u0dc3\u0dcf\u0daf\u0db1\u0dca\u0db1 </p>\n",
"<p>Dropout </p>\n": "<p>\u0dc4\u0dd0\u0dbd\u0dd3\u0db8 </p>\n",
"<p>Final graph attention layer where we average the heads </p>\n": "<p>\u0d85\u0db4\u0dd2\u0dc4\u0dd2\u0dc3\u0dca \u0dc3\u0dcf\u0db8\u0dcf\u0db1\u0dca\u0dba\u0dba \u0d91\u0dc4\u0dd2\u0daf\u0dd3 \u0d85\u0dc0\u0dc3\u0db1\u0dca \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb\u0dba </p>\n",
"<p>First graph attention layer </p>\n": "<p>\u0db4\u0dc5\u0db8\u0dd4\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dae\u0dcf\u0dbb \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb\u0dba </p>\n",
"<p>First graph attention layer where we concatenate the heads </p>\n": "<p>\u0d85\u0db4\u0dd2\u0dc4\u0dd2\u0dc3\u0dca concatenate \u0d91\u0dc4\u0dd2\u0daf\u0dd3 \u0db4\u0dc5\u0db8\u0dd4 \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb\u0dba </p>\n",
"<p>Output layer (without activation) for logits </p>\n": "<p>\u0db4\u0dd2\u0dc0\u0dd2\u0dc3\u0dd4\u0db8\u0dca\u0dc3\u0db3\u0dc4\u0dcf \u0db4\u0dca\u0dbb\u0dad\u0dd2\u0daf\u0dcf\u0db1 \u0dc3\u0dca\u0dae\u0dbb\u0dba (\u0dc3\u0d9a\u0dca\u0dbb\u0dd2\u0dba \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0d9a\u0dd2\u0db1\u0dca \u0dad\u0ddc\u0dbb\u0dc0) </p>\n",
"<p>Run the training </p>\n": "<p>\u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4\u0dc0\u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0dad\u0dca\u0db8\u0d9a \u0d9a\u0dbb\u0db1\u0dca\u0db1 </p>\n",
"<p>Set the model </p>\n": "<p>\u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0dc3\u0d9a\u0dc3\u0db1\u0dca\u0db1 </p>\n",
"<p>Start and watch the experiment </p>\n": "<p>\u0d85\u0dad\u0dca\u0dc4\u0daf\u0dcf\u0db6\u0dd0\u0dbd\u0dd3\u0db8 \u0d86\u0dbb\u0db8\u0dca\u0db7 \u0d9a\u0dbb \u0db1\u0dbb\u0db9\u0db1\u0dca\u0db1 </p>\n",
"<p>Whether to share weights for source and target nodes of edges </p>\n": "<p>\u0daf\u0dcf\u0dbb\u0dc0\u0dbd\u0db4\u0dca\u0dbb\u0db7\u0dc0\u0dba \u0dc3\u0dc4 \u0d89\u0dbd\u0d9a\u0dca\u0d9a\u0d9c\u0dad \u0db1\u0ddd\u0da9\u0dca \u0dc3\u0db3\u0dc4\u0dcf \u0db6\u0dbb \u0db6\u0dd9\u0daf\u0dcf \u0d9c\u0dad \u0dba\u0dd4\u0dad\u0dd4\u0daf \u0dba\u0db1\u0dca\u0db1 </p>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the features vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is the adjacency matrix of the form <span translate=no>_^_3_^_</span> or <span translate=no>_^_4_^_</span></li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span> \u0dc4\u0dd0\u0da9\u0dba\u0dda \u0dbd\u0d9a\u0dca\u0dc2\u0dab \u0daf\u0ddb\u0dc1\u0dd2\u0d9a \u0dc0\u0dda <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> \u0dba\u0db1\u0dd4 \u0d86\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0dda \u0d85\u0db1\u0dd4\u0d9a\u0dd8\u0dad\u0dd2\u0dba\u0dda \u0d85\u0db1\u0dd4\u0d9a\u0dd8\u0dad\u0dd2\u0dba <span translate=no>_^_3_^_</span> \u0dc4\u0ddd <span translate=no>_^_4_^_</span></li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the number of features per node </li>\n<li><span translate=no>_^_1_^_</span> is the number of features in the first graph attention layer </li>\n<li><span translate=no>_^_2_^_</span> is the number of classes </li>\n<li><span translate=no>_^_3_^_</span> is the number of heads in the graph attention layers </li>\n<li><span translate=no>_^_4_^_</span> is the dropout probability </li>\n<li><span translate=no>_^_5_^_</span> if set to True, the same matrix will be applied to the source and the target node of every edge</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span> node \u0d91\u0d9a\u0d9a\u0dca \u0db8\u0dad\u0db8 \u0d8a\u0da7 \u0d85\u0daf\u0dcf\u0dbd \u0dc0\u0dd2\u0dc1\u0dda\u0dc2\u0dcf\u0d82\u0d9c \u0dc3\u0d82\u0d9b\u0dca\u0dba\u0dcf\u0dc0 \u0dc0\u0dda </li>\n<li><span translate=no>_^_1_^_</span> \u0db4\u0dc5\u0db8\u0dd4 \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb\u0dba \u0dbd\u0d9a\u0dca\u0dc2\u0dab \u0dc3\u0d82\u0d9b\u0dca\u0dba\u0dcf\u0dc0 \u0dc0\u0dda </li>\n<li><span translate=no>_^_2_^_</span> \u0dba\u0db1\u0dd4 \u0db4\u0db1\u0dca\u0dad\u0dd2 \u0d9c\u0dab\u0db1 </li>\n<li><span translate=no>_^_3_^_</span> \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dae\u0dcf\u0dbb \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dc3\u0dca\u0dae\u0dbb \u0dc0\u0dbd \u0dc4\u0dd2\u0dc3\u0dca \u0d9c\u0dab\u0db1 </li>\n<li><span translate=no>_^_4_^_</span> \u0d85\u0dad\u0dc4\u0dd0\u0dbb \u0daf\u0dd0\u0db8\u0dd3\u0db8\u0dda \u0dc3\u0db8\u0dca\u0db7\u0dcf\u0dc0\u0dd2\u0dad\u0dcf\u0dc0 </li>\n<li><span translate=no>_^_5_^_</span> \u0dc3\u0dad\u0dca\u0dba \u0dbd\u0dd9\u0dc3 \u0dc3\u0d9a\u0dc3\u0dcf \u0d87\u0dad\u0dca\u0db1\u0db8\u0dca, \u0dc3\u0dd1\u0db8 \u0daf\u0dcf\u0dbb\u0dba\u0d9a\u0db8 \u0db4\u0dca\u0dbb\u0db7\u0dc0\u0dba\u0da7 \u0dc3\u0dc4 \u0d89\u0dbd\u0d9a\u0dca\u0d9a\u0d9c\u0dad \u0db1\u0ddd\u0da9\u0dba\u0da7 \u0d91\u0d9a\u0db8 \u0d85\u0db1\u0dd4\u0d9a\u0dd8\u0dad\u0dd2\u0dba \u0dba\u0ddc\u0daf\u0db1\u0dd4 \u0dbd\u0dd0\u0db6\u0dda</li></ul>\n",
"This trains is a Graph Attention Network v2 (GATv2) on Cora dataset": "\u0db8\u0dd9\u0db8 \u0daf\u0dd4\u0db8\u0dca\u0dbb\u0dd2\u0dba \u0d9a\u0ddd\u0dbb\u0dcf \u0daf\u0dad\u0dca\u0dad \u0dc3\u0db8\u0dd4\u0daf\u0dcf\u0dba \u0db8\u0dad \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0da2\u0dcf\u0dbd\u0dba v2 (GATV2) \u0dc0\u0dda",
"Train a Graph Attention Network v2 (GATv2) on Cora dataset": "\u0d9a\u0ddd\u0dbb\u0dcf \u0daf\u0dad\u0dca\u0dad \u0d9a\u0da7\u0dca\u0da7\u0dbd\u0dba \u0db8\u0dad \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0da2\u0dcf\u0dbd\u0dba v2 (GATV2) \u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4"
}
@@ -0,0 +1,27 @@
{
"<h1>Train a Graph Attention Network v2 (GATv2) on Cora dataset</h1>\n": "<h1>\u5728 Cora \u6570\u636e\u96c6\u4e0a\u8bad\u7ec3\u56fe\u6ce8\u610f\u529b\u7f51\u7edc v2 (Gatv2)</h1>\n",
"<h2>Configurations</h2>\n<p>Since the experiment is same as <a href=\"../gat/experiment.html\">GAT experiment</a> but with <a href=\"index.html\">GATv2 model</a> we extend the same configs and change the model.</p>\n": "<h2>\u914d\u7f6e</h2>\n<p>\u7531\u4e8e\u5b9e\u9a8c\u4e0e <a href=\"../gat/experiment.html\">GAT \u5b9e\u9a8c</a>\u76f8\u540c\uff0c\u4f46\u4f7f\u7528 G <a href=\"index.html\">ATv2 \u6a21\u578b</a>\uff0c\u6211\u4eec\u6269\u5c55\u4e86\u76f8\u540c\u7684\u914d\u7f6e\u5e76\u66f4\u6539\u4e86\u6a21\u578b\u3002</p>\n",
"<h2>Graph Attention Network v2 (GATv2)</h2>\n<p>This graph attention network has two <a href=\"index.html\">graph attention layers</a>.</p>\n": "<h2>Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 (GATv2)</h2>\n<p>\u8fd9\u4e2a\u56fe\u5f62\u5173\u6ce8\u7f51\u7edc\u6709\u4e24\u4e2a<a href=\"index.html\">\u56fe\u5f62\u5173\u6ce8\u5c42</a>\u3002</p>\n",
"<p> </p>\n": "<p></p>\n",
"<p> Create GATv2 model</p>\n": "<p>\u521b\u5efa GATv2 \u6a21\u578b</p>\n",
"<p>Activation function </p>\n": "<p>\u6fc0\u6d3b\u529f\u80fd</p>\n",
"<p>Activation function after first graph attention layer </p>\n": "<p>\u7b2c\u4e00\u4e2a\u56fe\u5f62\u5173\u6ce8\u5c42\u4e4b\u540e\u7684\u6fc0\u6d3b\u529f\u80fd</p>\n",
"<p>Adam optimizer </p>\n": "<p>Adam \u4f18\u5316\u5668</p>\n",
"<p>Apply dropout to the input </p>\n": "<p>\u5c06\u4e22\u5931\u5e94\u7528\u4e8e\u8f93\u5165</p>\n",
"<p>Calculate configurations. </p>\n": "<p>\u8ba1\u7b97\u914d\u7f6e\u3002</p>\n",
"<p>Create an experiment </p>\n": "<p>\u521b\u5efa\u5b9e\u9a8c</p>\n",
"<p>Create configurations </p>\n": "<p>\u521b\u5efa\u914d\u7f6e</p>\n",
"<p>Dropout </p>\n": "<p>\u8f8d\u5b66</p>\n",
"<p>Final graph attention layer where we average the heads </p>\n": "<p>\u6700\u540e\u4e00\u5f20\u56fe\u5173\u6ce8\u5c42\uff0c\u6211\u4eec\u5e73\u5747\u5934\u90e8</p>\n",
"<p>First graph attention layer </p>\n": "<p>\u7b2c\u4e00\u4e2a\u56fe\u5f62\u5173\u6ce8\u5c42</p>\n",
"<p>First graph attention layer where we concatenate the heads </p>\n": "<p>\u6211\u4eec\u8fde\u63a5\u5934\u90e8\u7684\u7b2c\u4e00\u4e2a\u56fe\u5f62\u6ce8\u610f\u5c42</p>\n",
"<p>Output layer (without activation) for logits </p>\n": "<p>logits \u7684\u8f93\u51fa\u5c42\uff08\u672a\u6fc0\u6d3b\uff09</p>\n",
"<p>Run the training </p>\n": "<p>\u8fd0\u884c\u8bad\u7ec3</p>\n",
"<p>Set the model </p>\n": "<p>\u8bbe\u7f6e\u6a21\u578b</p>\n",
"<p>Start and watch the experiment </p>\n": "<p>\u5f00\u59cb\u89c2\u770b\u5b9e\u9a8c</p>\n",
"<p>Whether to share weights for source and target nodes of edges </p>\n": "<p>\u662f\u5426\u5171\u4eab\u8fb9\u7684\u6e90\u8282\u70b9\u548c\u76ee\u6807\u8282\u70b9\u7684\u6743\u91cd</p>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the features vectors of shape <span translate=no>_^_1_^_</span> </li>\n<li><span translate=no>_^_2_^_</span> is the adjacency matrix of the form <span translate=no>_^_3_^_</span> or <span translate=no>_^_4_^_</span></li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u662f\u5f62\u72b6\u7684\u7279\u5f81\u5411\u91cf<span translate=no>_^_1_^_</span></li>\n<li><span translate=no>_^_2_^_</span>\u662f\u5f62\u5f0f\u7684\u90bb\u63a5\u77e9\u9635<span translate=no>_^_3_^_</span>\u6216<span translate=no>_^_4_^_</span></li></ul>\n",
"<ul><li><span translate=no>_^_0_^_</span> is the number of features per node </li>\n<li><span translate=no>_^_1_^_</span> is the number of features in the first graph attention layer </li>\n<li><span translate=no>_^_2_^_</span> is the number of classes </li>\n<li><span translate=no>_^_3_^_</span> is the number of heads in the graph attention layers </li>\n<li><span translate=no>_^_4_^_</span> is the dropout probability </li>\n<li><span translate=no>_^_5_^_</span> if set to True, the same matrix will be applied to the source and the target node of every edge</li></ul>\n": "<ul><li><span translate=no>_^_0_^_</span>\u662f\u6bcf\u4e2a\u8282\u70b9\u7684\u8981\u7d20\u6570</li>\n<li><span translate=no>_^_1_^_</span>\u662f\u7b2c\u4e00\u4e2a\u56fe\u5f62\u5173\u6ce8\u5c42\u4e2d\u7684\u8981\u7d20\u6570</li>\n<li><span translate=no>_^_2_^_</span>\u662f\u7c7b\u7684\u6570\u91cf</li>\n<li><span translate=no>_^_3_^_</span>\u662f\u56fe\u8868\u5173\u6ce8\u5c42\u4e2d\u7684\u5934\u90e8\u6570\u91cf</li>\n<li><span translate=no>_^_4_^_</span>\u662f\u8f8d\u5b66\u6982\u7387</li>\n<li><span translate=no>_^_5_^_</span>\u5982\u679c\u8bbe\u7f6e\u4e3a True\uff0c\u5219\u540c\u4e00\u77e9\u9635\u5c06\u5e94\u7528\u4e8e\u6bcf\u6761\u8fb9\u7684\u6e90\u8282\u70b9\u548c\u76ee\u6807\u8282\u70b9</li></ul>\n",
"This trains is a Graph Attention Network v2 (GATv2) on Cora dataset": "\u8fd9\u5217\u706b\u8f66\u662f Cora \u6570\u636e\u96c6\u4e0a\u7684 Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 (GATv2)",
"Train a Graph Attention Network v2 (GATv2) on Cora dataset": "\u5728 Cora \u6570\u636e\u96c6\u4e0a\u8bad\u7ec3\u56fe\u5f62\u6ce8\u610f\u529b\u7f51\u7edc v2 (GATv2)"
}
@@ -0,0 +1,4 @@
{
"<h1><a href=\"https://nn.labml.ai/graphs/gatv2/index.html\">Graph Attention Networks v2 (GATv2)</a></h1>\n<p>This is a <a href=\"https://pytorch.org\">PyTorch</a> implementation of the GATv2 operator from the paper <a href=\"https://arxiv.org/abs/2105.14491\">How Attentive are Graph Attention Networks?</a>.</p>\n<p>GATv2s work on graph data. A graph consists of nodes and edges connecting nodes. For example, in Cora dataset the nodes are research papers and the edges are citations that connect the papers.</p>\n<p>The GATv2 operator fixes the static attention problem of the standard GAT: since the linear layers in the standard GAT are applied right after each other, the ranking of attended nodes is unconditioned on the query node. In contrast, in GATv2, every node can attend to any other node.</p>\n<p>Here is <a href=\"https://nn.labml.ai/graphs/gatv2/experiment.html\">the training code</a> for training a two-layer GATv2 on Cora dataset. </p>\n": "<h1><a href=\"https://nn.labml.ai/graphs/gatv2/index.html\">\u30b0\u30e9\u30d5\u30fb\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30fb\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u30b9 v2 (GATv2)</a></h1>\n<p>\u3053\u308c\u306f\u3001\u300c<a href=\"https://arxiv.org/abs/2105.14491\">\u30b0\u30e9\u30d5\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u306f\u3069\u306e\u7a0b\u5ea6\u6ce8\u610f\u6df1\u3044\u306e\u304b</a>\uff1f\u300d<a href=\"https://pytorch.org\">\u3068\u3044\u3046\u8ad6\u6587\u306eGATv2\u6f14\u7b97\u5b50\u3092PyTorch\u3067\u5b9f\u88c5\u3057\u305f\u3082\u306e\u3067\u3059</a>\u3002</p>\u3002\n<p>GATv2\u306f\u30b0\u30e9\u30d5\u30c7\u30fc\u30bf\u3092\u51e6\u7406\u3057\u307e\u3059\u3002\u30b0\u30e9\u30d5\u306f\u3001\u30ce\u30fc\u30c9\u3068\u30ce\u30fc\u30c9\u3092\u63a5\u7d9a\u3059\u308b\u30a8\u30c3\u30b8\u3067\u69cb\u6210\u3055\u308c\u307e\u3059\u3002\u305f\u3068\u3048\u3070\u3001Cora\u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u3067\u306f\u3001\u30ce\u30fc\u30c9\u306f\u7814\u7a76\u8ad6\u6587\u3067\u3001\u7aef\u306f\u8ad6\u6587\u3092\u3064\u306a\u3050\u5f15\u7528\u3067\u3059</p>\u3002\n<p>GATv2 \u6f14\u7b97\u5b50\u306f\u3001\u6a19\u6e96 GAT \u306e\u9759\u7684\u6ce8\u610f\u306e\u554f\u984c\u3092\u89e3\u6c7a\u3057\u307e\u3059\u3002\u6a19\u6e96 GAT \u306e\u7dda\u5f62\u30ec\u30a4\u30e4\u30fc\u306f\u6b21\u3005\u306b\u9069\u7528\u3055\u308c\u308b\u305f\u3081\u3001\u53c2\u52a0\u30ce\u30fc\u30c9\u306e\u30e9\u30f3\u30af\u4ed8\u3051\u306f\u30af\u30a8\u30ea\u30ce\u30fc\u30c9\u3067\u6761\u4ef6\u4ed8\u3051\u3055\u308c\u307e\u305b\u3093\u3002\u5bfe\u7167\u7684\u306b\u3001GATv2\u3067\u306f\u3001\u3059\u3079\u3066\u306e\u30ce\u30fc\u30c9\u304c\u4ed6\u306e\u30ce\u30fc\u30c9\u306b\u63a5\u7d9a\u3067\u304d\u307e\u3059</p>\u3002\n<p>\u3053\u308c\u306f\u3001<a href=\"https://nn.labml.ai/graphs/gatv2/experiment.html\">Cora\u30c7\u30fc\u30bf\u30bb\u30c3\u30c8\u30672\u5c64GATv2\u3092\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u3059\u308b\u305f\u3081\u306e\u30c8\u30ec\u30fc\u30cb\u30f3\u30b0\u30b3\u30fc\u30c9\u3067\u3059</a>\u3002</p>\n",
"Graph Attention Networks v2 (GATv2)": "\u30b0\u30e9\u30d5\u30fb\u30a2\u30c6\u30f3\u30b7\u30e7\u30f3\u30fb\u30cd\u30c3\u30c8\u30ef\u30fc\u30af\u30b9 v2 (GATv2)"
}
@@ -0,0 +1,4 @@
{
"<h1><a href=\"https://nn.labml.ai/graphs/gatv2/index.html\">Graph Attention Networks v2 (GATv2)</a></h1>\n<p>This is a <a href=\"https://pytorch.org\">PyTorch</a> implementation of the GATv2 operator from the paper <a href=\"https://arxiv.org/abs/2105.14491\">How Attentive are Graph Attention Networks?</a>.</p>\n<p>GATv2s work on graph data. A graph consists of nodes and edges connecting nodes. For example, in Cora dataset the nodes are research papers and the edges are citations that connect the papers.</p>\n<p>The GATv2 operator fixes the static attention problem of the standard GAT: since the linear layers in the standard GAT are applied right after each other, the ranking of attended nodes is unconditioned on the query node. In contrast, in GATv2, every node can attend to any other node.</p>\n<p>Here is <a href=\"https://nn.labml.ai/graphs/gatv2/experiment.html\">the training code</a> for training a two-layer GATv2 on Cora dataset.</p>\n<p><a href=\"https://app.labml.ai/run/34b1e2f6ed6f11ebb860997901a2d1e3\"><span translate=no>_^_0_^_</span></a> </p>\n": "<h1><a href=\"https://nn.labml.ai/graphs/gatv2/index.html\">\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0da2\u0dcf\u0dbd v2 (GATV2)</a></h1>\n<p>\u0db8\u0dd9\u0dbaGATV2 \u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0d9a\u0dbb\u0dd4\u0d9c\u0dda <a href=\"https://pytorch.org\">PyTorch</a> \u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0dad\u0dca\u0db8\u0d9a \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0d9a\u0dd2 <a href=\"https://arxiv.org/abs/2105.14491\">\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dae\u0dcf\u0dbb \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dba\u0ddc\u0db8\u0dd4 \u0d9a\u0dbb\u0db1 \u0da2\u0dcf\u0dbd\u0dba\u0db1\u0dca \u0d9a\u0dd9\u0dad\u0dbb\u0db8\u0dca \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba\u0dd9\u0db1\u0dca \u0dc3\u0dd2\u0da7\u0dd2\u0db1\u0dc0\u0dcf\u0daf? </a>. </p>\n<p>GATV2s\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb \u0daf\u0dad\u0dca\u0dad \u0db8\u0dad \u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf \u0d9a\u0dbb\u0dba\u0dd2. \u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dae\u0dcf\u0dbb\u0dba\u0d9a\u0dca \u0db1\u0ddd\u0da9\u0dca \u0dc3\u0dc4 \u0daf\u0dcf\u0dbb \u0dc3\u0db8\u0dca\u0db6\u0db1\u0dca\u0db0 \u0d9a\u0dbb\u0db1 \u0db1\u0ddd\u0da9\u0dca \u0dc0\u0dbd\u0dd2\u0db1\u0dca \u0dc3\u0db8\u0db1\u0dca\u0dc0\u0dd2\u0dad \u0dc0\u0dda. \u0d8b\u0daf\u0dcf\u0dc4\u0dbb\u0dab\u0dba\u0d9a\u0dca \u0dbd\u0dd9\u0dc3, \u0d9a\u0ddd\u0dbb\u0dcf \u0daf\u0dad\u0dca\u0dad \u0d9a\u0da7\u0dca\u0da7\u0dbd\u0dba\u0dda \u0db1\u0ddd\u0da9\u0dca \u0db4\u0dbb\u0dca\u0dba\u0dda\u0dc2\u0dab \u0db4\u0dad\u0dca\u0dbb\u0dd2\u0d9a\u0dcf \u0dc0\u0db1 \u0d85\u0dad\u0dbb \u0daf\u0dcf\u0dbb \u0dba\u0db1\u0dd4 \u0db4\u0dad\u0dca\u0dbb\u0dd2\u0d9a\u0dcf \u0dc3\u0db8\u0dca\u0db6\u0db1\u0dca\u0db0 \u0d9a\u0dbb\u0db1 \u0d8b\u0db4\u0dd4\u0da7\u0dcf \u0daf\u0dd0\u0d9a\u0dca\u0dc0\u0dd3\u0db8\u0dca \u0dc0\u0dda. </p>\n<p>GATV2\u0d9a\u0dca\u0dbb\u0dd2\u0dba\u0dcf\u0d9a\u0dbb\u0dd4 \u0dc3\u0db8\u0dca\u0db8\u0dad GAT \u0dc4\u0dd2 \u0dc3\u0dca\u0dae\u0dd2\u0dad\u0dd2\u0d9a \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0dba\u0ddc\u0db8\u0dd4 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8\u0dda \u0d9c\u0dd0\u0da7\u0dc5\u0dd4\u0dc0 \u0db1\u0dd2\u0dc0\u0dd0\u0dbb\u0daf\u0dd2 \u0d9a\u0dbb\u0dba\u0dd2: \u0dc3\u0db8\u0dca\u0db8\u0dad GAT \u0dc4\u0dd2 \u0dbb\u0dda\u0d9b\u0dd3\u0dba \u0dc3\u0dca\u0dae\u0dbb \u0d91\u0d9a\u0dd2\u0db1\u0dd9\u0d9a\u0da7 \u0db4\u0dc3\u0dd4\u0dc0 \u0dba\u0ddc\u0daf\u0db1 \u0db6\u0dd0\u0dc0\u0dd2\u0db1\u0dca, \u0dc3\u0dc4\u0db7\u0dcf\u0d9c\u0dd3 \u0dc0\u0dd6 \u0db1\u0ddd\u0da9\u0dca \u0dc0\u0dbd \u0dc1\u0dca\u0dbb\u0dda\u0dab\u0dd2\u0d9c\u0dad \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc0\u0dd2\u0db8\u0dc3\u0dd4\u0db8\u0dca \u0db1\u0ddd\u0da9\u0dba \u0db8\u0dad \u0d9a\u0ddc\u0db1\u0dca\u0daf\u0dda\u0dc3\u0dd2 \u0dc0\u0dd2\u0dbb\u0dc4\u0dd2\u0dad\u0dc0 \u0db4\u0dc0\u0dad\u0dd3. \u0d8a\u0da7 \u0dc0\u0dd9\u0db1\u0dc3\u0dca\u0dc0, GATV2 \u0dc4\u0dd2, \u0dc3\u0dd1\u0db8 \u0db1\u0ddd\u0da9\u0dba\u0d9a\u0da7\u0db8 \u0dc0\u0dd9\u0db1\u0dad\u0dca \u0d95\u0db1\u0dd1\u0db8 \u0db1\u0ddd\u0da9\u0dba\u0d9a\u0da7 \u0dc3\u0dc4\u0db7\u0dcf\u0d9c\u0dd3 \u0dc0\u0dd2\u0dba \u0dc4\u0dd0\u0d9a\u0dd2\u0dba. </p>\n<p><a href=\"https://nn.labml.ai/graphs/gatv2/experiment.html\">\u0d9a\u0ddd\u0dbb\u0dcf \u0daf\u0dad\u0dca\u0dad \u0d9a\u0da7\u0dca\u0da7\u0dbd\u0dba\u0dda \u0dc3\u0dca\u0dae\u0dbb \u0daf\u0dd9\u0d9a\u0d9a GATV2 \u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 \u0d9a\u0dd2\u0dbb\u0dd3\u0db8 \u0dc3\u0db3\u0dc4\u0dcf \u0db4\u0dd4\u0dc4\u0dd4\u0dab\u0dd4 \u0d9a\u0dda\u0dad\u0dba</a> \u0db8\u0dd9\u0db1\u0dca\u0db1. </p>\n<p><a href=\"https://app.labml.ai/run/34b1e2f6ed6f11ebb860997901a2d1e3\"><span translate=no>_^_0_^_</span></a> </p>\n",
"Graph Attention Networks v2 (GATv2)": "\u0db4\u0dca\u0dbb\u0dc3\u0dca\u0dad\u0dcf\u0dbb\u0dba \u0d85\u0dc0\u0db0\u0dcf\u0db1\u0dba \u0da2\u0dcf\u0dbd v2 (GATV2)"
}
@@ -0,0 +1,4 @@
{
"<h1><a href=\"https://nn.labml.ai/graphs/gatv2/index.html\">Graph Attention Networks v2 (GATv2)</a></h1>\n<p>This is a <a href=\"https://pytorch.org\">PyTorch</a> implementation of the GATv2 operator from the paper <a href=\"https://arxiv.org/abs/2105.14491\">How Attentive are Graph Attention Networks?</a>.</p>\n<p>GATv2s work on graph data. A graph consists of nodes and edges connecting nodes. For example, in Cora dataset the nodes are research papers and the edges are citations that connect the papers.</p>\n<p>The GATv2 operator fixes the static attention problem of the standard GAT: since the linear layers in the standard GAT are applied right after each other, the ranking of attended nodes is unconditioned on the query node. In contrast, in GATv2, every node can attend to any other node.</p>\n<p>Here is <a href=\"https://nn.labml.ai/graphs/gatv2/experiment.html\">the training code</a> for training a two-layer GATv2 on Cora dataset. </p>\n": "<h1><a href=\"https://nn.labml.ai/graphs/gatv2/index.html\">Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 (Gatv2)</a></h1>\n<p>\u8fd9\u662f <a href=\"https://pytorch.org\">PyTorch</a> \u5bf9 Gatv2 \u8fd0\u7b97\u7b26\u7684\u5b9e\u73b0\uff0c\u6458\u81ea\u300a<a href=\"https://arxiv.org/abs/2105.14491\">\u56fe\u6ce8\u610f\u529b\u7f51\u7edc\u6709\u591a\u4e13\u5fc3\uff1f</a>\u300b</p>\u3002\n<p>Gatv2 \u5904\u7406\u56fe\u8868\u6570\u636e\u3002\u56fe\u7531\u8282\u70b9\u548c\u8fde\u63a5\u8282\u70b9\u7684\u8fb9\u7ec4\u6210\u3002\u4f8b\u5982\uff0c\u5728 Cora \u6570\u636e\u96c6\u4e2d\uff0c\u8282\u70b9\u662f\u7814\u7a76\u8bba\u6587\uff0c\u8fb9\u7f18\u662f\u8fde\u63a5\u8bba\u6587\u7684\u5f15\u6587\u3002</p>\nG@@ <p>atv2 \u8fd0\u7b97\u7b26\u4fee\u590d\u4e86\u6807\u51c6 GAT \u7684\u9759\u6001\u6ce8\u610f\u529b\u95ee\u9898\uff1a\u7531\u4e8e\u6807\u51c6 GAT \u4e2d\u7684\u7ebf\u6027\u5c42\u662f\u7d27\u63a5\u5e94\u7528\u7684\uff0c\u56e0\u6b64\u6709\u4eba\u503c\u5b88\u8282\u70b9\u7684\u6392\u540d\u4e0d\u53d7\u67e5\u8be2\u8282\u70b9\u7684\u9650\u5236\u3002\u76f8\u6bd4\u4e4b\u4e0b\uff0c\u5728 Gatv2 \u4e2d\uff0c\u6bcf\u4e2a\u8282\u70b9\u90fd\u53ef\u4ee5\u7ba1\u7406\u4efb\u4f55\u5176\u4ed6\u8282\u70b9\u3002</p>\n<p>\u4ee5\u4e0b\u662f<a href=\"https://nn.labml.ai/graphs/gatv2/experiment.html\">\u5728 Cora \u6570\u636e\u96c6\u4e0a\u8bad\u7ec3\u53cc\u5c42 Gatv2 \u7684\u8bad\u7ec3\u4ee3\u7801</a>\u3002</p>\n",
"Graph Attention Networks v2 (GATv2)": "Graph \u6ce8\u610f\u529b\u7f51\u7edc v2 (GATv2)"
}