{"id":333,"date":"2015-08-19T22:07:22","date_gmt":"2015-08-19T22:07:22","guid":{"rendered":"http:\/\/www.marekrei.com\/blog\/?p=333"},"modified":"2019-09-27T23:34:58","modified_gmt":"2019-09-27T23:34:58","slug":"26-things-i-learned-in-the-deep-learning-summer-school","status":"publish","type":"post","link":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/","title":{"rendered":"26 Things I Learned in the Deep Learning Summer School"},"content":{"rendered":"<p>In the beginning of August I got the chance to attend the Deep Learning Summer School in Montreal. It consisted of 10 days of talks from some of the most well-known\u00a0neural network researchers. During this time I learned a lot, way more\u00a0than I could ever fit into a blog post. Instead of trying to pass on 60 hours worth of neural network knowledge, I\u00a0have made a list of small interesting nuggets of information that I was able to summarise in a paragraph.<\/p>\n<p>At the moment of writing, the <a href=\"https:\/\/sites.google.com\/site\/deeplearningsummerschool\/schedule\">summer school website<\/a> is still online, along with all the presentation slides. All of the information and most of the illustrations come from these slides and are the work of their\u00a0original authors.\u00a0The talks in the summer school were filmed as well, hopefully they will also find their way to the web.<\/p>\n<p><strong>Update<\/strong>: <a href=\"http:\/\/videolectures.net\/deeplearning2015_montreal\/\">the Deep Learning Summer School videos are now online<\/a>.<\/p>\n<p>Alright, let&#8217;s get started.<\/p>\n<h2>1. The need for distributed representations<\/h2>\n<p>During his first talk, Yoshua Bengio said &#8220;This is my most important slide&#8221;. You can see that slide below:<a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-334\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png\" alt=\"dlss-3aug2015\" width=\"1100\" height=\"850\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png 1100w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015-150x116.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015-300x232.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015-1024x791.png 1024w\" sizes=\"auto, (max-width: 1100px) 100vw, 1100px\" \/><\/a><\/p>\n<p>Let&#8217;s say you have a classifier that needs to detect people that are male\/female, have glasses or\u00a0don&#8217;t have glasses, and are tall\/short. With non-distributed representations, you are dealing with 2*2*2=8 different classes of people. In order to train an accurate classifier, you need to have enough training data for each of these 8 classes. However, with distributed representations, each of these properties could be captured by a different dimension. This means that even if your classifier has never encountered tall men with glasses, it would be able to detect them, because it has learned to detect gender, glasses and height independently from all the other examples.<br \/>\n<!--more--><\/p>\n<h2>2. Local minima are not a problem in high dimensions<\/h2>\n<p>The team of Yoshua Bengio have experimentally found that when optimising the parameters of high-dimensional neural nets, there effectively are no local minima. Instead, there are saddle points which are local minima in some dimensions but not all. This means that training can slow down quite a lot in these points, until the network figures out how to escape, but as long as we&#8217;re willing to wait long enough then it will find a way.<\/p>\n<p>Below is a graph demonstrating a network during training, oscillating between two states: approaching a saddle point and then escaping it.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug20152.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-340\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug20152.png\" alt=\"dlss-3aug20152\" width=\"867\" height=\"461\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug20152.png 867w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug20152-150x80.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug20152-300x160.png 300w\" sizes=\"auto, (max-width: 867px) 100vw, 867px\" \/><\/a><\/p>\n<p>Given one specific dimension, there is some small probability \\(p\\) with which a point is a local minimum, but not a global minimum, in that dimension. Now, the probability of a point in a 1000-dimensional space being an incorrect local minimum in <strong>all<\/strong> of these would be \\(p^{1000}\\), which is just astronomically small. However, the probability of it being a local minimum in <strong>some<\/strong> of these dimensions is actually quite high. And when we get these minima in many dimensions at once, then training can appear to be stuck until it finds the right direction.<\/p>\n<p>In addition, this probability \\(p\\) will increase as the loss function\u00a0gets closer to the global minimum. \u00a0This means that if we do ever end up at a genuine local minimum, then for all intents and purposes it will be close enough to the global minimum that it will not matter.<\/p>\n<h2>3. Derivatives derivatives derivatives<\/h2>\n<p>Leon Bottou had some useful tables with activation functions, loss functions, and their corresponding derivatives. I&#8217;ll keep these here for later.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-337\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou1.png\" alt=\"bottou1\" width=\"838\" height=\"349\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou1.png 838w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou1-150x62.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou1-300x125.png 300w\" sizes=\"auto, (max-width: 838px) 100vw, 838px\" \/><\/a><\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou2.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-338\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou2.png\" alt=\"bottou2\" width=\"870\" height=\"332\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou2.png 870w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou2-150x57.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/bottou2-300x114.png 300w\" sizes=\"auto, (max-width: 870px) 100vw, 870px\" \/><\/a><\/p>\n<p>&nbsp;<\/p>\n<p><strong>Update<\/strong>: As pointed out by commenters, the min and max functions in the ramp formula should be switched.<\/p>\n<h2>4. Weight initialisation strategy<\/h2>\n<p>The current recommended strategy for initialising weights in a neural network is to sample values \\(W_{i,j}^{(k)}\\)\u00a0uniformly from \\([-b,b]\\), where<\/p>\n<p style=\"text-align: center;\">\\(b = \\sqrt{\\frac{6}{H_k + H_{k+1}}}\\)<\/p>\n<p>\\(H_k\\) and \\(H_{k+1}\\) are the sizes of hidden layers before and after the weight matrix.<\/p>\n<p>Recommended by\u00a0Hugo Larochelle, published by\u00a0Glorot &amp; Bengio\u00a0(2010).<\/p>\n<h2>5. Neural net training tricks<\/h2>\n<p>A few practical suggestions from\u00a0Hugo Larochelle:<\/p>\n<ul>\n<li>Normalise real-valued data. Subtract the mean and divide by standard deviation.<\/li>\n<li>Decrease the learning rate during training.<\/li>\n<li>Can update using mini-batches &#8211; the gradient is more stable.<\/li>\n<li>Can use momentum, to get through plateaus.<\/li>\n<\/ul>\n<h2>6. Gradient checking<\/h2>\n<p>If you implemented your backprop by hand and it&#8217;s not working, then there&#8217;s roughly 99% chance\u00a0that the gradient\u00a0calculation has\u00a0a bug. Use gradient checking to identify the issue. The idea is to use the definition of a gradient: how much will the model error change, if we increase a specific weight by a small amount.<\/p>\n<p style=\"text-align: center;\">\\(\\frac{\\partial f(x)}{\\partial x} \\approx\u00a0\\frac{f(x+\\epsilon) &#8211; f(x-\\epsilon)}{2\\epsilon}\\)<\/p>\n<p>A more in-depth explanation is available here: <a href=\"http:\/\/ufldl.stanford.edu\/wiki\/index.php\/Gradient_checking_and_advanced_optimization\">Gradient checking and advanced optimization<\/a><\/p>\n<h2>7. Motion tracking<\/h2>\n<p>Human motion tracking can be done with impressive accuracy. Below are examples from the paper\u00a0<a href=\"http:\/\/www.uoguelph.ca\/~gwtaylor\/publications\/cvpr2010\/gwtaylor_cvpr2010.pdf\">Dynamical Binary Latent Variable Models for 3D Human Pose Tracking<\/a> by Graham Taylor et al. (2010). The method uses\u00a0conditional restricted\u00a0Boltzmann machines.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/gwtaylor_cvpr2010.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-356\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/gwtaylor_cvpr2010.png\" alt=\"gwtaylor_cvpr2010\" width=\"811\" height=\"359\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/gwtaylor_cvpr2010.png 811w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/gwtaylor_cvpr2010-150x66.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/gwtaylor_cvpr2010-300x133.png 300w\" sizes=\"auto, (max-width: 811px) 100vw, 811px\" \/><\/a><\/p>\n<h2>8. Syntax or no syntax? (aka, &#8220;is syntax a thing?&#8221;)<\/h2>\n<p>Chris Manning and\u00a0Richard Socher have put a lot of effort into developing compositional models that combine neural embeddings with more traditional parsing\u00a0approaches. This culminated with a\u00a0<a href=\"http:\/\/nlp.stanford.edu\/~socherr\/EMNLP2013_RNTN.pdf\">Recursive Neural Tensor Network<\/a>\u00a0(Socher et al., 2013), which uses both additive and multiplicative interactions to combine word meanings along a parse tree.<\/p>\n<p>And then, the model was beaten (by quite a margin) by the\u00a0<a href=\"http:\/\/arxiv.org\/abs\/1405.4053\">Paragraph Vector<\/a> (Le &amp; Mikolov, 2014), which knows absolutely nothing about the sentence structure or syntax. Chris Manning referred to this result as &#8220;a defeat for creating &#8216;good&#8217; compositional vectors&#8221;.<\/p>\n<p>However, more recent work using parse trees has again surpassed this result.\u00a0<a href=\"http:\/\/www.cs.cornell.edu\/~oirsoy\/files\/nips14drsv.pdf\">Irsoy &amp; Cardie (NIPS, 2014)<\/a> managed to beat paragraph vectors by going &#8220;deep&#8221; with their networks in multiple dimensions. Finally, <a href=\"https:\/\/aclweb.org\/anthology\/P\/P15\/P15-1150.pdf\">Tai et al. (ACL, 2015)<\/a> have improved the results again by combining LSTMs with parse trees.<\/p>\n<p>The accuracies\u00a0of these models on the Stanford 5-class sentiment dataset are as follows:<\/p>\n<p>[table width=&#8221;500&#8243; colwidth=&#8221;100|50&#8243; colalign=&#8221;left|center&#8221;]<br \/>\nMethod, Accuracy<br \/>\nRNTN (Socher et al. 2013), 45.7<br \/>\nParagraph Vector (Le &amp; Mikolov 2014),\u00a048.7<br \/>\nDRNN (Irsoy &amp; Cardie 2014), 49.8<br \/>\nTree LSTM (Tai et al. 2015), 50.9<br \/>\n[\/table]<\/p>\n<p>So it seems that, at the moment, models using the parse tree are beating simpler approaches. I&#8217;m curious to see if and when the next syntax-free approach\u00a0will emerge that will advance this race. After all, the goal of many neural models is not to discard the underlying grammar, but to implicitly capture it in the same network.<\/p>\n<h2>9. Distributed vs distributional<\/h2>\n<p>Chris Manning himself cleared up the confusion between the two words.<\/p>\n<p><strong>Distributed<\/strong>: A concept is represented as continuous activation levels in a number of elements. Like a dense word embedding, as opposed to 1-hot vectors.<\/p>\n<p><strong>Distributional<\/strong>:\u00a0Meaning is represented by contexts of use. Word2vec is distributional, but so are count-based word vectors, as we use the contexts of the word to model the meaning.<\/p>\n<h2>10. The state of dependency parsing<\/h2>\n<p>Comparison of dependency parsers on the Penn Treebank:<\/p>\n<p>[table colalign=&#8221;left|center|center&#8221;]<br \/>\nParser,\u00a0Unlabelled Accuracy, Labelled Acccuracy,\u00a0Speed (sent\/s)<br \/>\nMaltParser, 89.8, 87.2, 469<br \/>\nMSTParser, 91.4, 88.1, 10<br \/>\nTurboParser, 92.3, 89.6, 8<br \/>\n<a href=\"http:\/\/cs.stanford.edu\/~danqi\/papers\/emnlp2014.pdf\">Stanford Neural Dependency Parser<\/a>, 92.0, 89.7, 654<br \/>\nGoogle,\u00a094.3, 92.4, ?<br \/>\n[\/table]<\/p>\n<p>The last result is from Google &#8220;pulling out all the stops&#8221;, by\u00a0putting massive amounts of resources into training the Stanford neural parser.<\/p>\n<h2>11. Theano<\/h2>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/14439811.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-370 size-full aligncenter\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/14439811.png\" alt=\"\" width=\"399\" height=\"289\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/14439811.png 399w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/14439811-150x109.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/14439811-300x217.png 300w\" sizes=\"auto, (max-width: 399px) 100vw, 399px\" \/><\/a><\/p>\n<p>Well, I knew a bit about <a href=\"http:\/\/deeplearning.net\/software\/theano\/\">Theano<\/a> before, but I learned a whole lot more during the summer school. And it is pretty awesome.<\/p>\n<p>Since\u00a0Theano\u00a0originates from\u00a0Montreal, it was especially helpful to be able to ask questions directly from the people who are developing it.<\/p>\n<p>Most of the information that was presented is available online, in the form of <a href=\"https:\/\/github.com\/mila-udem\/summerschool2015\">interactive python tutorials<\/a>.<\/p>\n<h2>12. Nvidia Digits<\/h2>\n<p>Nvidia has a toolkit called <a href=\"https:\/\/developer.nvidia.com\/digits\">Digits<\/a> that trains and visualises complex neural network models without needing to write any code.\u00a0And they&#8217;re selling <a href=\"https:\/\/developer.nvidia.com\/devbox\">DevBox<\/a> &#8211; a machine customised for running Digits and other deep learning software (Theano, Caffe, etc). It comes with 4 Titan X GPUs and currently costs $15,000.<\/p>\n<h2>13. Fuel<\/h2>\n<p><a href=\"https:\/\/github.com\/mila-udem\/fuel\">Fuel<\/a> is a toolkit that manages iteration over your datasets &#8211; it can split them into minibatches, manage shuffling, apply various preprocessing steps, etc. There are prebuilt functions for some established datasets, such as MNIST, CIFAR-10, and Google&#8217;s 1B Word corpus.\u00a0It is mainly designed for use with <a href=\"http:\/\/github.com\/mila-udem\/blocks\">Blocks<\/a>, a toolkit that simplifies network construction with Theano.<\/p>\n<h2>14. Multimodal linguistic regularities<\/h2>\n<p>Remember &#8220;king &#8211; man + woman = queen&#8221;? Turns out that works with images as well (Kiros et al., 2015).<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/talk_Montreal_part2_pdf.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-371\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/talk_Montreal_part2_pdf.png\" alt=\"talk_Montreal_part2_pdf\" width=\"1788\" height=\"1164\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/talk_Montreal_part2_pdf.png 1788w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/talk_Montreal_part2_pdf-150x98.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/talk_Montreal_part2_pdf-300x195.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/talk_Montreal_part2_pdf-1024x667.png 1024w\" sizes=\"auto, (max-width: 1788px) 100vw, 1788px\" \/><\/a><\/p>\n<h2>15.\u00a0Taylor series approximation<\/h2>\n<p>When we are at point \\(x_0\\) and take a step to \\(x\\), then we can estimate the function value in the new location by knowing the derivatives, using the Taylor series approximation.<\/p>\n<p style=\"text-align: center;\">\\(<br \/>\nf(x) = f(x_0) + (x &#8211; x_0)f'(x) + \\frac{1}{2}\u00a0(x &#8211; x_0)^2 f&#8221;(x) + &#8230;<br \/>\n\\)<\/p>\n<p>Similarly, we can estimate\u00a0the loss of a function, when we update parameters \\(\\theta_0\\) to \\(\\theta\\).<\/p>\n<p style=\"text-align: center;\">\\(<br \/>\nJ(\\theta)\u00a0=J(\\theta_0) + (\\theta &#8211; \\theta_0)^T g + \\frac{1}{2}\u00a0(\\theta &#8211; \\theta_0)^T\u00a0H(\\theta &#8211; \\theta_0) + &#8230;<br \/>\n\\)<\/p>\n<p>where \\(g\\) contains the derivatives with respect to \\(\\theta\\), and \\(H\\) is the Hessian with second order derivatives\u00a0with respect to \\(\\theta\\).<\/p>\n<p>This is the second-order Taylor approximation, but we could increase the accuracy by adding even higher-order derivatives.<\/p>\n<h2>16. Computational intensity<\/h2>\n<p>Adam Coates presented a strategy for analysing the speed of\u00a0matrix operations on a GPU. It&#8217;s a simplified model that says your time is spent on either reading\/writing to memory or doing calculations. It assumes you can do both in parallel so\u00a0we are interested in which one of them takes more time.<\/p>\n<p>Let&#8217;s say we are multiplying a matrix with a vector:<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_1.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-376 size-medium aligncenter\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_1-300x209.png\" alt=\"dlss_systems_1\" width=\"300\" height=\"209\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_1-300x209.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_1-150x105.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_1.png 972w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/a><\/p>\n<p>If \\(M=1024\\) and \\(N=512\\), then the number of bytes we need to read and store is:<\/p>\n<p style=\"text-align: center;\">\\( 4\\text{ bytes }\\times (1024\u00a0\\times\u00a0512 + 512 + 1024) = 2.1e6\\text{ bytes} \\)<\/p>\n<p>And the number of calculations\u00a0we need to do is:<\/p>\n<p style=\"text-align: center;\">\\(2\\times 1024\\times 512 = 1e6\\text{ FLOPs}\\)<\/p>\n<p>If we have a GPU that can do\u00a06 TFLOP\/s and has memory bandwidth of\u00a0 300GB\/s, then the\u00a0total running time will be:<\/p>\n<p style=\"text-align: center;\">\\(\\text{max}\\{2.1e6\\text{ bytes }\/ (300e9\\text{ bytes}\/s), 1e6\\text{ FLOPs} \/ (6e12\\text{ FLOP}\/s) \\} \\\\<br \/>\n= \\text{max}\\{ 7\\mu\u00a0s, 0.16\\mu s \\} \\)<\/p>\n<p>This means the process is bounded by the \\(7\\mu s\\) spent on copying to\/from the\u00a0memory, and\u00a0getting a faster GPU would not make any difference. As you can probably guess, this situation gets better with bigger matrices\/vectors, and when doing matrix-matrix operations.<\/p>\n<p>Adam also described the idea of calculating the intensity of an operation:<\/p>\n<p style=\"text-align: center;\">Intensity = (# arithmetic ops) \/ (# bytes to load or store)<\/p>\n<p>In the previous scenario, this would be<\/p>\n<p style=\"text-align: center;\">Intensity = (1E6 FLOPs) \/ (2.1E6 bytes) = 0.5 FLOPs\/bytes<\/p>\n<p>Low intensity means the system is bottlenecked on memory, and high intensity means it&#8217;s bottlenecked by the GPU speed. This can be visualised, in order to find which of the two needs to improve in order to speed up the whole system, and where the sweet spot lies.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_2.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-377\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_2.png\" alt=\"dlss_systems_2\" width=\"1572\" height=\"609\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_2.png 1572w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_2-150x58.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_2-300x116.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_2-1024x397.png 1024w\" sizes=\"auto, (max-width: 1572px) 100vw, 1572px\" \/><\/a><\/p>\n<h2>17. Minibatches<\/h2>\n<p>Continuing from the intensity calculations, one way of increasing the intensity of your network\u00a0(in order to be limited by computation instead of memory), is to process data in minibatches. This avoids some memory operations, and GPUs are great at processing large matrices in parallel.<\/p>\n<p>However, increasing the batch size too much will probably start to hurting the training algorithm and converging can take longer. It&#8217;s important to find a good balance in order to get the best results in the least amount of time.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_3.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-388\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_3.png\" alt=\"dlss_systems_3\" width=\"1316\" height=\"682\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_3.png 1316w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_3-150x78.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_3-300x155.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss_systems_3-1024x531.png 1024w\" sizes=\"auto, (max-width: 1316px) 100vw, 1316px\" \/><\/a><\/p>\n<h2>18. Training on adversarial examples<\/h2>\n<p>It was recently revealed that\u00a0neural networks are easily tricked by adversarial examples. In the example below, the image on the left is correctly classified as a goldfish. However, if we apply the noise pattern shown in the middle, resulting in the image on the right, the classifier becomes convinced this is a picture of a daisy. The image is from\u00a0Andrej Karpathy&#8217;s blog post\u00a0<a href=\"http:\/\/karpathy.github.io\/2015\/03\/30\/breaking-convnets\/\">&#8220;Breaking Linear Classifiers on ImageNet&#8221;<\/a>, and you can read more about it there.\u00a0<a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/fish.jpeg\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-395\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/fish.jpeg\" alt=\"fish\" width=\"1033\" height=\"327\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/fish.jpeg 1033w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/fish-150x47.jpeg 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/fish-300x95.jpeg 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/fish-1024x324.jpeg 1024w\" sizes=\"auto, (max-width: 1033px) 100vw, 1033px\" \/><\/a><\/p>\n<p>The noise pattern isn&#8217;t random though &#8211; the noise is carefully calculated, in order to trick the network. But the point remains: the image on the right is clearly still a goldfish and not a daisy.<\/p>\n<p>Apparently strategies like ensemble models, voting after multiple saccades, and unsupervised pretraining have all failed against this vulnerability. Applying heavy regularisation helps, but\u00a0not before ruining the accuracy on the clean data.<\/p>\n<p>Ian Goodfellow presented the idea of training on these adversarial examples.\u00a0They can be automatically generated and added to the training set. The results below show that in addition to helping with the adversarial cases, this also improves accuracy on the clean examples.<a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/goodfellow_adv.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-396\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/goodfellow_adv.png\" alt=\"goodfellow_adv\" width=\"1166\" height=\"780\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/goodfellow_adv.png 1166w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/goodfellow_adv-150x100.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/goodfellow_adv-300x201.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/goodfellow_adv-1024x685.png 1024w\" sizes=\"auto, (max-width: 1166px) 100vw, 1166px\" \/><\/a>Finally, we can improve this further by penalising the KL-divergence between the original predicted distribution and the predicted distribution on the adversarial example. This optimises the network to be more robust, and to predict similar class distributions for similar (adversarial) images.<\/p>\n<h2>19. Everything is language modelling<\/h2>\n<p>Phil Blunsom presented the idea that almost all NLP can be structured as a language model. We can do this by concatenating the output to the input and trying to predict the probability of the whole sequence.<\/p>\n<p>Translation:<\/p>\n<p style=\"text-align: center;\">\\(P(\\text{Les chiens aiment les os || Dogs love bones})\\)<\/p>\n<p style=\"text-align: left;\">Question answering:<\/p>\n<p style=\"text-align: center;\">\\(P(\\text{What do dogs love? || bones .})\\)<\/p>\n<p style=\"text-align: left;\">Dialogue:<\/p>\n<p style=\"text-align: center;\">\\(P(\\text{How are you? || Fine thanks. And you?})\\)<\/p>\n<p>The latter two need to be additionally conditioned on some world knowledge. The second part doesn&#8217;t even need to be words, but could be labels or some structured output like dependency relations.<\/p>\n<h2>20. SMT had a rough start<\/h2>\n<p>When Frederick Jelinek and his team at IBM submitted one of the first papers on statistical machine translation to COLING in 1988, they got the following anonymous review:<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt2.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-402 size-medium aligncenter\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt2-270x300.png\" alt=\"blunsom-lm-mt2\" width=\"270\" height=\"300\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt2-270x300.png 270w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt2-135x150.png 135w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt2.png 872w\" sizes=\"auto, (max-width: 270px) 100vw, 270px\" \/><\/a><\/p>\n<p><em>The validity of a statistical (information theoretic) approach to MT has indeed been\u00a0<\/em><em>recognized, as the authors mention, by Weaver as early as 1949. And was universally\u00a0<\/em><em>recognized as mistaken by 1950 (cf. Hutchins, MT \u2013 Past, Present, Future, Ellis\u00a0<\/em><em>Horwood, 1986, p. 30ff and references therein). The crude force of computers is not\u00a0<\/em><em>science. The paper is simply beyond the scope of COLING.<\/em><\/p>\n<h2>21. The state of Neural Machine Translation<\/h2>\n<p>Apparently a very simple neural model can produce surprisingly good results. An example of translating from Chinese to English, from Phil Blunsom&#8217;s slides:<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-403\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt.png\" alt=\"blunsom-lm-mt\" width=\"924\" height=\"684\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt.png 924w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt-150x111.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/blunsom-lm-mt-300x222.png 300w\" sizes=\"auto, (max-width: 924px) 100vw, 924px\" \/><\/a><\/p>\n<p>In this model, the vectors for the Chinese words are simply added together to form a sentence vector. The decoder\u00a0consists of a conditional language model which takes the sentence vector, together with vectors from the two recently generated English words, and generates the next word in the translation.<\/p>\n<p>However, neural models are still not outperforming the very best traditional MT systems. They do come very close though. Results from <a href=\"http:\/\/papers.nips.cc\/paper\/5346-sequence-to-sequence-learning-with-neural-networks.pdf\">&#8220;Sequence to Sequence Learning<\/a>\u00a0<a href=\"http:\/\/papers.nips.cc\/paper\/5346-sequence-to-sequence-learning-with-neural-networks.pdf\">with Neural Networks&#8221;<\/a>\u00a0by Sutskever et al. (2014):<\/p>\n<p>[table]<br \/>\nModel,\u00a0BLEU score<br \/>\nBaseline, 33.30<br \/>\nBest WMT&#8217;14 result,\u00a037.0<br \/>\nScoring with\u00a05 LSTMs,\u00a036.5<br \/>\nOracle (upper bound),\u00a0\u223c45<br \/>\n[\/table]<\/p>\n<p><strong>Update:<\/strong> <a href=\"https:\/\/twitter.com\/stanfordnlp\">@stanfordnlp<\/a> pointed out that there are some recent\u00a0results where the neural model does indeed outperform the state-of-the-art traditional MT system. Check out &#8220;<a href=\"http:\/\/arxiv.org\/pdf\/1508.04025.pdf\">Effective Approaches to Attention-based Neural Machine Translation<\/a>&#8221; (Luong et. al., 2015).<\/p>\n<h2>22. MetaMind classifier demo<\/h2>\n<p>Richard Socher demonstrated the MetaMind\u00a0<a href=\"https:\/\/www.metamind.io\/vision\/train\">image classification demo<\/a>, which you can train yourself by uploading images. I trained a classifier to detect Edison and Einstein (couldn&#8217;t find enough unique images of Tesla). 5 example images for both classes, testing on one held out image each. Seemed to work pretty well.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/edison_vs_einstein.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-406\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/edison_vs_einstein.png\" alt=\"edison_vs_einstein\" width=\"1158\" height=\"417\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/edison_vs_einstein.png 1158w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/edison_vs_einstein-150x54.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/edison_vs_einstein-300x108.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/edison_vs_einstein-1024x369.png 1024w\" sizes=\"auto, (max-width: 1158px) 100vw, 1158px\" \/><\/a><\/p>\n<h2>23. Optimising gradient updates<\/h2>\n<p>Mark Schmidt gave two presentations about numerical optimisation in different scenarios.<\/p>\n<p>In a <strong>deterministic<\/strong> gradient method we calculate the gradient over the whole dataset and then apply the update. The iteration cost is linear with the dataset size.<\/p>\n<p>In\u00a0<strong>stochastic<\/strong> gradient methods we calculate the gradient on one datapoint and then apply the update. The iteration cost is independent of the dataset size.<\/p>\n<p>Each iteration of the stochastic gradient descent is much faster, but it usually takes many more iterations to train the network, as this graph illustrates:<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/2015_DLSS_ConvexOptimization.png\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-408 size-medium aligncenter\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/2015_DLSS_ConvexOptimization-300x233.png\" alt=\"2015_DLSS_ConvexOptimization\" width=\"300\" height=\"233\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/2015_DLSS_ConvexOptimization-300x233.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/2015_DLSS_ConvexOptimization-150x117.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/2015_DLSS_ConvexOptimization-1024x796.png 1024w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/2015_DLSS_ConvexOptimization.png 1088w\" sizes=\"auto, (max-width: 300px) 100vw, 300px\" \/><\/a><\/p>\n<p>In order to get the best of both worlds, we can use batching. More specifically, we could do one pass of the dataset with stochastic gradient descent, in order to quickly get on the right track, and then start increasing the batch size. The gradient error decreases as the batch size increases, although eventually the iteration cost will become dependent on the dataset size again.<\/p>\n<p>Stochastic Average Gradient (SAG) is a method that gets around this, providing a linear convergence rate with only 1 gradient per iteration. Unfortunately, it is not feasible for large neural networks, as it needs to remember the gradient updates\u00a0for every datapoint, leading to large memory requirements.\u00a0Stochastic Variance-Reduced Gradient (SVRG) is another method that reduces this memory cost, and\u00a0only needs 2 gradient calculations per iteration (plus occasional full passes).<\/p>\n<p>Mark said a student of his implemented a variety of optimisation methods (AdaGrad, momentum, SAG, etc). When asked, what he would use in a black box neural network system, the student said two methods: <a href=\"http:\/\/arxiv.org\/pdf\/1412.6606.pdf\">Streaming SVRG<\/a> (Frostig et al., 2015), and a method they haven&#8217;t published yet.<\/p>\n<h2>24. Theano profiling<\/h2>\n<p>If you put &#8220;profile=True&#8221; into THEANO_FLAGS, it will analyse your program, showing a breakdown of how much is spent on each operation. Very handy for finding bottlenecks.<\/p>\n<h2>25.\u00a0Adversarial nets framework<\/h2>\n<p>Following on from Ian Goodfellow&#8217;s talk on adversarial examples, Yoshua Bengio talked about having two systems competing\u00a0with each other.<\/p>\n<p>System D is a discriminative system that aims to classify between real data and artificially generated data.<\/p>\n<p>System G is a generative system, that tries to generate artificial data, which D would incorrectly classify as real.<\/p>\n<p>As we train one, the other needs to get better as well. In practice this does work, although\u00a0the step needs to be quite small to make sure D can keep up with G. Below are some examples from &#8220;<a href=\"http:\/\/arxiv.org\/abs\/1506.05751\">Deep Generative Image Models using a\u00a0Laplacian Pyramid of Adversarial Networks<\/a>&#8221; &#8211; a more advanced version of this model which tries to generate images of churches.<\/p>\n<p><a href=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/denton_generating_images.png\"><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-409\" src=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/denton_generating_images.png\" alt=\"denton_generating_images\" width=\"1456\" height=\"802\" srcset=\"https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/denton_generating_images.png 1456w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/denton_generating_images-150x83.png 150w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/denton_generating_images-300x165.png 300w, https:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/denton_generating_images-1024x564.png 1024w\" sizes=\"auto, (max-width: 1456px) 100vw, 1456px\" \/><\/a><\/p>\n<h2>26. arXiv.org numbering<\/h2>\n<p>The arXiv number contains the year and month of the submission, followed by the sequence number. So paper\u00a01508.03854 was number\u00a03854 in August 2015. Good to know.<\/p>\n<h2><\/h2>\n","protected":false},"excerpt":{"rendered":"<p>In the beginning of August I got the chance to attend the Deep Learning Summer School in Montreal. It consisted of 10 days of talks&hellip;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-333","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v23.7 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>26 Things I Learned in the Deep Learning Summer School - Marek Rei<\/title>\n<meta name=\"description\" content=\"Here is a list of small interesting nuggets of information from the summer school, regarding neural networks and deep learning.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"26 Things I Learned in the Deep Learning Summer School - Marek Rei\" \/>\n<meta property=\"og:description\" content=\"Here is a list of small interesting nuggets of information from the summer school, regarding neural networks and deep learning.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/\" \/>\n<meta property=\"og:site_name\" content=\"Marek Rei\" \/>\n<meta property=\"article:published_time\" content=\"2015-08-19T22:07:22+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2019-09-27T23:34:58+00:00\" \/>\n<meta property=\"og:image\" content=\"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png\" \/>\n<meta name=\"author\" content=\"Marek\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Marek\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"15 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/\",\"url\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/\",\"name\":\"26 Things I Learned in the Deep Learning Summer School - Marek Rei\",\"isPartOf\":{\"@id\":\"https:\/\/www.marekrei.com\/blog\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#primaryimage\"},\"image\":{\"@id\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#primaryimage\"},\"thumbnailUrl\":\"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png\",\"datePublished\":\"2015-08-19T22:07:22+00:00\",\"dateModified\":\"2019-09-27T23:34:58+00:00\",\"author\":{\"@id\":\"https:\/\/www.marekrei.com\/blog\/#\/schema\/person\/a145eb0a06ed4acf5b0f84a24b7a1191\"},\"description\":\"Here is a list of small interesting nuggets of information from the summer school, regarding neural networks and deep learning.\",\"breadcrumb\":{\"@id\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#primaryimage\",\"url\":\"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png\",\"contentUrl\":\"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/www.marekrei.com\/blog\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"26 Things I Learned in the Deep Learning Summer School\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/www.marekrei.com\/blog\/#website\",\"url\":\"https:\/\/www.marekrei.com\/blog\/\",\"name\":\"Marek Rei\",\"description\":\"Thoughts on Machine Learning and Natural Language Processing\",\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/www.marekrei.com\/blog\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Person\",\"@id\":\"https:\/\/www.marekrei.com\/blog\/#\/schema\/person\/a145eb0a06ed4acf5b0f84a24b7a1191\",\"name\":\"Marek\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.marekrei.com\/blog\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/48a65414bfda6485aaa0703e548de0ed25292b5fe0d979ed8c28ad83cf5a82c0?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/48a65414bfda6485aaa0703e548de0ed25292b5fe0d979ed8c28ad83cf5a82c0?s=96&d=mm&r=g\",\"caption\":\"Marek\"},\"url\":\"https:\/\/www.marekrei.com\/blog\/author\/marek\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"26 Things I Learned in the Deep Learning Summer School - Marek Rei","description":"Here is a list of small interesting nuggets of information from the summer school, regarding neural networks and deep learning.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/","og_locale":"en_US","og_type":"article","og_title":"26 Things I Learned in the Deep Learning Summer School - Marek Rei","og_description":"Here is a list of small interesting nuggets of information from the summer school, regarding neural networks and deep learning.","og_url":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/","og_site_name":"Marek Rei","article_published_time":"2015-08-19T22:07:22+00:00","article_modified_time":"2019-09-27T23:34:58+00:00","og_image":[{"url":"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png"}],"author":"Marek","twitter_misc":{"Written by":"Marek","Est. reading time":"15 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/","url":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/","name":"26 Things I Learned in the Deep Learning Summer School - Marek Rei","isPartOf":{"@id":"https:\/\/www.marekrei.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#primaryimage"},"image":{"@id":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#primaryimage"},"thumbnailUrl":"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png","datePublished":"2015-08-19T22:07:22+00:00","dateModified":"2019-09-27T23:34:58+00:00","author":{"@id":"https:\/\/www.marekrei.com\/blog\/#\/schema\/person\/a145eb0a06ed4acf5b0f84a24b7a1191"},"description":"Here is a list of small interesting nuggets of information from the summer school, regarding neural networks and deep learning.","breadcrumb":{"@id":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#primaryimage","url":"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png","contentUrl":"http:\/\/www.marekrei.com\/blog\/wp-content\/uploads\/2015\/08\/dlss-3aug2015.png"},{"@type":"BreadcrumbList","@id":"https:\/\/www.marekrei.com\/blog\/26-things-i-learned-in-the-deep-learning-summer-school\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.marekrei.com\/blog\/"},{"@type":"ListItem","position":2,"name":"26 Things I Learned in the Deep Learning Summer School"}]},{"@type":"WebSite","@id":"https:\/\/www.marekrei.com\/blog\/#website","url":"https:\/\/www.marekrei.com\/blog\/","name":"Marek Rei","description":"Thoughts on Machine Learning and Natural Language Processing","potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.marekrei.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Person","@id":"https:\/\/www.marekrei.com\/blog\/#\/schema\/person\/a145eb0a06ed4acf5b0f84a24b7a1191","name":"Marek","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.marekrei.com\/blog\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/48a65414bfda6485aaa0703e548de0ed25292b5fe0d979ed8c28ad83cf5a82c0?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/48a65414bfda6485aaa0703e548de0ed25292b5fe0d979ed8c28ad83cf5a82c0?s=96&d=mm&r=g","caption":"Marek"},"url":"https:\/\/www.marekrei.com\/blog\/author\/marek\/"}]}},"_links":{"self":[{"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/posts\/333","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/comments?post=333"}],"version-history":[{"count":83,"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/posts\/333\/revisions"}],"predecessor-version":[{"id":1303,"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/posts\/333\/revisions\/1303"}],"wp:attachment":[{"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/media?parent=333"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/categories?post=333"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.marekrei.com\/blog\/wp-json\/wp\/v2\/tags?post=333"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}