<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>On The Lambda</title>
  
  <subtitle>another blog by a data scientist</subtitle>
  <link href="https://onthelambda.com/atom.xml" rel="self"/>
  
  <link href="https://onthelambda.com/"/>
  <updated>2022-07-16T23:59:56.556Z</updated>
  <id>https://onthelambda.com/</id>
  
  <author>
    <name>Tony Fischetti</name>
    
  </author>
  
  <generator uri="https://hexo.io/">Hexo</generator>
  
  <entry>
    <title>Data validation with the assertr package</title>
    <link href="https://onthelambda.com/2017/03/20/data-validation-with-the-assertr-package/"/>
    <id>https://onthelambda.com/2017/03/20/data-validation-with-the-assertr-package/</id>
    <published>2017-03-20T14:03:11.000Z</published>
    <updated>2022-07-16T23:59:56.556Z</updated>
    
    <content type="html"><![CDATA[<p><em>Version 2.0 of my data set validation package <code>assertr</code> hit CRAN just this weekend. It has some pretty great improvements over version 1. For those new to the package, what follows is a short and new introduction. For those who are already using <code>assertr</code>, the text below will point out the improvements.</em></p><span id="more"></span><p>I can (and have) go on and on about the treachery of messy&#x2F;bad datasets. Though its substantially less exciting than… pretty much everything else, I believe (proportional to the heartache and stress it causes) we don’t spend enough time talking about it or building solutions around it. No matter how new and fancy your ML algorithm is, it’s success is predicated upon a properly sanitized dataset. If you are using bad data, your approach will fail—either flagrantly (best case), or unnoticeably (considerably more probable and considerably more pernicious).</p><p><code>assertr</code> is a R package to help you identify common dataset errors. More specifically, it helps you easily spell out your assumptions about how the data should look and alert you of any deviation from those assumptions.</p><p>I’ll return to this point later in the post when we have more background, but I want to be up front about the goals of the package; <code>assertr</code> is not (and can never be) a “one-stop shop” for all of your data validation needs. The specific kind of checks individuals or teams have to perform any particular dataset are often far too idiosyncratic to ever be exhaustively addressed by a single package (although, the <code>assertive</code> meta-package may come very close!) But all of these checks will reuse motifs and follow the same patterns. So, instead, I’m trying to sell <code>assertr</code> as a way of thinking about dataset validations—a set of common dataset validation <em>actions</em>. If we think of these actions as <em>verbs</em>, you could say that <code>assertr</code> attempts to impose a grammar of error checking for datasets.</p><p>In my experience, the overwhelming majority of data validation tasks fall into only five different patterns:</p><ul><li><em>For every element in a column</em>, you want to make sure it fits certain criteria. Examples of this strain of error checking would be to make sure every element is a valid credit card number, or fits a certain regex pattern, or represents a date between two other dates. <code>assertr</code> calls this verb <code>assert</code>.</li><li><em>For every element in a column</em>, you want to make sure certain criteria are met <strong>but the criteria can only be decided only <em>after</em> looking at the entire column as a whole</strong>. For example, testing whether each element is within <em>n</em> standard deviations of the mean of the elements requires computation on the elements prior to inform the criteria to check for. <code>assertr</code> calls this verb <code>insist</code>.</li><li><em>For every row of a dataset</em>, you want to make sure certain assumptions hold. Examples include ensuring that no row has more than <em>n</em> number of missing values, or that a group of columns are jointly unique and never duplicated. <code>assertr</code> calls this verb <code>assert_rows</code>.</li><li><em>For every row of a dataset</em>, you want to make sure certain assumptions hold <strong>but the criteria can only be decided only <em>after</em> looking at the entire column as a whole</strong>. This closely mirrors the distinction between <code>assert</code> and <code>insist</code>, but for entire rows (not individual elements). An example of using this would be checking to make sure that the <a href="https://en.wikipedia.org/wiki/Mahalanobis_distance">Mahalanobis distance</a> between each row and all other rows are within <em>n</em> number of standard deviations of the mean distance. <code>assertr</code> calls this verb <code>insist_rows</code>.</li><li>You want to check some property of the dataset as a whole object. Examples include making sure the dataset has more than <em>n</em> columns, making sure the dataset has some specified column names, etc… <code>assertr</code> calls this last verb <code>verify</code>.</li></ul><p>Some of this might sound a little complicated, but I promise this is a worthwhile way to look at dataset validation. Now we can begin with an example of what can be achieved with these verbs. The following example is borrowed from the package vignette and README…</p><p>Pretend that, before finding the average miles per gallon for each number of engine cylinders in the <code>mtcars</code> dataset, we wanted to confirm the following dataset assumptions…*   that it has the columns <code>mpg</code>, <code>vs</code>, and <code>am</code></p><ul><li>that the dataset contains more than 10 observations</li><li>that the column for ‘miles per gallon’ (mpg) is a positive number</li><li>that the column for ‘miles per gallon’ (mpg) does not contain a datum that is outside 4 standard deviations from its mean</li><li>that the <code>am</code> and <code>vs</code> columns (automatic&#x2F;manual and v&#x2F;straight engine, respectively) contain 0s and 1s only</li><li>each row contains at most 2 NAs</li><li>each row is unique jointly between the <code>mpg</code>, <code>am</code>, and <code>wt</code> columns</li><li>each row’s mahalanobis distance is within 10 median absolute deviations of all the distances (for outlier detection)</li></ul><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br></pre></td><td class="code"><pre><span class="line">library<span class="punctuation">(</span>dplyr<span class="punctuation">)</span></span><br><span class="line">library<span class="punctuation">(</span>assertr<span class="punctuation">)</span></span><br><span class="line"></span><br><span class="line">mtcars <span class="operator">%&gt;%</span></span><br><span class="line">  verify<span class="punctuation">(</span>has_all_names<span class="punctuation">(</span><span class="string">&quot;mpg&quot;</span><span class="punctuation">,</span> <span class="string">&quot;vs&quot;</span><span class="punctuation">,</span> <span class="string">&quot;am&quot;</span><span class="punctuation">,</span> <span class="string">&quot;wt&quot;</span><span class="punctuation">)</span><span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  verify<span class="punctuation">(</span>nrow<span class="punctuation">(</span>.<span class="punctuation">)</span> <span class="operator">&gt;</span> <span class="number">10</span><span class="punctuation">)</span> <span class="operator">%&gt;%</span> verify<span class="punctuation">(</span>mpg <span class="operator">&gt;</span> <span class="number">0</span><span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  insist<span class="punctuation">(</span>within_n_sds<span class="punctuation">(</span><span class="number">4</span><span class="punctuation">)</span><span class="punctuation">,</span> mpg<span class="punctuation">)</span> <span class="operator">%&gt;%</span> assert<span class="punctuation">(</span>in_set<span class="punctuation">(</span><span class="number">0</span><span class="punctuation">,</span><span class="number">1</span><span class="punctuation">)</span><span class="punctuation">,</span> am<span class="punctuation">,</span> vs<span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  assert_rows<span class="punctuation">(</span>num_row_NAs<span class="punctuation">,</span> within_bounds<span class="punctuation">(</span><span class="number">0</span><span class="punctuation">,</span><span class="number">2</span><span class="punctuation">)</span><span class="punctuation">,</span> everything<span class="punctuation">(</span><span class="punctuation">)</span><span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  assert_rows<span class="punctuation">(</span>col_concat<span class="punctuation">,</span> is_uniq<span class="punctuation">,</span> mpg<span class="punctuation">,</span> am<span class="punctuation">,</span> wt<span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  insist_rows<span class="punctuation">(</span>maha_dist<span class="punctuation">,</span> within_n_mads<span class="punctuation">(</span><span class="number">10</span><span class="punctuation">)</span><span class="punctuation">,</span> everything<span class="punctuation">(</span><span class="punctuation">)</span><span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  group_by<span class="punctuation">(</span>cyl<span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  summarise<span class="punctuation">(</span>avg.mpg<span class="operator">=</span>mean<span class="punctuation">(</span>mpg<span class="punctuation">)</span><span class="punctuation">)</span> </span><br></pre></td></tr></table></figure><p>Before <code>assertr</code> version 2, the pipeline would immediately terminate at the first failure. Sometimes this is a good thing. However, sometimes we’d like to run a dataset through our entire suite of checks and record all failures. The latest version includes the <code>chain_start</code> and <code>chain_end</code> functions; all assumptions within a chain (below a call to <code>chain_start</code> and above <code>chain_end</code>) will run from beginning to end and accumulate errors along the way. At the end of the chain, a specific action can be taken but the default is to halt execution and display a comprehensive report of what failed including line numbers and the offending datum, where applicable.</p><p>Another major improvement since the last version of <code>assertr</code> of CRAN is that <code>assertr</code> errors are now S3 classes (instead of dumb strings). Additionally, the behavior of each assertion statement on success (no error) and failure can now be flexibly customized. For example, you can now tell <code>assertr</code> to just return TRUE and FALSE instead of returning the data passed in or halting execution, respectively. Alternatively, you can instruct <code>assertr</code> to just give a warning instead of throwing a fatal error. For more information on this, see <code>help(&quot;success_and_error_functions&quot;)</code></p><h3 id="Beyond-these-examples"><a href="#Beyond-these-examples" class="headerlink" title="Beyond these examples"></a>Beyond these examples</h3><p>Since the package was initially published on CRAN (almost exactly two years ago) many people have asked me how they should go about using <code>assertr</code> to test a particular assumption (and I’m very happy to help if you have one of your own, dear reader!) In every single one of these cases, I’ve been able to express it as an incantation using one of these 5 verbs. It also underscored, to me, that creating specialized functions for every need is a pipe dream. There is, however, two good pieces of news.</p><p>The first is that there is another package, <a href="https://cran.r-project.org/package=assertive"><code>assertive</code></a> (vignette here) that greatly enhances the <code>assertr</code> experience. The predicates (functions that start with “is_”) from this (meta)package can be used in <code>assertr</code> pipelines just as easily as <code>assertr</code>’s own predicates. And <code>assertive</code> has an enormous amount of them! Some specialized and exciting examples include <code>is_hex_color</code>, <code>is_ip_address</code>, and <code>is_isbn_code</code>!</p><p>The second is if <code>assertive</code> doesn’t have what you’re looking for, with just a little bit of studying the <code>assertr</code> grammar, you can whip up your own predicates with relative ease. Using some these basic constructs and a little effort, I’m confident that the grammar is expressive enough to completely adapt to your needs.</p><p><em>If this package interests you, I urge you to read the most recent package vignette <a href="https://cran.r-project.org/package=assertr/vignettes/assertr.html">here</a>. If you’re a <code>assertr</code> old-timer, I point you to <a href="https://cran.r-project.org/package=assertr/NEWS">this NEWS file</a> that list the changes from the previous version.</em></p>]]></content>
    
    
    <summary type="html">&lt;p&gt;&lt;em&gt;Version 2.0 of my data set validation package &lt;code&gt;assertr&lt;/code&gt; hit CRAN just this weekend. It has some pretty great improvements over version 1. For those new to the package, what follows is a short and new introduction. For those who are already using &lt;code&gt;assertr&lt;/code&gt;, the text below will point out the improvements.&lt;/em&gt;&lt;/p&gt;</summary>
    
    
    
    <category term="R" scheme="https://onthelambda.com/categories/R/"/>
    
    
    <category term="R" scheme="https://onthelambda.com/tags/R/"/>
    
    <category term="dplyr" scheme="https://onthelambda.com/tags/dplyr/"/>
    
    <category term="magrittr" scheme="https://onthelambda.com/tags/magrittr/"/>
    
    <category term="packages" scheme="https://onthelambda.com/tags/packages/"/>
    
    <category term="CRAN" scheme="https://onthelambda.com/tags/CRAN/"/>
    
    <category term="software" scheme="https://onthelambda.com/tags/software/"/>
    
  </entry>
  
  <entry>
    <title>The Bayesian approach to ridge regression</title>
    <link href="https://onthelambda.com/2016/10/30/the-bayesian-approach-to-ridge-regression/"/>
    <id>https://onthelambda.com/2016/10/30/the-bayesian-approach-to-ridge-regression/</id>
    <published>2016-10-30T20:58:09.000Z</published>
    <updated>2022-07-17T00:01:52.687Z</updated>
    
    <content type="html"><![CDATA[<p>In a TODO <a href="http://www.onthelambda.com/2015/08/19/kickin-it-with-elastic-net-regression/">previous post</a>, we demonstrated that ridge regression (a form of regularized linear regression that attempts to shrink the beta coefficients toward zero) can be super-effective at combating overfitting and lead to a greatly more generalizable model. This approach to regularization used penalized maximum likelihood estimation (for which we used the amazing <code>glmnet</code> package). There is, however, another approach… an equivalent approach… but one that allows us greater flexibility in model construction and lends itself more easily to an intuitive interpretation of the uncertainty of our beta coefficient estimates. I’m speaking, of course, of the bayesian approach.</p><span id="more"></span><p>As it turns out, careful selection of the type and shape of our prior distributions with respect to the coefficients can mimic different types of frequentist linear model regularization. For ridge regression, we use normal priors of varying width.</p><p>Though it can be shown analytically that shifting the width of normal priors on the beta coefficients is equivalent to L2 penalized maximum likelihood estimation, the math is scary and hard to follow. In this post, we are going to be taking a computational approach to demonstrating the equivalence of the bayesian approach and ridge regression.</p><p>This post is going to be a part of a multi-post series investigating other bayesian approaches to linear model regularization including lasso regression facsimiles and hybrid approaches.</p><h3 id="mtcars"><a href="#mtcars" class="headerlink" title="mtcars"></a>mtcars</h3><p>We are going to be using the venerable <code>mtcars</code> dataset for this demonstration because (a) it’s multicollinearity and high number of potential predictors relative to its sample size lends itself fairly well to ridge regression, and (b) we used it in the TODO <a href="http://www.onthelambda.com/2015/08/19/kickin-it-with-elastic-net-regression/">elastic net blog post</a> :)</p><p>Before you lose interest… here! have a figure! An explanation will follow.</p><p><a href="/images/mtcars-loocv-mse-1024x884.png"><img src="/images/mtcars-loocv-mse-1024x884.png" alt="mtcars-loocv-mse"></a></p><p>After scaling the predictor variables to be 0-centered and have a standard deviation of 1, I described a model predicting <code>mpg</code> using all available predictors and placed normal priors on the beta coefficients with a standard deviation for each value from 0.05 to 5 (by 0.025). To fit the model, instead of MCMC estimation via JAGS or Stan, I used quadratic approximation performed by the awesome <code>rethinking</code> package written by Richard McElreath written for his excellent book, <a href="http://xcelab.net/rm/statistical-rethinking/">Statistical Rethinking</a>. Quadratic approximation uses an optimization algorithm to find the maximum a priori (MAP) point of the posterior distribution and approximates the rest of the posterior with a normal distribution about the MAP estimate. I use this method chiefly because as long as it took to run these simulations using quadratic approximation, it would have taken many orders of magnitude longer to use MCMC. Various spot checks confirmed that the quadratic approximation was comparable to the posterior as told by Stan.</p><p>As you can see from the figure, as the prior on the coefficients gets tighter, the model performance (as measured by the leave-one-out cross-validated mean squared error) improves—at least until the priors become too strong to be influenced sufficiently by the evidence. The ribbon about the MSE is the 95% credible interval (using a normal likelihood). I know, I know… it’s pretty damn wide.</p><p>The dashed vertical line is at the prior width that minimizes the LOOCV MSE. The minimum MSE is, for all practical purposes, <em>identical</em> to that of the highest performing ridge regression model using <code>glmnet</code>. This is good.</p><p>Another really fun thing to do with the results is to visualize the movement of the beta coefficient estimates and different penalties. The figure below depicts this. Again, the dashed vertical line is the highest performing prior width.</p><p><a href="/images/mtcars-coef-shrinkage-1024x884.png"><img src="/images/mtcars-coef-shrinkage-1024x884.png" alt="mtcars-coef-shrinkage"></a></p><p>One last thing: we’ve heretofore only demonstrated that the bayesian approach can perform as well as the L2 penalized MLE… but it’s conceivable that it achieves this by finding a completely different coefficient vector. The figure below shows the same figure as above but I overlaid the coefficient estimates (for each predictor) of the top-performing <code>glmnet</code> model. These are shown as the dashed colored horizontal lines.</p><p><a href="/images/mtcars-coef-shrinkage-net-overlay.png"><img src="/images/mtcars-coef-shrinkage-net-overlay.png" alt="mtcars-coef-shrinkage-net-overlay"></a></p><p>These results are pretty exciting! (if you’re the type to not get invited to parties). Notice that, at the highest performing prior width, the coefficients of the bayesian approach and the <code>glmnet</code> approach are virtually identical.</p><p>Sooooo, not only did the bayesian variety produce an equivalently generalizable model (as evinced by equivalent cross-validated MSEs) but also yielded a vector of beta coefficient estimates nearly identical to those estimated by <code>glmnet</code>. This suggests that both the bayesian approach and <code>glmnet</code>‘s approach, using different methods, regularize the model via the same underlying mechanism.</p><p>A drawback of the bayesian approach is that its solution takes many orders of magnitude more time to arrive at. Two advantages of the Bayesian approach are (a) the ability to study the posterior distributions of the coefficient estimates and ease of interpretation that they allows, and (b) the enhanced flexibility in model design and the ease by which you can, for example, swap out likelihood functions or construct more complicated hierarchal models.</p><p>If you are even the least bit interested in this, I urge you to look <a href="https://github.com/tonyfischetti/bayesian-regularization/blob/master/l2-regularization/bayes-reg.R">at the code</a> (in <a href="https://github.com/tonyfischetti/bayesian-regularization">this git repository</a>) because (a) I worked really hard on it and, (b) it demonstrates cool use of meta-programming, parallelization, and progress bars… if I do say so myself :)</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;In a TODO &lt;a href=&quot;http://www.onthelambda.com/2015/08/19/kickin-it-with-elastic-net-regression/&quot;&gt;previous post&lt;/a&gt;, we demonstrated that ridge regression (a form of regularized linear regression that attempts to shrink the beta coefficients toward zero) can be super-effective at combating overfitting and lead to a greatly more generalizable model. This approach to regularization used penalized maximum likelihood estimation (for which we used the amazing &lt;code&gt;glmnet&lt;/code&gt; package). There is, however, another approach… an equivalent approach… but one that allows us greater flexibility in model construction and lends itself more easily to an intuitive interpretation of the uncertainty of our beta coefficient estimates. I’m speaking, of course, of the bayesian approach.&lt;/p&gt;</summary>
    
    
    
    <category term="R" scheme="https://onthelambda.com/categories/R/"/>
    
    <category term="bayesian methods" scheme="https://onthelambda.com/categories/bayesian-methods/"/>
    
    
    <category term="R" scheme="https://onthelambda.com/tags/R/"/>
    
    <category term="research" scheme="https://onthelambda.com/tags/research/"/>
    
    <category term="statistics" scheme="https://onthelambda.com/tags/statistics/"/>
    
    <category term="datavis" scheme="https://onthelambda.com/tags/datavis/"/>
    
    <category term="data mining" scheme="https://onthelambda.com/tags/data-mining/"/>
    
    <category term="machine learning" scheme="https://onthelambda.com/tags/machine-learning/"/>
    
    <category term="bayesian" scheme="https://onthelambda.com/tags/bayesian/"/>
    
    <category term="mcmc" scheme="https://onthelambda.com/tags/mcmc/"/>
    
  </entry>
  
  <entry>
    <title>Using Python decorators to be a lazy programmer: a case study</title>
    <link href="https://onthelambda.com/2016/07/08/using-python-decorators-to-be-a-lazy-programmer-a-case-study/"/>
    <id>https://onthelambda.com/2016/07/08/using-python-decorators-to-be-a-lazy-programmer-a-case-study/</id>
    <published>2016-07-08T19:31:04.000Z</published>
    <updated>2022-07-17T00:04:50.054Z</updated>
    
    <content type="html"><![CDATA[<p>Decorators are considered one of the more advanced features of python and it will often be the last topic in a python class or introductory book. It will, unfortunately, also be one that trips up many beginning or even intermediate python programmers. Those who stick it out and work through it, though, will be <a href="http://i.imgur.com/er0tV3O.jpg">handsomely rewarded</a> for their hard work.</p><span id="more"></span><p>Known by those in-the-know, decorators are tools to make your python code beautiful, more concise, well-written, and elegant–but did you know you can use decorators to be a lazy programmer?!</p><p>I just recently <a href="https://github.com/tonyfischetti/artsy-artwork-dl">whipped up a script</a> that I’m using to help my expand and organize my photo library of my favorite pieces of art. This script that takes the URL of an art piece from <a href="http://www.artsy.net/">artsy.net</a>, downloads the image of the artwork and renames the downloaded file to follow a filename template (given as a CLI arg) based on the artist’s name, the title of the piece, and the date of completion. For example, if we wanted to download an image of the <a href="https://en.wikipedia.org/wiki/Fountain_(Duchamp)">Fountain</a>, and have the file automatically named: <code>Marcel Duchamp - Fountain - 1917.jpg</code>, one can use the following command:</p><figure class="highlight shell"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">python artsy-dl.py &quot;https://www.artsy.net/artwork/marcel-duchamp-fountain-1&quot; &quot;%a - %t - %d&quot;</span><br></pre></td></tr></table></figure><p>because <code>%a</code> is automatically replaced with the artist’s name (that is extracted from the webpage), <code>%t</code> is automatically replaced with the title of the piece, and <code>%d</code> is the date of the piece.</p><p>In this post, I’ll be demonstrating the use of decorators to the end doing the bare minimum and how I saved myself from having to write tedious (but important) error-checking code. But before that, you, dear reader, have to be clear on what a decorator is. The section that follows is perhaps the greatest intro to decorators ever.</p><p>(If you are already familiar with decorators, you can skip to the section called “The problem”–though you may want to at least skim this section.)</p><h3 id="What-are-decorators"><a href="#What-are-decorators" class="headerlink" title="What are decorators?"></a>What are decorators?</h3><p>Here’s the skinny on Python decorators. Grokking decorators necessitates an intuitive understanding of three concepts:</p><ul><li>People often speak of functions as being “first-class citizens” in Python. By this they mean that functions are values that can be assigned to variables, returned from (other) functions, and passed as an argument to (still other) functions.</li><li>When a function (we’ll call it <code>outer</code>) returns another function (we’ll call this <code>inner</code>), the inner function “closes over” (remembers) variables defined in the enclosing scope (<code>outer</code>‘s scope). If the returned function is stored in a variable and called at a later time, it still remembers the variable(s) from the enclosing scope—even if it is called long after the outer function finishes running and it’s variables otherwise lost. The <code>inner</code> function that is returned is known as a “closure”.</li><li>A language that supports closures affords us the unique opportunity to easily add or modify the behavior of a function <code>a</code> by creating a function <code>b</code> that takes function <code>a</code> as an argument, and returns a function <code>c</code> which does something, and then calls function <code>a</code>. Function <code>c</code> can now be used in place of function <code>a</code>–it is essentially function <code>a</code> plus some extra functionality: it is a “decorated” version of function <code>a</code>.</li></ul><p>In order to concretize these concepts, let’s see an example of a decorator complete with an illustration of the motivation behind creating it and the cognitive steps taken toward it’s finished state. Unlike some decorator tutorials, this lesson will not patronize you, dear reader, by designing a overly-simple decorator with no practical worth (it’s always been my thought that this pedagogical strategy most often backfires). Instead, we’ll create, you and I, a decorator of actual utility.</p><p>Suppose we wanted to time the execution of a function. Wanting something with a little more precision than a stopwatch, we decide to use the <code>time</code> module:</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">import</span> time</span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_a_second</span>():</span><br><span class="line">    time.sleep(<span class="number">1</span>)</span><br><span class="line"></span><br><span class="line">start_time = time.time()</span><br><span class="line">sleep_for_a_second()</span><br><span class="line">end_time = time.time()</span><br><span class="line"><span class="built_in">print</span>(<span class="string">&quot;It took &#123;0:.2f&#125; seconds&quot;</span>.<span class="built_in">format</span>(end_time-start_time)) <span class="comment">#&gt; It took 1.00 seconds</span></span><br></pre></td></tr></table></figure><p>This is ok, but if we want to time the execution of many different functions, this will result in a lot of repeated code. Being champions of the <a href="https://en.wikipedia.org/wiki/Don%27t_repeat_yourself">DRY principle</a>, we decide it would be better to put this in a function:</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">time_a_function</span>(<span class="params">func</span>):</span><br><span class="line">    start_time = time.time()</span><br><span class="line">    func()</span><br><span class="line">    end_time = time.time()</span><br><span class="line">    <span class="built_in">print</span>(<span class="string">&quot;It took &#123;0:.2f&#125; seconds&quot;</span>.<span class="built_in">format</span>(end_time-start_time))</span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_a_second</span>():</span><br><span class="line">    time.sleep(<span class="number">1</span>)</span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_two_seconds</span>():</span><br><span class="line">    time.sleep(<span class="number">2</span>)</span><br><span class="line"></span><br><span class="line">time_a_function(sleep_for_a_second) <span class="comment">#&gt; It took 1.00 seconds</span></span><br><span class="line">time_a_function(sleep_for_two_seconds) <span class="comment">#&gt; It took 2.00 seconds</span></span><br></pre></td></tr></table></figure><p><code>time_a_function</code> is a function that takes the function we want to time as an argument.</p><p>We just timed two functions. Notice how we can now time an arbitrary number of functions with no extra code.</p><p>But there’s an issue with this approach. We’ve hitherto been timing functions that take no arguments. How would we time a function that takes one or more arguments?</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_n_seconds</span>(<span class="params">n</span>):</span><br><span class="line">    time.sleep(n)</span><br><span class="line"></span><br><span class="line">time_a_function(sleep_for_n_seconds(<span class="number">5</span>)) <span class="comment">#&gt; TypeError: &#x27;NoneType&#x27; object is not callable</span></span><br></pre></td></tr></table></figure><p>Nope. Before, we were passing the variable that holds the function to <code>time_a_function</code>, but the above incantation evaluates <code>sleep_for_n_seconds(5)</code>, passes it’s <code>None</code> return <em>value</em> to <code>time_a_function</code> and because <code>time_a_function</code> can’t call it (because it’s not a function), we get an error. So how <em>are</em> we going to time <code>sleep_for_n_seconds</code>?</p><p>The solution is to make a function that takes a function that returns a function that takes an argument (n) and performs the function and times it and use the returned function in place of the original (whew!).</p><p>In other words:</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">timer_decoration</span>(<span class="params">func</span>):</span><br><span class="line">  <span class="keyword">def</span> <span class="title function_">new_fn</span>(<span class="params">n</span>):</span><br><span class="line">    start_time = time.time()</span><br><span class="line">    func(n)</span><br><span class="line">    end_time = time.time() <span class="built_in">print</span>(<span class="string">&quot;It took &#123;0:.2f&#125; seconds&quot;</span>.<span class="built_in">format</span>(end_time-start_time))</span><br><span class="line">  <span class="keyword">return</span> new_fn</span><br><span class="line"></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_n_seconds</span>(<span class="params">n</span>):</span><br><span class="line">  time.sleep(n)</span><br><span class="line"></span><br><span class="line">sleep_for_n_seconds = timer_decoration(sleep_for_n_seconds)</span><br><span class="line">sleep_for_n_seconds(<span class="number">3</span>) <span class="comment">#&gt; It took 3.01 seconds</span></span><br></pre></td></tr></table></figure><p>Study this code carefully. You’ve just wrote a decorator. If you are confused it–as always with programming–helps if you type the code out yourself (no copy-and-pasting!) and play around with it.</p><p>Though it’s not terrible unwieldy otherwise, python gives us a nice elegant way to tag a particular function with a decorator so that the function is automatically decorated (i.e. doesn’t require us to replace the original function).</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">@timer_decoration</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_n_seconds</span>(<span class="params">n</span>):</span><br><span class="line">    time.sleep(n)</span><br><span class="line"></span><br><span class="line"><span class="comment"># now the reassignment of &#x27;sleep_for_n_seconds&#x27; is unnecessary</span></span><br><span class="line">sleep_for_n_seconds(<span class="number">3</span>)</span><br><span class="line"><span class="comment">#&gt; It took 3.00 seconds</span></span><br></pre></td></tr></table></figure><p>But what happens if we try to decorate the original <code>sleep_for_a_second</code> function?</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line"><span class="meta">@timer_decoration</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_a_second</span>():</span><br><span class="line">    time.sleep(<span class="number">1</span>)</span><br><span class="line"></span><br><span class="line"><span class="comment"># sleep_for_a_second()</span></span><br><span class="line"><span class="comment">#&gt; TypeError: new_fn() missing 1 required positional argument: &#x27;n&#x27;</span></span><br></pre></td></tr></table></figure><p><code>sleep_for_a_second</code> is now expecting 1 argument :(. We can generalize our decorator to handle functions that take an arbitrary number of arguments with <code>\*args</code> and <code>\*\*kargs</code>…</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br><span class="line">25</span><br><span class="line">26</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">timer_decoration</span>(<span class="params">func</span>):</span><br><span class="line">    <span class="keyword">def</span> <span class="title function_">new_fn</span>(<span class="params">*args, **kargs</span>):</span><br><span class="line">        start_time = time.time()</span><br><span class="line">        func(*args, **kargs)</span><br><span class="line">        end_time = time.time()</span><br><span class="line">        <span class="built_in">print</span>(<span class="string">&quot;It took &#123;0:.2f&#125; seconds&quot;</span>.<span class="built_in">format</span>(end_time-start_time))</span><br><span class="line">    <span class="keyword">return</span> new_fn</span><br><span class="line"></span><br><span class="line"><span class="meta">@timer_decoration</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_n_seconds</span>(<span class="params">n</span>):</span><br><span class="line">    time.sleep(n)</span><br><span class="line"></span><br><span class="line"><span class="meta">@timer_decoration</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_a_second</span>():</span><br><span class="line">    time.sleep(<span class="number">1</span>)</span><br><span class="line"></span><br><span class="line"><span class="meta">@timer_decoration</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_k_seconds</span>(<span class="params">k=<span class="number">1</span></span>):</span><br><span class="line">    time.sleep(k)</span><br><span class="line"></span><br><span class="line">sleep_for_n_seconds(<span class="number">3</span>)</span><br><span class="line"><span class="comment">#&gt; It took 3.00 seconds</span></span><br><span class="line">sleep_for_a_second()</span><br><span class="line"><span class="comment">#&gt; It took 1.00 seconds</span></span><br><span class="line">sleep_for_k_seconds(k=<span class="number">4</span>)</span><br><span class="line"><span class="comment">#&gt; It took 4.00 seconds</span></span><br></pre></td></tr></table></figure><p>Ace!</p><p>Finally, let’s rewrite our decorator to support returning the return value of the decorated function.</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">timer_decoration</span>(<span class="params">func</span>):</span><br><span class="line">    <span class="keyword">def</span> <span class="title function_">new_fn</span>(<span class="params">*args, **kargs</span>):</span><br><span class="line">        start_time = time.time()</span><br><span class="line">        ret_val = func(*args, **kargs)</span><br><span class="line">        end_time = time.time()</span><br><span class="line">        <span class="built_in">print</span>(<span class="string">&quot;It took &#123;0:.2f&#125; seconds&quot;</span>.<span class="built_in">format</span>(end_time-start_time))</span><br><span class="line">        <span class="keyword">return</span> ret_val</span><br><span class="line">    <span class="keyword">return</span> new_fn</span><br><span class="line"></span><br><span class="line"><span class="meta">@timer_decoration</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">sleep_for_a_second_p</span>():</span><br><span class="line">    time.sleep(<span class="number">1</span>)</span><br><span class="line">    <span class="keyword">return</span> <span class="literal">True</span></span><br><span class="line"></span><br><span class="line"><span class="built_in">print</span>(sleep_for_a_second_p())</span><br><span class="line"><span class="comment">#&gt; It took 1.01 seconds</span></span><br><span class="line"><span class="comment">#&gt; True</span></span><br></pre></td></tr></table></figure><p>Note that our decorator is now generalized enough to be used with any function… no matter what it’s return type is… no matter what arguments it takes…</p><p>It doesn’t matter what the function is, the function’s behavior remains the same except now it is “decorated” with functionality that times it.</p><p>This is just one example of a decorator with obvious generalized utility. You can also use decorators to perform <a href="https://wiki.python.org/moin/PythonDecoratorLibrary#Memoize">memoization</a>, <a href="https://wiki.python.org/moin/PythonDecoratorLibrary#Type_Enforcement_.28accepts.2Freturns.29">static-typing-like type enforcement of function signatures</a>, <a href="https://wiki.python.org/moin/PythonDecoratorLibrary#Retry">automatically retry functions that failed</a>, and <a href="https://wiki.python.org/moin/PythonDecoratorLibrary#Lazy_Thunkify">simulate non-strict evaluation</a>.</p><h3 id="The-problem"><a href="#The-problem" class="headerlink" title="The problem"></a>The problem</h3><p>To review, I wrote a script that takes the URL of an artwork on artsy.net, downloads the image, and then names the file in accordance with a user-supplied format string that uses info about the artwork. With the help of the <code>requests</code>, <code>lxml</code>, and <code>wget</code> modules, a script to do this can be coded relatively quickly. The problem, though–which is common for scripts that talk to the web that don’t do error-checking–is that the script is brittle. Without error-checking, any malformed URL, network interruption, invalid output path, or weird edge case like an artwork without a title, will result in an unsightly error message and lengthy stack trace. Besides being aesthetically objectionable, if anyone else is using your script, you will look like an incompetent software engineer. So you have to bite the bullet and error-check.</p><p>The problem with error-checking is</p><ul><li>If all possible errors are checked for separately and individually and handled appropriately (this is good practice), it will result in code often <em>many</em> times longer than the original code. Only a small fraction of the code will be the actual interesting logic of the program–most of it will now be mindless conditionals.</li><li>It’s difficult for someone without training (me) to anticipate every possible error.</li><li>It takes a lot of work and I’m lazy</li></ul><p>So my usual M.O. is to wrap each component in a try&#x2F;except block (with no specificity in the exception), print an error message, and terminate execution…</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">try</span>:</span><br><span class="line">    &lt;brittle code&gt;</span><br><span class="line"><span class="keyword">except</span>:</span><br><span class="line">    sys.exit(<span class="string">&quot;&lt;brittle code&gt; broke&quot;</span>)</span><br></pre></td></tr></table></figure><p>Except I don’t even do that. Instead of wrapping each component in a try&#x2F;except with its own error message, I just wind up try&#x2F;excepting <code>main</code> once. This cuts down typing and carpal tunnel is a real thing…</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">main</span>():</span><br><span class="line">    &lt;literally everything&gt;</span><br><span class="line"></span><br><span class="line"><span class="keyword">try</span>:</span><br><span class="line">    main()</span><br><span class="line"><span class="keyword">except</span>:</span><br><span class="line">    sys.exit(<span class="string">&quot;whoopsie daisy&quot;</span>)</span><br></pre></td></tr></table></figure><h3 id="Decorators-to-the-rescue"><a href="#Decorators-to-the-rescue" class="headerlink" title="Decorators to the rescue"></a>Decorators to the rescue</h3><p>Okkkaaaayyyyy… if we absolutely <em>must</em> do some modicum of error checking around each component (so the user has some kind of clue as to why it the script’s usage failed) we can write a decorator to do this for us. The following is an excerpt of <a href="https://github.com/tonyfischetti/artsy-artwork-dl/blob/d6be495654360553420aae8d7d8e399957b938a2/artsy-dl.py">the script as of commit d6be4956543</a>:</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br><span class="line">13</span><br><span class="line">14</span><br><span class="line">15</span><br><span class="line">16</span><br><span class="line">17</span><br><span class="line">18</span><br><span class="line">19</span><br><span class="line">20</span><br><span class="line">21</span><br><span class="line">22</span><br><span class="line">23</span><br><span class="line">24</span><br><span class="line">25</span><br></pre></td><td class="code"><pre><span class="line">...</span><br><span class="line"><span class="comment"># the decorator</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">cop_out</span>(<span class="params">f</span>):</span><br><span class="line">    <span class="keyword">def</span> <span class="title function_">inner</span>(<span class="params">*args, **kargs</span>):</span><br><span class="line">        <span class="keyword">try</span>:</span><br><span class="line">            <span class="keyword">return</span> f(*args, **kargs)</span><br><span class="line">        <span class="keyword">except</span>:</span><br><span class="line">            sys.exit(<span class="string">&quot;\nThe function &lt;&#123;&#125;&gt; failed\n&quot;</span>.<span class="built_in">format</span>(f.__name__))</span><br><span class="line">    <span class="keyword">return</span> inner</span><br><span class="line"></span><br><span class="line"><span class="meta">@cop_out</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">get_command_line_arguments</span>():</span><br><span class="line">    <span class="keyword">return</span> sys.argv[<span class="number">1</span>], sys.argv[<span class="number">2</span>]</span><br><span class="line"></span><br><span class="line"><span class="meta">@cop_out</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">download_webpage</span>(<span class="params">url</span>):</span><br><span class="line">    r = requests.get(url)</span><br><span class="line">    <span class="keyword">if</span> r.status_code != <span class="number">200</span>:</span><br><span class="line">        <span class="keyword">raise</span> Exception</span><br><span class="line">    <span class="keyword">return</span> r</span><br><span class="line"></span><br><span class="line"><span class="meta">@cop_out</span></span><br><span class="line"><span class="keyword">def</span> <span class="title function_">parse_webpage</span>(<span class="params">requests_object</span>):</span><br><span class="line">    <span class="keyword">return</span> lxml.html.fromstring(requests_object.text)</span><br><span class="line">....</span><br></pre></td></tr></table></figure><p>Note how we use <code>f.__name__</code> to get the name of the function that was decorated. This allows us to add the support for specialized (at the function level, at least) error messages for free!</p><p>Now, if the user calls the script with too few arguments, the program will print <code>The function &lt;get_command_line_arguments&gt; failed</code>. If you give it a real URL but not to an artsy.net artwork, it’ll say <code>The function &lt;extract_artist_name_from_webpage&gt; failed</code>. If you give it a made-up URL, it’ll say <code>The function &lt;download_webpage&gt; failed</code>, etc…</p><p>Sure, beyond the function level, you don’t know <em>why</em> it failed, but anything is better than nothing and your users shouldn’t be so bossy and entitled.</p><p>But one more thing… if you looked at the code, you’ll notice that my function names are descriptive… maybe <em>too</em> long and descriptive. The use of prose-like descriptive function names (certainly by the standards of <a href="http://hackage.haskell.org/package/base-4.9.0.0/docs/src/Data.Foldable.html#mapM_">Haskell programmers</a>) was no accident. Although it may <em>seem</em> like an uncharacteristically diligent and conscientious decision on my part, it was actually to facilitate further laziness. Consider the following tweak to the decorator:</p><figure class="highlight python"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br></pre></td><td class="code"><pre><span class="line"><span class="keyword">def</span> <span class="title function_">cop_out</span>(<span class="params">f</span>):</span><br><span class="line">    <span class="keyword">def</span> <span class="title function_">inner</span>(<span class="params">*args, **kargs</span>):</span><br><span class="line">        <span class="keyword">try</span>:</span><br><span class="line">            <span class="keyword">return</span> f(*args, **kargs)</span><br><span class="line">        <span class="keyword">except</span>:</span><br><span class="line">            message = f.__name__.replace(<span class="string">&quot;_&quot;</span>, <span class="string">&quot; &quot;</span>)</span><br><span class="line">            sys.exit(<span class="string">&quot;\nFailed to &#123;&#125;\n&quot;</span>.<span class="built_in">format</span>(message))</span><br><span class="line">    <span class="keyword">return</span> inner</span><br></pre></td></tr></table></figure><p>Consider how this generates error messages that appear to individualized…</p><figure class="highlight plaintext"><figcaption><span>4on</span></figcaption><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br></pre></td><td class="code"><pre><span class="line">$ python artsy-dl.py &quot;https://www.artsy.net/artwork/jean-michel-basquiat-untitled-33211237&quot;</span><br><span class="line">Failed to get command line arguments</span><br><span class="line"></span><br><span class="line">$ python3.4 artsy-dl.py &quot;https://www.artsy.net/artwork/jean-michel-BAHHSKIIIAAHT-untitled-33211237&quot;</span><br><span class="line">                        &quot;%a/%t - %d&quot;</span><br><span class="line">Failed to download webpage</span><br><span class="line"></span><br><span class="line">$ python2.7 artsy-dl.py &quot;https://www.artsy.net/artist/jean-michel-basquiat&quot;</span><br><span class="line">                         &quot;%a/%t - %d&quot;</span><br><span class="line">Failed to extract artist name from webpage</span><br></pre></td></tr></table></figure><p>So there you have it! Decorators can be used for legitimate, elegant solutions but can also be employed–virtually for free–to give the illusion that you are a caring software engineer and meticulous with your error checking.</p><h4 id="PS"><a href="#PS" class="headerlink" title="PS"></a>PS</h4><p>If you’re a potential employer or client, I’m just kidding–I’m very diligent about error checking. This piece is satire. I promise.</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;Decorators are considered one of the more advanced features of python and it will often be the last topic in a python class or introductory book. It will, unfortunately, also be one that trips up many beginning or even intermediate python programmers. Those who stick it out and work through it, though, will be &lt;a href=&quot;http://i.imgur.com/er0tV3O.jpg&quot;&gt;handsomely rewarded&lt;/a&gt; for their hard work.&lt;/p&gt;</summary>
    
    
    
    <category term="python" scheme="https://onthelambda.com/categories/python/"/>
    
    
    <category term="python" scheme="https://onthelambda.com/tags/python/"/>
    
    <category term="software" scheme="https://onthelambda.com/tags/software/"/>
    
    <category term="art" scheme="https://onthelambda.com/tags/art/"/>
    
    <category term="functional programming" scheme="https://onthelambda.com/tags/functional-programming/"/>
    
  </entry>
  
  <entry>
    <title>Computational foreign language learning: a study in Spanish verbs usage</title>
    <link href="https://onthelambda.com/2016/06/30/computational-foreign-language-learning-a-study-in-spanish-verbs-usage/"/>
    <id>https://onthelambda.com/2016/06/30/computational-foreign-language-learning-a-study-in-spanish-verbs-usage/</id>
    <published>2016-06-30T20:59:55.000Z</published>
    <updated>2022-07-22T17:08:43.805Z</updated>
    
    <content type="html"><![CDATA[<p><em><strong>Abstract</strong>: I did some computer-y stuff to construct a personal Spanish text corpus and create a Spanish verb study guide specifically tailored to the <a href="https://en.wikipedia.org/wiki/Variety_(linguistics)">linguistic variety</a> of Spanish I intend to consume and produce. It worked fairly well. It also revealed a (in some small way) generalizable depiction of the relative frequencies of Spanish verb tenses and moods. This technique may prove to be extremely beneficial to Spanish-language pedagogy. If you’re uninterested in my motivations or procedure, you can skip to the section labeled “results”.</em></p><span id="more"></span><p>As regular readers of this blog may be aware, one of my favorite activities is marshaling the skills that I use as a computational scientist to study the humanities. For example, in a <a href="http://www.onthelambda.com/2013/10/29/linguistics-meet-evolutionary-biology/">previous post</a>, we saw how principles from phylogenetic systematics helped textual critics reconstruct the original manuscript for “The Canterbury Tales”; <a href="http://www.onthelambda.com/2014/02/20/how-to-fake-a-sophisticated-knowledge-of-wine-with-markov-chains/">in another</a>, we deployed techniques first used to study physics to the end of fooling vineyards into retweeting fake, computer-generated wine reviews.</p><p>For this post, I used both tools from computational linguistics and some good-old-fashioned data wrangling (web-scraping, parsing texts, etc…) to create a custom-fit Spanish verb study guide.</p><h3 id="The-problems"><a href="#The-problems" class="headerlink" title="The problems"></a>The problems</h3><h4 id="Problem-1"><a href="#Problem-1" class="headerlink" title="Problem #1"></a>Problem #1</h4><p>Although foreign language immersion is the almost certainly the best learning path for most types of foreign language learners, no reasonable student without an lavish budget for traveling can expect to get by without having to do some rote memorization. In the context of Spanish verbs, this either means unguided memorization of a dictionary or consultation of a list of the most commonly used Spanish verbs. But, even if you could trust that the most-popular-verbs list was compiled in a principled manner, there are vast regional and sub-culture-specific variations in verb frequency. For example, the verb <em>coger</em> means “to take” in Spain but in Central America it’s… it’s a… pretty vulgar verb. It stands to reason that there are pretty enormous differences in this verb’s popularity across regions, contexts, and <a href="https://en.wikipedia.org/wiki/Register_(sociolinguistics)">registers</a>. Depending on which region’s dialect you prioritize familiarity with, and depending on how raggle-taggle the people you intend to roll with are—or the media you intend to consume—a one-size-fits-all verb list might let you down.</p><h4 id="Problem-2"><a href="#Problem-2" class="headerlink" title="Problem #2"></a>Problem #2</h4><p>English isn’t a very <a href="https://en.wikipedia.org/wiki/Inflection">inflective</a> language—the tense (or person, mood, aspect, etc…) is largely determined, not through verb conjugation, but via <a href="https://en.wikipedia.org/wiki/Periphrasis">periphrasis</a>, the use of personal pronouns, and other auxiliary words. This is in stark comparison to Spanish, a highly-inflective, relatively <a href="https://en.wikipedia.org/wiki/Synthetic_language">synthetic language</a> where the verb’s conjugation betrays its tense, person, mood, and aspect—all in one word! This linguistic elegance is a learning obstacle, since one verb might be written in a little under 60 different ways (6 persons * (4 tenses in the indicative mood + 3 tenses in the subjunctive mood + 1 imperative mood)). This pedagogical nightmare is partially allayed by careful prioritization of some tenses and moods, over others—at least initially. For example, a Spanish-language learner almost always learns the commonly-used and versatile present indicative tense first. But beyond the next few obvious choices, the order in which these tenses should be prioritized is not clear and (probably) dependent on how and where you expect to use and consume the language. Further complicating things, there are entire persons (here’s looking to you, <em>vosotros</em>) that are very uncommon in most Spanish-speaking countries.</p><h3 id="The-solution"><a href="#The-solution" class="headerlink" title="The solution"></a>The solution</h3><p>The solution to this problem is to create a personal corpus of Spanish text, containing examples of the types of text you expect to consume and produce. Then, the verbs need to be identified, have their mood, tense, and person recorded, and converted into infinitive form (for frequency tabulation). The relative frequencies of the persons, mood, and tenses—as well as the frequencies of the verbs (in infinitive form)—will inform the creation of a Spanish verb study guide specifically catered to type of linguistic variety the learner intends to employ. Whether the learner’s primary interest in learning Spanish is to be able to bond with a new family member over their love of Mexican <em>telenovelas</em> or to read and understand <em>Don Quixote</em> in its entirety, this approach will hasten the learner’s sense of accomplishment with respect to cookie-cutter verb study guides, increase learner satisfaction, and increase the likelihood of the learner actually achieving language mastery. I mean, as a learner myself, I would be discouraged if I felt like the main payoff of studying Spanish is to read and understand books that are very obviously juvenile or primary meant for pedagogical purposes. I want to read Márquez and I want to read him <em>now</em>!</p><h3 id="The-corpus"><a href="#The-corpus" class="headerlink" title="The corpus"></a>The corpus</h3><p>For my particular corpus, I chose a whole mess of books (most of which I’ve read—and loved—in English) that I’m interested in reading in the original language. These include <em>Rayuelas</em> and <em>Final De Juego</em> by Julio Cortázar (my favorite short story writer), <em>Cien Años De Soledad</em> by Gabriel García Márquez (generally considered to be a masterpiece), <em>Darios de Motocicleta</em> by Che Guevara, <em>Ficciones</em> by Jorge Luis Borges, and <em>La Cuidad De Las Bestias</em> by Isabel Allende. These texts were obtained electronically—legitimately!—and I used various ad-hoc regexes to remove formatting and conversion-from-PDF-to-text) artifacts.</p><p>My interest in Spanish isn’t only for consuming literature, though; I wanted to include other sources of text, like movie scripts (I planned on <em>Lo Que le Pasó a Santiago</em>, generally considered to be one of the best Puerto Rican films), but I couldn’t find the script online. I also wanted to include the lyrics to my favorite Spanish-language bands (Soda Stereo, El Ultimo Vecíno, Décima Víctima, Caifenes, Shakira, Millie Quezada, …) but the tool I used to identify the verbs in the corpus often choked on these texts. Why, you ask?…</p><h3 id="Parts-of-speech-tagging"><a href="#Parts-of-speech-tagging" class="headerlink" title="Parts-of-speech tagging"></a>Parts-of-speech tagging</h3><p><em>references are at the bottom of the post</em></p><p>Parts-of-speech tagging (hereafter, ‘POS tagging’) is when you go through a text and, for each word, identify the which part of speech (verb, noun, adjective, etc…) the word functions as.</p><p>This is a non-trivial task because the same word can function as different parts-of-speech depending on the context. Take the following sentence, for example, which is an expanded and modified version of a sentence that is used as an example in <a href="https://class.coursera.org/nlp/lecture/149">this video</a></p><p><em>Fruit flies like bananas</em></p><p>So, taken individually, <em>all</em> words in this sentence can function as multiple parts of speech. Take “like” for instance; it can be a noun (“<em>my status got mad likes</em>“), a verb (“<em>I like your status</em>“), a quotative (“<em>I was like, ‘I enjoyed your status</em>‘“), conjunction (_“I updated my status like the world depended on it_”), a preposition (“<em>I wrote my status like Nathaniel Hawthorne</em>“). Depending on how colloquial the text in question is, “like” can even be used as a discourse marker (“<em>I’m, like, scared of ghosts, Scoob</em>“). As a standalone word, “like” can serve the purpose of 6 different parts of speech.</p><p>But even looking at the entire sentence as a whole, the parts-of-speech for each word is ambiguous.</p><p>Concretely, the sentence can be interpreted as (a) “<em>fruit flies</em> (noun) <em>like</em> (verb) <em>bananas</em> (noun)”, (b) “<em>fruit</em> (noun) <em>flies</em> (verb) <em>like</em> (preposition) <em>bananas</em> (noun) [do]”, or even (c) “<em>fruit</em> (noun) <em>flies</em> (verb) <em>like</em> (conjunction(?)) <em>bananas</em> (adjective)”—using the colloquial meaning of the word <em>bananas</em> meaning <em>“crazy”</em>.</p><p>Note that the POS tag for one word is conditional on the POS tags of other words: whether flies is a noun or a verb affects whether bananas is interpretable as a adjective.</p><p>Because this task isn’t easy, this job used to be left to humans to perform. Now, various techniques allow for this to be done programmatically to a high degree of accuracy. We’ll go through a few of them, ending with the sophisticated method employed by the POS tagger that we will be using, the <a href="http://nlp.stanford.edu/software/tagger.shtml">Stanford Parts-of-speech tagger</a>.</p><h4 id="Unigram-tagging"><a href="#Unigram-tagging" class="headerlink" title="Unigram tagging"></a>Unigram tagging</h4><p>A training corpus with the POS tags for each word is read and, for each unique word, the number of times it is used as one of the various parts of speech is tallied. When a word is encountered in untagged text, the tagger chooses the part-of-speech that the word is most commonly used as in the training text. If the word encountered was not in the training text at all, it defaults to a noun. Somehow, this context-free elementary method can yield accuracies of 90%-94% <a href="http://www.aclweb.org/anthology/P98-1029">(Brill &amp; Wu, 1998)</a>. When Brill and Wu used this method with&#x2F;on the famous <a href="https://catalog.ldc.upenn.edu/LDC99T42">Penn Treebank Wall Street Journal corpus</a> with a 80%&#x2F;20% training&#x2F;testing split, it achieved 93.3% accuracy.</p><h4 id="n-gram-tagging"><a href="#n-gram-tagging" class="headerlink" title="n-gram tagging"></a>n-gram tagging</h4><p>Using an <em>n</em>-gram model, the tag of a particular word is assumed to be conditionally dependent on the tag of the preceding <em>n-1</em> words. For example, in a bigram model, the tag of the current word is guessed from the current word, <em>and</em> the tag of the previous word. A trigram model uses tag information from the previous <em>two</em> words, in concert with the conditional probability of a particular tag <em>given</em> a certain word. The unigram tagger is a special case of the _n_-gram tagger where <em>n</em> is 1. It’s not hard to see that _n_-gram tagging will offer an enormous accuracy improvement.</p><p>If this reminds you of the Markov chains that we made use of in the previous post on <a href="http://www.onthelambda.com/2014/02/20/how-to-fake-a-sophisticated-knowledge-of-wine-with-markov-chains/">computer-generating wine reviews</a>, then you have a good eye. N-gram tagging is a type of <a href="https://en.wikipedia.org/wiki/Hidden_Markov_model">Hidden Markov Model (HMM)</a>. What makes HMMs different than simple Markov models is that the states themselves (the POS tags) are not directly observable; the observable portion of each state are the actual words—and the words are only a probabilistic function of the state.</p><p>In addition to testing a unigram model, Brill and Wu also tested this technique’s ability on the WSJ corpus. In particular, they used a trigram tagger—with a twist. <a href="https://www.aclweb.org/anthology/J/J93/J93-2006.pdf">Weischedel, Ralph, et al (1993)</a> noted that the suffix of a word (<em>-ed</em>, <em>-s</em>, <em>-ing</em>, <em>-ion</em>, <em>-ly</em>, etc…) strongly influenced the probability that the word served as a particular part of speech. When this information was wielded to help classify unknown words, it greatly improved accuracy outcomes. When Brill and Wu used this method with a trigram tagger against the WSJ corpus, the technique yielded an 96.4% accuracy rate.</p><h4 id="Maximum-Entropy-models"><a href="#Maximum-Entropy-models" class="headerlink" title="Maximum Entropy models"></a>Maximum Entropy models</h4><p>Maximum Entropy models are a lot like—insofar as they are equivalent to—multinomial logistic regression models that attempt to model the probability of a given tag class given various predictor variables, or features. Maximum entropy models can use features such as the current word, the previous word, the previous word’s tag, etc…—like would a HMM—but also features like whether the word contains a number, whether the word is capitalized, etc… An optimization algorithm called Generalized Iterative Scaling selects the feature weights that maximize the likelihood function.</p><p><a href="http://www.aclweb.org/anthology/W96-0213">Ratnaparkhi (1996)</a> tested a straightforward maximum entropy model on the WSJ corpus and noted that it yielded an accuracy of 96.6%. Four years after that, <a href="http://ilpubs.stanford.edu:8090/459/1/2000-39.pdf">Toutanova et al. (2000)</a> published a paper in which they show that by adding additional features like whether the word is capitalized <em>and</em> in the middle of a sentence and non-local features that look 8 words back for a <a href="https://en.wikipedia.org/wiki/Modal_verb">modal verb</a> (for disambiguating base form verbs and non-3rd person singular present verbs) they can achieve a WSJ accuracy of 96.8%. This is the benefit of the Maximum Entropy model approach—you can arbitrarily add features (within reason) without necessarily knowing <em>how</em> those features contribute the the probabilities of tag outputs.</p><p>Three years after that, <a href="http://nlp.stanford.edu/pubs/tagging.pdf">Toutanova et al. (2003)</a> achieved a 97.2% accuracy rate on the WSJ corpus by (a) adding features for the words <em>following</em> the word currently being tagged, and (b) using regularization to combat overfitting as a result of using <em>many</em> features—many of which probably only weakly contribute information of the probability of the current word’s tag class. Their regularization technique involved placing a zero-centered Gaussian prior on the feature weights and is mathematically tantamount to the L2 regularization that we saw in <a href="http://www.onthelambda.com/2015/08/19/kickin-it-with-elastic-net-regression/">this previous blog post.</a> This state-of-the-art tagger is the one on which the Stanford tagger we use is based.</p><p><em>There is another famous type of POS tagger called Transformation-Based tagger. In contrast to all the others that were mentioned above, this is not a probabilistic&#x2F;stochastic model and is, instead, based on rules and knowledge. I won’t describe it here because it’s very different and this post is already too long but I should mention that it can score a 96.6% on the on WSJ corpus <a href="http://www.aclweb.org/anthology/P98-1029">(Brill et al., 1998)</a>.</em></p><h3 id="The-procedure"><a href="#The-procedure" class="headerlink" title="The procedure"></a>The procedure</h3><p><em>These steps assume a POSIX compliant system and some command-line proficiency</em></p><p><em>The filenames are links and you can find a <a href="https://github.com/tonyfischetti/spanish-verb-research">repo with all the code here</a></em></p><ul><li>Downloaded full version of the <a href="http://nlp.stanford.edu/software/tagger.shtml">Stanford Parts-of-speech tagger</a></li><li>Ran the tagger on the text, put each tag on a separate line, and filtered for verbs only. The parts-of-speech were identified using <a href="http://nlp.stanford.edu/software/spanish-faq.shtml#tagset">this tagset</a>. As you can see, the verbs all start with the letter “v”. This can be achieved by the following incantation:</li></ul><figure class="highlight bash"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">./stanford-postagger.sh models/spanish.tagger THE_BOOK.txt perl -pe <span class="string">&#x27;s/ /\\n/g&#x27;</span> grep <span class="string">&#x27;_v&#x27;</span> &gt; tmp</span><br></pre></td></tr></table></figure><p>If this causes you problems, you might want to try to give the tagger (which runs in multicore!) more memory; try adding <code>-Xmx2048M</code> as a argument in the <code>java</code> command in <code>./stanford-postagger.sh</code>—this will give it 2GBs to work with.</p><ul><li>For each work, I ran this.py on it, which parsed the stanford tag and made it in nice tab delimited format:</li></ul><figure class="highlight bash"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">./stanford-output-to-nice-tsv.py &lt; tmp &gt; ./output-verbs/THE_BOOK.txt</span><br></pre></td></tr></table></figure><ul><li>Catted all of them together into <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/all.txt">all.txt</a>–a monstrous text file with 84,437 words that the tagger interpreted as verbs:</li></ul><figure class="highlight bash"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line"><span class="built_in">cat</span> rayuelas.txt final-de-juego.txt darios-de-motocicleta.txt cien-anos-de-soledad.txt ficciones.txt la-cuidad-de-las-bestias.txt &gt; all.txt</span><br></pre></td></tr></table></figure><p>Now we need to get the infinitives, but in order to prioritize which we should get the infinitives for, and not have to repeat conjugated verbs, we need to get the uniques…</p><ul><li>So I ran</li></ul><figure class="highlight bash"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line"><span class="built_in">cat</span> all.txt perl -pe <span class="string">&#x27;s/(.+?)\\t.\*/\\1/g&#x27;</span> &gt; all-verbs.txt</span><br></pre></td></tr></table></figure><p> to get a list of only verbs (no mood or tense)</p><ul><li>I wanted to get a list of unique verbs sorted by the number of occurrences; this would normally be a job for the <code>sort uniq -c</code>. <em>Desafortunademente</em>, this command fails. It turns out that unicode can represent (for example) <em>habría</em> in <a href="https://en.wikipedia.org/wiki/Precomposed_character">at least two different ways</a>. For this reason, we have to use the python script <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/process-all-verbs.py">process-all-verbs.py</a> which uses the <code>unicodedata</code> module to normalize the verbs and <em>then</em> count them.</li></ul><figure class="highlight bash"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">./process-all-verbs.py <span class="built_in">tee</span> all-verbs-count.txt</span><br></pre></td></tr></table></figure><p>Ok, <em>now</em> were ready to get infinitive forms for these verbs. We are going to do this by programmatically making request to translate the word to the (excellent) website <a href="http://www.spanishdict.com/">Span¡shD!ct.com</a>. What we want can be extracted from the returned HTML via CSS selectors.</p><ul><li><a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/get-infinitives.py">get-infinitives.py</a> goes through each line of <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/all-verbs-count.txt">all-verbs-count.txt</a> and constructs the url to query the website with. It then uses the CSS selector “.mismatch” for information about the verb.</li></ul><p>In the best case scenario, it says something like “___“ is the ____ form of _____ in the ____“. Sometimes, there’s more than one possible person or tense so it says “____ represents different conjugations of the verb _____“</p><p>In either case, we get the infinitive. If it fails, we record it and move on. It waits between 1 and 2 seconds between each verb. After every 20, it dumps the JSON so that in case something bad happens I could just load the intermediate results and restart.</p><ul><li>You can see that the SpanishDict infinitive conversion systematically failed for certain words. For example, it interpreted inflected verbs like <em>he</em>, <em>dice</em>, and <em>era</em> as English words to translate, not Spanish words to provide information for. In other cases, it interpreted a verb’s past participle (<em>aburrir</em> -&gt; <em>aburrido</em> (“to bore”)) as an adjective (“boring”). I manually filled in many of the ones that failed using equal parts regex and black magic. This went into <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/finished-supplemented.json">finished-supplemented.json</a>.</li><li>Finally, we need to inner join <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/all.txt">all.txt</a><code>to the information in</code><a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/finished-supplemented.json">finished-supplemented.json</a>. The <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/combine.py">combine.py</a> script does this:</li></ul><figure class="highlight bash"><table><tr><td class="gutter"><pre><span class="line">1</span><br></pre></td><td class="code"><pre><span class="line">./combine.py <span class="built_in">tee</span> tagged-plus-infinitives.txt [/code]</span><br></pre></td></tr></table></figure><p>The tab-delimited <a href="https://raw.githubusercontent.com/tonyfischetti/spanish-verb-research/master/tagged-plus-infinitives.txt">tagged-plus-infinitives.txt</a> in now ready to be consumed for analysis.</p><h3 id="Some-numbers"><a href="#Some-numbers" class="headerlink" title="Some numbers"></a>Some numbers</h3><ul><li><em>Rayuelas</em> - 203,197 words - 29,882 verbs <em>Final de juego</em> - 54,303 words - 8,160 verbs <em>Darios de Motocicleta</em> - 53,804 words - 6,557 verbs <em>Cien Años de Soledad</em> - 15,4381 words - 20,987 verbs <em>Ficciones</em> - 48,845 words - 5,769 verbs <em>La Cuidad De Las Bestias</em> - 94,075 words - 13,082 verbs</li><li>There were 84,437 words that the tagger identified as verbs in all.</li><li>There were 13,972 unique conjugated verbs.</li><li>After the first try with SpanishDict, for only 6,852 verbs did we have the infinitives. This greatly increased with the black magic alluded to in the previous section.</li><li>I went from 84,437 to 71,378 verbs when I inner joined with the verbs that I was able to find infinitives for.</li></ul><h3 id="The-results"><a href="#The-results" class="headerlink" title="The results"></a>The results</h3><p><a href="/images/SpanishVerbPlot.png"><img src="/images/SpanishVerbPlot.png" alt="Figure 1: Proportion of Spanish verb moods and tenses in corpus"></a></p><p class="tony-caption">Proportion of Spanish verb moods and tenses in corpus</p><p>The results were rather fascinating.</p><p>These were the 14 most common <em>conjugated</em> verbs:</p><table>  <tr>    <th>conjugated verb</th>    <th>count</th>    <th>percent</th>  </tr>  <tr>    <td>había</td> <td>2599</td> <td>3.64</td>  </tr>  <tr>    <td>era</td> <td>2396</td><td>3.36</td>  </tr>  <tr>    <td>es</td> <td>2303</td><td>3.23</td>  </tr>  <tr>    <td>dijo</td> <td>1763</td><td>2.47</td>  </tr>  <tr>    <td>estaba</td> <td>1169</td><td>1.64</td>  </tr>  <tr>    <td>fue</td> <td>816</td><td>1.14</td>  </tr>  <tr>    <td>ser</td> <td>606</td><td>0.85</td>  </tr>  <tr>    <td>habían</td> <td>517</td><td>0.72</td>  </tr>  <tr>    <td>hay</td> <td>512</td><td>0.72</td>  </tr>  <tr>    <td>tenía</td> <td>467</td><td>0.65</td>  </tr></table><p>(you can see the <a href="https://github.com/tonyfischetti/spanish-verb-research/blob/master/most-common-conjugated-verbs.csv">full spreadsheet here</a>)</p><p>With this information alone, this whole endeavor was worth it. Sure, most of the verbs in this list aren’t that much of a surprise, but there are two pieces of information that could prove really helpful to me. The first is that 4 verbs in the top 15 are forms of the verb <em>haber</em> (“to have”)—including the very first one, which accounts for 3.6% of <em>all conjugated verbs in the corpus</em>. This is a verb that I was, heretofore, relatively unfamiliar with.</p><p>In contrast to <em>tener</em> (which also means “to have”), <em>haber</em> is often used as an auxiliary verb as it would in such english sentences as “I <strong>have</strong> to go to the dentist”, “I <strong>had</strong> all but lost it” (past perfect tense), “<strong>there is</strong> a freeze-up coming”. Because of it’s ubiquitous usage as an auxiliary word (like its being used in all sentences in the <a href="https://en.wikipedia.org/wiki/Perfect_(grammar)">perfect</a> mood), I should get more familiar with this verb and its conjugations if I ever hope to read these works of literature.</p><p>The second important piece of information for me was that a <em>majority</em> of the verbs in the top 14 were in the imperfect tense (a type of past tense). Now, I think I may have been concentrating too much on the preterite tense (<em>another</em> past tense) in comparison. Next, these were the 14 most common verbs when put into infinitive form:</p><table>  <tr>    <th>infinitive</th>    <th>count</th>    <th>perc</th>  </tr>  <tr>    <td>ser</td>    <td>8066</td>    <td>11.3</td>  </tr>  <tr>    <td>haber</td>    <td>5461</td>    <td>7.65</td>  </tr>  <tr>    <td>estar</td>    <td>2746</td>    <td>3.85</td>  </tr>  <tr>    <td>decir</td>    <td>2734</td>    <td>3.83</td>  </tr>  <tr>    <td>tener</td>    <td>1774</td>    <td>2.49</td>  </tr>  <tr>    <td>hacer</td>    <td>1757</td>    <td>2.46</td>  </tr>  <tr>    <td>ir</td>    <td>1721</td>    <td>2.41</td>  </tr>  <tr>    <td>poder</td>    <td>1614</td>    <td>2.26</td>  </tr>  <tr>    <td>ver</td>    <td>1336</td>    <td>1.87</td>  </tr>  <tr>    <td>dar</td>    <td>1210</td>    <td>1.7</td>  </tr></table><p>(you can see the <a href="https://github.com/tonyfischetti/spanish-verb-research/blob/master/most-common-infinitives.csv">full spreadsheet here</a>)</p><p>To me, there wasn’t really anything unexpected here except for maybe <em>pasar</em> (to happen) and <em>parecer</em> (to seem), which I was, up until this point—relatively unfamiliar with in spite of the fact that they are used in a number of frequently spoken expressions like <em>¿Que pasó?</em> (“What happened?”) and <em>¿Que te parece?</em> (~”What do you think?”).</p><p>Finally, <a href="http://www.onthelambda.com/wp-content/uploads/2016/06/SpanishVerbPlot.png">figure 1</a> is a plot which depicts the proportions in which each mood and tense occur. The large vertical bars show the relative proportions of each mood (I count the Infinitive, Gerund, and Participle as moods) in descending order; they are Indicative (65%), Infinitive (20%), Subjunctive (4%), Participle (4%), Gerund (3%), and Imperative (1%). Each vertical bar is further broken down by the proportion of each tense <em>within that mood</em> (sorted, with the most frequently used on the bottom. For example, the present tense is the most common tense in the indicative mood and accounts for 26% of all mood&#x2F;tense pairs. The Infinitive, Participle, and Imperative moods (to the extent that there are actually moods) have only one tense (to the extent that they can be said to have tenses).</p><p>These results were most surprising to me; for one, I was (again) reminded that I should probably hold nailing down the imperfect tense with as much or more importance as I do with the preterite tense. Second, I was surprised that usage of the future tense was far eclipsed by gerund, participle, and both subjective tenses—in spite of the fact that I use it quite often in my texts to my friends and my internal monologue. Of course, this—and other insights—may just be artifacts of the particular body of literature I chose for my corpus (see next section).</p><p><strong>Limitations:</strong></p><p>Although this was a wildly fun project that yielded interesting and extremely practical insights, there are a number of important caveats to be aware of when interpreting these results. </p><p>First is a generalizability issue; the results indicate the verb popularity and mood&#x2F;tense breakdowns for <em>just 6 pieces of Spanish literature</em>. Because of this, the corpus is heavily dominated by the writing style of the included authors—at least some of whom have a very idiosyncratic writing style. Additionally, as with most literature, all of the non-short-stories in my corpus were told in the past tense (usually by a third person omniscient narrator). This past tense bias is very clearly non-representative of everyday spoken Spanish (of course, it was never meant to be representative of that). This problem could have been, at least partially, alleviated via the inclusion of more prosaic Spanish from movie scripts and blogs—if only they POS tagged correctly!!</p><p>Speaking of tagging correctly, the second issue is one of the correctness of the POS tags. The best POS taggers (Stanford is certainly one) can, at best, achieve an accuracy of 97%. Although this is an incredible feat of computational linguistics and the product of many many years of research, it is important to put this in the proper perspective. Recall that the rudimentary unigram tagger can achieve a 90%-94% accuracy rate (b) the 97% accuracy rate decreases as the testing corpus diverges in style from the training corpus. Especially because of Cortázar—who (at least in English translations) employs highly unusual sentence structure and often straight-up grammatically-incorrect non-human-parsable sentences—this fact must be kept in mind; unless the Spanish model that comes with Stanford was trained with Surrealist literature (it wasn’t!), tag accuracy will suffer.</p><h4 id="References"><a href="#References" class="headerlink" title="References"></a>References</h4><p><a href="http://www.aclweb.org/anthology/P98-1029">Brill, Eric, and Jun Wu. “Classifier combination for improved lexical disambiguation.” Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics-Volume 1. Association for Computational Linguistics, 1998.</a></p><p><a href="http://www.aclweb.org/anthology/W96-0213">Ratnaparkhi, Adwait. “A maximum entropy model for part-of-speech tagging.” Proceedings of the conference on empirical methods in natural language processing. Vol. 1. 1996.</a></p><p><a href="http://ilpubs.stanford.edu:8090/459/1/2000-39.pdf">Toutanova, Kristina, and Christopher D. Manning. “Enriching the knowledge sources used in a maximum entropy part-of-speech tagger.” Proceedings of the 2000 Joint SIGDAT conference on Empirical methods in natural language processing and very large corpora: held in conjunction with the 38th Annual Meeting of the Association for Computational Linguistics-Volume 13. Association for Computational Linguistics, 2000.</a></p><p><a href="http://nlp.stanford.edu/pubs/tagging.pdf">Toutanova, Kristina, et al. “Feature-rich part-of-speech tagging with a cyclic dependency network.” Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology-Volume 1. Association for Computational Linguistics, 2003.</a></p><p><a href="https://www.aclweb.org/anthology/J/J93/J93-2006.pdf">Weischedel, Ralph, et al. “Coping with ambiguity and unknown words through probabilistic models.” Computational linguistics 19.2 (1993): 361-382.</a></p>]]></content>
    
    
    <summary type="html">&lt;p&gt;&lt;em&gt;&lt;strong&gt;Abstract&lt;/strong&gt;: I did some computer-y stuff to construct a personal Spanish text corpus and create a Spanish verb study guide specifically tailored to the &lt;a href=&quot;https://en.wikipedia.org/wiki/Variety_(linguistics)&quot;&gt;linguistic variety&lt;/a&gt; of Spanish I intend to consume and produce. It worked fairly well. It also revealed a (in some small way) generalizable depiction of the relative frequencies of Spanish verb tenses and moods. This technique may prove to be extremely beneficial to Spanish-language pedagogy. If you’re uninterested in my motivations or procedure, you can skip to the section labeled “results”.&lt;/em&gt;&lt;/p&gt;</summary>
    
    
    
    <category term="natural language processing" scheme="https://onthelambda.com/categories/natural-language-processing/"/>
    
    
    <category term="R" scheme="https://onthelambda.com/tags/R/"/>
    
    <category term="linguistics" scheme="https://onthelambda.com/tags/linguistics/"/>
    
    <category term="spanish" scheme="https://onthelambda.com/tags/spanish/"/>
    
    <category term="natural language processing" scheme="https://onthelambda.com/tags/natural-language-processing/"/>
    
    <category term="pedagogy" scheme="https://onthelambda.com/tags/pedagogy/"/>
    
    <category term="python" scheme="https://onthelambda.com/tags/python/"/>
    
    <category term="research" scheme="https://onthelambda.com/tags/research/"/>
    
    <category term="science" scheme="https://onthelambda.com/tags/science/"/>
    
    <category term="statistics" scheme="https://onthelambda.com/tags/statistics/"/>
    
    <category term="unix" scheme="https://onthelambda.com/tags/unix/"/>
    
  </entry>
  
  <entry>
    <title>Genre-based Music Recommendations Using Open Data (and the problem with recommender systems)</title>
    <link href="https://onthelambda.com/2016/01/11/genre-based-music-recommendations-using-open-data-and-the-problem-with-recommender-systems/"/>
    <id>https://onthelambda.com/2016/01/11/genre-based-music-recommendations-using-open-data-and-the-problem-with-recommender-systems/</id>
    <published>2016-01-12T03:43:47.000Z</published>
    <updated>2022-07-22T17:31:15.010Z</updated>
    
    <content type="html"><![CDATA[<p>After a long 12 months of pouring my soul into it, my book, <a href="http://bit.ly/data-analysis-with-r">Data Analysis with R</a>, was finally published. After the requisite 2-4 day breather, I started thinking about how I was going to get back into the swing of regular blog posts and decided that the easier and softer way is to cannibalize and expand on an example in the book.</p><span id="more"></span><p>In the chapter “Sources of Data” I show how to consume web data of different formats in R. The motivating example is to build a simple recommendation system that uses user-supplied “tags” (genres&#x2F;labels) submitted to <a href="http://www.last.fm/">Last.fm</a> and <a href="http://musicbrainz.org/">MusicBrainz</a> to quantify musical artist “similarity”. The example in the book stops at the construction and sorting of the similarity matrix but, in this post, we’re going to make a really fly D3 visualization of the musical similarity network and provide recommendations in the tooltips. The code, including the Javascript and HTML, I used for this post was hastily thrown into <a href="https://github.com/tonyfischetti/genre-based-music-recommendations">a git repo</a> and is available here. If you’re uninterested in the detailed methodology, I suggest you skip to the section labeled “Outcome”.</p><h3 id="Methodology"><a href="#Methodology" class="headerlink" title="Methodology"></a>Methodology</h3><p>Although in the book tags from both Last.fm and MusicBrainz are used, we’ll just be using Last.fm here. (In additional contrast to the book, the code here is, as you might imagine, substantially faster-paced.)</p><p>The first step is to make a character vector of all the artists that you’d like to be included. If you were building a real system, you’d probably want all Last.fm artists. Since we’re not, I just used 70 of my most played artists on <a href="http://www.last.fm/user/statethatiamin">my Last.fm</a>. Since I got the list straight from the source, I didn’t have to worry that any of the API requests would return “No Artist Found”.</p><p>The following is a function that takes an artist and returns the properly formatted Last.fm API call to get the tags in JSON format.</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">create_artist_query_url_lfm <span class="operator">&lt;-</span> <span class="keyword">function</span><span class="punctuation">(</span>artist_name<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">  prefix <span class="operator">&lt;-</span> <span class="string">&quot;http://ws.audioscrobbler.com/2.0/?method=artist.gettoptags&amp;artist=&quot;</span></span><br><span class="line">  postfix <span class="operator">&lt;-</span> <span class="string">&quot;&amp;api_key=c2e57923a25c03f3d8b317b3c8622b43&amp;format=json&quot;</span></span><br><span class="line">  encoded_artist <span class="operator">&lt;-</span> URLencode<span class="punctuation">(</span>artist_name<span class="punctuation">)</span> <span class="built_in">return</span><span class="punctuation">(</span>paste0<span class="punctuation">(</span>prefix<span class="punctuation">,</span> encoded_artist<span class="punctuation">,</span> postfix<span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line"><span class="punctuation">&#125;</span></span><br></pre></td></tr></table></figure><p><a href="http://ws.audioscrobbler.com/2.0/?method=artist.gettoptags&artist=La%20Banda%20Gorda&api_key=c2e57923a25c03f3d8b317b3c8622b43&format=json">This is an example</a> of the JSON payload from my favorite merengue artist.</p><p>We only want the tag names–curiously, attempts to factor in degree of tag fit (the “count” attribute) resulted in (what I interpreted as) substantially poorer recommendations.</p><p>The following is a function that will return a vector of all the tags.</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line">library<span class="punctuation">(</span>jsonlite<span class="punctuation">)</span></span><br><span class="line"></span><br><span class="line">get_tag_frame_lfm <span class="operator">&lt;-</span> <span class="keyword">function</span><span class="punctuation">(</span>an_artist<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">  print<span class="punctuation">(</span>paste0<span class="punctuation">(</span><span class="string">&quot;Attempting to fetch: &quot;</span><span class="punctuation">,</span> an_artist<span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line">  artist_url <span class="operator">&lt;-</span> create_artist_query_url_lfm<span class="punctuation">(</span>an_artist<span class="punctuation">)</span></span><br><span class="line">  json <span class="operator">&lt;-</span> fromJSON<span class="punctuation">(</span>artist_url<span class="punctuation">)</span> <span class="built_in">return</span><span class="punctuation">(</span>as.vector<span class="punctuation">(</span>json<span class="operator">$</span>toptags<span class="operator">$</span>tag<span class="punctuation">[</span><span class="punctuation">,</span><span class="string">&quot;name&quot;</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line"><span class="punctuation">&#125;</span></span><br></pre></td></tr></table></figure><p>Since the above function is <a href="https://en.wikipedia.org/wiki/Referential_transparency">referentially transparent</a>, and it involves using resources that aren’t yours, it’s a good idea to <a href="https://en.wikipedia.org/wiki/Memoization">memoize</a> the function so that if you (accidentally or otherwise) call the function with the same artist, the function will return the cached result instead of making the web request again. This can be achieved quite easily with the <a href="https://cran.r-project.org/web/packages/memoise/index.html"><code>memoise</code></a> package.</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line">library<span class="punctuation">(</span>memoise<span class="punctuation">)</span></span><br><span class="line"></span><br><span class="line">mem_get_tag_frame_lfm <span class="operator">&lt;-</span> memoise<span class="punctuation">(</span>get_tag_frame_lfm<span class="punctuation">)</span></span><br></pre></td></tr></table></figure><p>To get the tags from all the artists in our custom ARTIST_LIST vector..</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br></pre></td><td class="code"><pre><span class="line">artists_tags <span class="operator">&lt;-</span> sapply<span class="punctuation">(</span>ARTIST_LIST<span class="punctuation">,</span> mem_get_tag_frame_lfm<span class="punctuation">)</span></span><br><span class="line"><span class="built_in">names</span><span class="punctuation">(</span>artists_tags<span class="punctuation">)</span> <span class="operator">&lt;-</span> ARTIST_LIST</span><br></pre></td></tr></table></figure><p>To get a list of all pairs of artists to compute the similarity for, we can use the <code>combn</code> function to create a 2 by 2,415 character matrix of all possible combinations (choose 2). Let’s get that into a 2,415 by 2 <code>data.frame</code> with the name “artist1” and “artist2”…</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br></pre></td><td class="code"><pre><span class="line">cmbs <span class="operator">&lt;-</span> combn<span class="punctuation">(</span>ARTIST_LIST<span class="punctuation">,</span> <span class="number">2</span><span class="punctuation">)</span></span><br><span class="line">comparisons <span class="operator">&lt;-</span> data.frame<span class="punctuation">(</span>t<span class="punctuation">(</span>cmbs<span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line"><span class="built_in">names</span><span class="punctuation">(</span>comparisons<span class="punctuation">)</span> <span class="operator">&lt;-</span> <span class="built_in">c</span><span class="punctuation">(</span><span class="string">&quot;artist1&quot;</span><span class="punctuation">,</span> <span class="string">&quot;artist2&quot;</span><span class="punctuation">)</span></span><br></pre></td></tr></table></figure><p>The similarity metric we’ll be using is simple as all get-out: <a href="https://en.wikipedia.org/wiki/Jaccard_index">the Jaccard index</a>. Assuming we put the tags from both artists into two sets, it is the cardinality of the sets’ intersection divided by the sets’ union…</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line">jaccard_index <span class="operator">&lt;-</span> <span class="keyword">function</span><span class="punctuation">(</span>tags1<span class="punctuation">,</span> tags2<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">  <span class="built_in">length</span><span class="punctuation">(</span>intersect<span class="punctuation">(</span>tags1<span class="punctuation">,</span> tags2<span class="punctuation">)</span><span class="punctuation">)</span><span class="operator">/</span><span class="built_in">length</span><span class="punctuation">(</span>union<span class="punctuation">(</span>tags1<span class="punctuation">,</span> tags2<span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line"><span class="punctuation">&#125;</span></span><br><span class="line"></span><br><span class="line">comparisons<span class="operator">$</span>similarity <span class="operator">&lt;-</span> apply<span class="punctuation">(</span>comparisons<span class="punctuation">,</span> <span class="number">1</span><span class="punctuation">,</span> <span class="keyword">function</span><span class="punctuation">(</span>arow<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">  jaccard_index<span class="punctuation">(</span>artists_tags<span class="punctuation">[[</span>unlist<span class="punctuation">(</span>arow<span class="punctuation">[</span><span class="number">1</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">]</span><span class="punctuation">]</span><span class="punctuation">,</span> artists_tags<span class="punctuation">[[</span>unlist<span class="punctuation">(</span>arow<span class="punctuation">[</span><span class="number">2</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">]</span><span class="punctuation">]</span><span class="punctuation">)</span></span><br><span class="line"><span class="punctuation">&#125;</span><span class="punctuation">)</span></span><br></pre></td></tr></table></figure><p>Now we’ve added a new column to our previously 2,415 by 2 <code>data.frame</code>, “similarity” that contains the Jaccard index.</p><p>Our D3 visualization expects a JSON with two top level attributes: “nodes” and “links”. The “nodes” attribute is an array of <em>x</em> number of 5 key-value pairs (where <em>x</em> is the number of nodes). The 5 keys are “name” (the name of the artist) “group” (a number that affects the coloring of the node in the visualization that we will be setting to “1”), and “first”, “second”, and “third”, which are the top 3 most similar artists and will serve as the recommendations that pop-up in a tool-tip when you mouse over an artist node in the visualization.</p><p>This is some code to get the top 3 most similar artists. It takes the 2,415 by 3 <code>comparisons</code> data.frame, the number of “most similar artists” to return, an artist, and an arbitrary threshold for “similar-ness” as arguments. Any similarity below this threshold will not be considered a viable recommendation.</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br></pre></td><td class="code"><pre><span class="line">library<span class="punctuation">(</span>dplyr<span class="punctuation">)</span></span><br><span class="line"></span><br><span class="line">get_top_n <span class="operator">&lt;-</span> <span class="keyword">function</span><span class="punctuation">(</span>comparisons<span class="punctuation">,</span> N<span class="punctuation">,</span> artist<span class="punctuation">,</span> threshold<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">  comparisons <span class="operator">%&lt;&gt;%</span></span><br><span class="line">    filter<span class="punctuation">(</span>artist1<span class="operator">==</span>artist artist2<span class="operator">==</span>artist<span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">    arrange<span class="punctuation">(</span>desc<span class="punctuation">(</span>similarity<span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line">  other_artist <span class="operator">&lt;-</span> ifelse<span class="punctuation">(</span>comparisons<span class="operator">$</span>similarity<span class="operator">&gt;</span>threshold<span class="punctuation">,</span></span><br><span class="line">                    ifelse<span class="punctuation">(</span>comparisons<span class="operator">$</span>artist1<span class="operator">==</span>artist<span class="punctuation">,</span></span><br><span class="line">                           comparisons<span class="operator">$</span>artist2<span class="punctuation">,</span> comparisons<span class="operator">$</span>artist1<span class="punctuation">)</span><span class="punctuation">,</span> <span class="string">&quot;None&quot;</span><span class="punctuation">)</span></span><br><span class="line">  <span class="built_in">return</span><span class="punctuation">(</span>other_artist<span class="punctuation">[</span><span class="number">1</span><span class="operator">:</span>N<span class="punctuation">]</span><span class="punctuation">)</span></span><br><span class="line"><span class="punctuation">&#125;</span></span><br></pre></td></tr></table></figure><p>The inner <code>ifelse</code> clause has to handle the fact that the “similar” artist can be in the first column or the second column. The outer <code>ifelse</code> returns “None” for every <code>similarity</code> value that is not above the threshold.</p><p>Let’s make the <code>data.frame</code> that will serve as the “nodes” attribute in the final JSON…</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br></pre></td><td class="code"><pre><span class="line">nodes <span class="operator">&lt;-</span> sapply<span class="punctuation">(</span>ARTIST_LIST<span class="punctuation">,</span> <span class="keyword">function</span><span class="punctuation">(</span>x<span class="punctuation">)</span> get_top_n<span class="punctuation">(</span>comparisons<span class="punctuation">,</span> <span class="number">3</span><span class="punctuation">,</span> x<span class="punctuation">,</span> <span class="number">0.25</span><span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line">nodes <span class="operator">&lt;-</span> data.frame<span class="punctuation">(</span>t<span class="punctuation">(</span>nodes<span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line"><span class="built_in">names</span><span class="punctuation">(</span>nodes<span class="punctuation">)</span> <span class="operator">&lt;-</span> <span class="built_in">c</span><span class="punctuation">(</span><span class="string">&quot;first&quot;</span><span class="punctuation">,</span> <span class="string">&quot;second&quot;</span><span class="punctuation">,</span> <span class="string">&quot;third&quot;</span><span class="punctuation">)</span></span><br><span class="line">nodes<span class="operator">$</span>name <span class="operator">&lt;-</span> row.names<span class="punctuation">(</span>nodes<span class="punctuation">)</span></span><br><span class="line">row.names<span class="punctuation">(</span>nodes<span class="punctuation">)</span> <span class="operator">&lt;-</span> <span class="literal">NULL</span></span><br><span class="line">nodes<span class="operator">$</span>group <span class="operator">&lt;-</span> 1</span><br></pre></td></tr></table></figure><p>For the other top-level JSON attribute, “links”, we need an array of <em>y</em> number of 5 key-value pairs where <em>y</em> is the number of sufficiently strong similarities between the artists. The 5 keys are “node1” (the name of the first artist), “source” (the 0-indexed index of the artist with respect to the array in the “nodes” attribute), “node2” (the name of the second artist), “target” (the index of the second artist) and “weight”, which is the degree of similarity between the two artists; this will translate into thicker “edges” in the similarity graph.</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br></pre></td><td class="code"><pre><span class="line"><span class="comment"># find the 0-indexed index</span></span><br><span class="line"></span><br><span class="line">lookup_number <span class="operator">&lt;-</span> <span class="keyword">function</span><span class="punctuation">(</span>name<span class="punctuation">)</span></span><br><span class="line">  which<span class="punctuation">(</span>name<span class="operator">==</span>ARTIST_LIST<span class="punctuation">)</span><span class="operator">-</span><span class="number">1</span></span><br><span class="line"></span><br><span class="line">strong_links <span class="operator">&lt;-</span> comparisons <span class="operator">%&gt;%</span></span><br><span class="line">  filter<span class="punctuation">(</span>similarity <span class="operator">&gt;</span> <span class="number">0.25</span><span class="punctuation">)</span> <span class="operator">%&gt;%</span></span><br><span class="line">  rename<span class="punctuation">(</span>node1 <span class="operator">=</span> artist1<span class="punctuation">,</span> node2 <span class="operator">=</span> artist2<span class="punctuation">,</span> weight<span class="operator">=</span>similarity<span class="punctuation">)</span></span><br><span class="line"></span><br><span class="line">strong_links<span class="operator">$</span>source <span class="operator">&lt;-</span> sapply<span class="punctuation">(</span>strong_links<span class="operator">$</span>node1<span class="punctuation">,</span> lookup_number<span class="punctuation">)</span></span><br><span class="line">strong_links<span class="operator">$</span>target <span class="operator">&lt;-</span> sapply<span class="punctuation">(</span>strong_links<span class="operator">$</span>node2<span class="punctuation">,</span> lookup_number<span class="punctuation">)</span></span><br></pre></td></tr></table></figure><p>Finally, we can create the properly formatted JSON and send it to the file “artists.json” thusly…</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br></pre></td><td class="code"><pre><span class="line">object <span class="operator">&lt;-</span> <span class="built_in">list</span><span class="punctuation">(</span><span class="string">&quot;nodes&quot;</span><span class="operator">=</span>nodes<span class="punctuation">,</span> <span class="string">&quot;links&quot;</span><span class="operator">=</span>strong_links<span class="punctuation">)</span></span><br><span class="line">sink<span class="punctuation">(</span><span class="string">&quot;artists.json&quot;</span><span class="punctuation">)</span></span><br><span class="line">toJSON<span class="punctuation">(</span>object<span class="punctuation">,</span> dataframe<span class="operator">=</span><span class="string">&quot;rows&quot;</span><span class="punctuation">,</span> pretty<span class="operator">=</span><span class="literal">TRUE</span><span class="punctuation">)</span></span><br><span class="line">sink<span class="punctuation">(</span><span class="punctuation">)</span></span><br></pre></td></tr></table></figure><h3 id="Outcome"><a href="#Outcome" class="headerlink" title="Outcome"></a>Outcome</h3><p><a href="/images/music-similarity.png"><img src="/images/music-similarity.png" alt="Musical Similarity Network"></a></p><p>Using “artists.json” and the “index.html” that can be found <a href="https://github.com/tonyfischetti/genre-based-music-recommendations/blob/master/index.html">here</a>, <a href="http://statethatiamin.onlythisrose.com/genre-based-music-recommendations/">the similarity graph looks a little like this.</a> (Make sure you scroll to see the whole thing.)</p><p>For illustrative purposes, I pre-labeled the artists’ “group” with labels that correspond to what I view as the artist’s primary genre. This is why the nodes in the linked visualization have different colors. Note that, independently, the genres that I indicated tend to cluster together in the network. For example, Reggae (light green), Hip-Hop (green), and Punk (orange) all form almost completely connected graphs, though unconnected to each other (disjoint subgraphs). Indie rock (blue), post-punk (light blue) and classic rock (light orange) together form a rather tightly-connected subgraph. Curiously, the Sex Pistols (that I labeled “Punk”) are not part of the Punk cluster but part of the Indie-rock&#x2F;post-punk&#x2F;classic-rock component. There are three orphan nodes (no edges), “Johann Sebastian Bach”, “<a href="http://www.last.fm/music/P:ano">P:ano</a>“, and “<a href="http://www.last.fm/music/No+Kids">No Kids</a>“. Bach is orphaned because he’s the only Baroque artist in my top 70 artists :( –P:ano and No Kids are obscure… you’ve probably never heard of them.</p><p>The recommendations, prima facie, appear to be on point. For example, without direct knowledge of association, “KRS-One” recommends “Boogie Down Productions” (the group that KRS-One comes from) most highly. Similarly, “The Smiths” and “Morrissey” recommend each other, and “De La Soul” and “A Tribe Called Quest” (part of a positive, Afrocentric hip-hop collective known as <a href="https://en.wikipedia.org/wiki/Native_Tongues">the Native Tongues</a> together with Queen Latifah, et al.) recommend each other.</p><p>Appropriately, Joy Division and New Order, whose Jaccard index of band members is 0.6 but whose music style is somewhat distinct, don’t recommend each other. Lastly, subgenred artists appear to recommend other artists in the subgenre. For example, goth band “The Sisters of Mercy” appropriately recommends other goth-esque bands “Bauhaus”, “And Also The Trees”, and “Joy Division”.</p><h3 id="Afterword"><a href="#Afterword" class="headerlink" title="Afterword"></a>Afterword</h3><p>Using this similarity measure to drive recommendations seems successful. It should be noted, though, that my ability to assess the effectiveness of using the Jaccard index as the sole arbiter of musical similarity is hampered; judging an algorithm on the basis that the system recommends other bands that I <em>necessarily</em> like is prejudicial, to say the least.</p><p>This stands even if the system makes good theoretical sense. This still stands even if the system, quite independently, indicates that associated acts—that are objectively and incontrovertibly similar—are good recommendations.</p><p>This raises a larger question on how to accurately measure the effectiveness of recommender systems; do you tell people what they want to hear, or do you pledge allegiance to a particular theoretical interpretation of similarity? If it’s the latter, how do you iterate and improve the system? If it’s the former, is your only criterion for success positive user-provided feedback?</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;After a long 12 months of pouring my soul into it, my book, &lt;a href=&quot;http://bit.ly/data-analysis-with-r&quot;&gt;Data Analysis with R&lt;/a&gt;, was finally published. After the requisite 2-4 day breather, I started thinking about how I was going to get back into the swing of regular blog posts and decided that the easier and softer way is to cannibalize and expand on an example in the book.&lt;/p&gt;</summary>
    
    
    
    <category term="R" scheme="https://onthelambda.com/categories/R/"/>
    
    
    <category term="R" scheme="https://onthelambda.com/tags/R/"/>
    
    <category term="research" scheme="https://onthelambda.com/tags/research/"/>
    
    <category term="D3" scheme="https://onthelambda.com/tags/D3/"/>
    
    <category term="datavis" scheme="https://onthelambda.com/tags/datavis/"/>
    
    <category term="graph theory" scheme="https://onthelambda.com/tags/graph-theory/"/>
    
    <category term="music" scheme="https://onthelambda.com/tags/music/"/>
    
    <category term="open data" scheme="https://onthelambda.com/tags/open-data/"/>
    
    <category term="recommender systems" scheme="https://onthelambda.com/tags/recommender-systems/"/>
    
  </entry>
  
  <entry>
    <title>Kickin&#39; it with elastic net regression</title>
    <link href="https://onthelambda.com/2015/08/19/kickin-it-with-elastic-net-regression/"/>
    <id>https://onthelambda.com/2015/08/19/kickin-it-with-elastic-net-regression/</id>
    <published>2015-08-20T00:47:19.000Z</published>
    <updated>2022-07-23T00:40:02.186Z</updated>
    
    <content type="html"><![CDATA[<p>With the kind of data that I usually work with, overfitting regression models can be a huge problem if I’m not careful. Ridge regression is a really effective technique for thwarting overfitting. It does this by penalizing the L2 norm (euclidean distance) of the coefficient vector which results in “shrinking” the beta coefficients. The aggressiveness of the penalty is controlled by a parameter <mjx-container class="MathJax" jax="SVG"><svg style="vertical-align: -0.027ex;" xmlns="http://www.w3.org/2000/svg" width="1.319ex" height="1.597ex" role="img" focusable="false" viewBox="0 -694 583 706"><g stroke="currentColor" fill="currentColor" stroke-width="0" transform="scale(1,-1)"><g data-mml-node="math"><g data-mml-node="mi"><path data-c="1D706" d="M166 673Q166 685 183 694H202Q292 691 316 644Q322 629 373 486T474 207T524 67Q531 47 537 34T546 15T551 6T555 2T556 -2T550 -11H482Q457 3 450 18T399 152L354 277L340 262Q327 246 293 207T236 141Q211 112 174 69Q123 9 111 -1T83 -12Q47 -12 47 20Q47 37 61 52T199 187Q229 216 266 252T321 306L338 322Q338 323 288 462T234 612Q214 657 183 657Q166 657 166 673Z"></path></g></g></g></svg></mjx-container>.</p><span id="more"></span><p>Lasso regression is a related regularization method. Instead of using the L2 norm, though, it penalizes the L1 norm (manhattan distance) of the coefficient vector.</p><p>Because it uses the L1 norm, some of the coefficients will shrink to zero while lambda increases. A similar effect would be achieved in Bayesian linear regression using a Laplacian prior (strongly peaked at zero) on each of the beta coefficients. Because some of the coefficients shrink to zero, the lasso doubles as a crackerjack feature selection technique in addition to a solid shrinkage method. This property gives it a leg up on ridge regression. On the other hand, the lasso will occasionally achieve poor results when there’s a high degree of collinearity in the features and ridge regression will perform better. Further, the L1 norm is underdetermined when the number of predictors exceeds the number of observations while ridge regression can handle this.</p><p>Elastic net regression is a hybrid approach that blends both penalization of the L2 <em>and</em> L1 norms. Specifically, elastic net regression minimizes the following…</p><p><mjx-container class="MathJax" jax="SVG" display="true"><svg style="vertical-align: -0.566ex;" xmlns="http://www.w3.org/2000/svg" width="33.059ex" height="2.565ex" role="img" focusable="false" viewBox="0 -883.9 14611.9 1133.9"><g stroke="currentColor" fill="currentColor" stroke-width="0" transform="scale(1,-1)"><g data-mml-node="math"><g data-mml-node="mo"><path data-c="2016" d="M133 736Q138 750 153 750Q164 750 170 739Q172 735 172 250T170 -239Q164 -250 152 -250Q144 -250 138 -244L137 -243Q133 -241 133 -179T132 250Q132 731 133 736ZM329 739Q334 750 346 750Q353 750 361 744L362 743Q366 741 366 679T367 250T367 -178T362 -243L361 -244Q355 -250 347 -250Q335 -250 329 -239Q327 -235 327 250T329 739Z"></path></g><g data-mml-node="mi" transform="translate(500,0)"><path data-c="1D466" d="M21 287Q21 301 36 335T84 406T158 442Q199 442 224 419T250 355Q248 336 247 334Q247 331 231 288T198 191T182 105Q182 62 196 45T238 27Q261 27 281 38T312 61T339 94Q339 95 344 114T358 173T377 247Q415 397 419 404Q432 431 462 431Q475 431 483 424T494 412T496 403Q496 390 447 193T391 -23Q363 -106 294 -155T156 -205Q111 -205 77 -183T43 -117Q43 -95 50 -80T69 -58T89 -48T106 -45Q150 -45 150 -87Q150 -107 138 -122T115 -142T102 -147L99 -148Q101 -153 118 -160T152 -167H160Q177 -167 186 -165Q219 -156 247 -127T290 -65T313 -9T321 21L315 17Q309 13 296 6T270 -6Q250 -11 231 -11Q185 -11 150 11T104 82Q103 89 103 113Q103 170 138 262T173 379Q173 380 173 381Q173 390 173 393T169 400T158 404H154Q131 404 112 385T82 344T65 302T57 280Q55 278 41 278H27Q21 284 21 287Z"></path></g><g data-mml-node="mo" transform="translate(1212.2,0)"><path data-c="2212" d="M84 237T84 250T98 270H679Q694 262 694 250T679 230H98Q84 237 84 250Z"></path></g><g data-mml-node="mi" transform="translate(2212.4,0)"><path data-c="1D44B" d="M42 0H40Q26 0 26 11Q26 15 29 27Q33 41 36 43T55 46Q141 49 190 98Q200 108 306 224T411 342Q302 620 297 625Q288 636 234 637H206Q200 643 200 645T202 664Q206 677 212 683H226Q260 681 347 681Q380 681 408 681T453 682T473 682Q490 682 490 671Q490 670 488 658Q484 643 481 640T465 637Q434 634 411 620L488 426L541 485Q646 598 646 610Q646 628 622 635Q617 635 609 637Q594 637 594 648Q594 650 596 664Q600 677 606 683H618Q619 683 643 683T697 681T738 680Q828 680 837 683H845Q852 676 852 672Q850 647 840 637H824Q790 636 763 628T722 611T698 593L687 584Q687 585 592 480L505 384Q505 383 536 304T601 142T638 56Q648 47 699 46Q734 46 734 37Q734 35 732 23Q728 7 725 4T711 1Q708 1 678 1T589 2Q528 2 496 2T461 1Q444 1 444 10Q444 11 446 25Q448 35 450 39T455 44T464 46T480 47T506 54Q523 62 523 64Q522 64 476 181L429 299Q241 95 236 84Q232 76 232 72Q232 53 261 47Q262 47 267 47T273 46Q276 46 277 46T280 45T283 42T284 35Q284 26 282 19Q279 6 276 4T261 1Q258 1 243 1T201 2T142 2Q64 2 42 0Z"></path></g><g data-mml-node="mi" transform="translate(3064.4,0)"><path data-c="1D6FD" d="M29 -194Q23 -188 23 -186Q23 -183 102 134T186 465Q208 533 243 584T309 658Q365 705 429 705H431Q493 705 533 667T573 570Q573 465 469 396L482 383Q533 332 533 252Q533 139 448 65T257 -10Q227 -10 203 -2T165 17T143 40T131 59T126 65L62 -188Q60 -194 42 -194H29ZM353 431Q392 431 427 419L432 422Q436 426 439 429T449 439T461 453T472 471T484 495T493 524T501 560Q503 569 503 593Q503 611 502 616Q487 667 426 667Q384 667 347 643T286 582T247 514T224 455Q219 439 186 308T152 168Q151 163 151 147Q151 99 173 68Q204 26 260 26Q302 26 349 51T425 137Q441 171 449 214T457 279Q457 337 422 372Q380 358 347 358H337Q258 358 258 389Q258 396 261 403Q275 431 353 431Z"></path></g><g data-mml-node="mo" transform="translate(3630.4,0)"><path data-c="2016" d="M133 736Q138 750 153 750Q164 750 170 739Q172 735 172 250T170 -239Q164 -250 152 -250Q144 -250 138 -244L137 -243Q133 -241 133 -179T132 250Q132 731 133 736ZM329 739Q334 750 346 750Q353 750 361 744L362 743Q366 741 366 679T367 250T367 -178T362 -243L361 -244Q355 -250 347 -250Q335 -250 329 -239Q327 -235 327 250T329 739Z"></path></g><g data-mml-node="mo" transform="translate(4352.7,0)"><path data-c="2B" d="M56 237T56 250T70 270H369V420L370 570Q380 583 389 583Q402 583 409 568V270H707Q722 262 722 250T707 230H409V-68Q401 -82 391 -82H389H387Q375 -82 369 -68V230H70Q56 237 56 250Z"></path></g><g data-mml-node="mi" transform="translate(5352.9,0)"><path data-c="1D706" d="M166 673Q166 685 183 694H202Q292 691 316 644Q322 629 373 486T474 207T524 67Q531 47 537 34T546 15T551 6T555 2T556 -2T550 -11H482Q457 3 450 18T399 152L354 277L340 262Q327 246 293 207T236 141Q211 112 174 69Q123 9 111 -1T83 -12Q47 -12 47 20Q47 37 61 52T199 187Q229 216 266 252T321 306L338 322Q338 323 288 462T234 612Q214 657 183 657Q166 657 166 673Z"></path></g><g data-mml-node="mo" transform="translate(5935.9,0)"><path data-c="5B" d="M118 -250V750H255V710H158V-210H255V-250H118Z"></path></g><g data-mml-node="mo" transform="translate(6213.9,0)"><path data-c="28" d="M94 250Q94 319 104 381T127 488T164 576T202 643T244 695T277 729T302 750H315H319Q333 750 333 741Q333 738 316 720T275 667T226 581T184 443T167 250T184 58T225 -81T274 -167T316 -220T333 -241Q333 -250 318 -250H315H302L274 -226Q180 -141 137 -14T94 250Z"></path></g><g data-mml-node="mn" transform="translate(6602.9,0)"><path data-c="31" d="M213 578L200 573Q186 568 160 563T102 556H83V602H102Q149 604 189 617T245 641T273 663Q275 666 285 666Q294 666 302 660V361L303 61Q310 54 315 52T339 48T401 46H427V0H416Q395 3 257 3Q121 3 100 0H88V46H114Q136 46 152 46T177 47T193 50T201 52T207 57T213 61V578Z"></path></g><g data-mml-node="mo" transform="translate(7325.1,0)"><path data-c="2212" d="M84 237T84 250T98 270H679Q694 262 694 250T679 230H98Q84 237 84 250Z"></path></g><g data-mml-node="mi" transform="translate(8325.3,0)"><path data-c="1D6FC" d="M34 156Q34 270 120 356T309 442Q379 442 421 402T478 304Q484 275 485 237V208Q534 282 560 374Q564 388 566 390T582 393Q603 393 603 385Q603 376 594 346T558 261T497 161L486 147L487 123Q489 67 495 47T514 26Q528 28 540 37T557 60Q559 67 562 68T577 70Q597 70 597 62Q597 56 591 43Q579 19 556 5T512 -10H505Q438 -10 414 62L411 69L400 61Q390 53 370 41T325 18T267 -2T203 -11Q124 -11 79 39T34 156ZM208 26Q257 26 306 47T379 90L403 112Q401 255 396 290Q382 405 304 405Q235 405 183 332Q156 292 139 224T121 120Q121 71 146 49T208 26Z"></path></g><g data-mml-node="mo" transform="translate(8965.3,0)"><path data-c="29" d="M60 749L64 750Q69 750 74 750H86L114 726Q208 641 251 514T294 250Q294 182 284 119T261 12T224 -76T186 -143T145 -194T113 -227T90 -246Q87 -249 86 -250H74Q66 -250 63 -250T58 -247T55 -238Q56 -237 66 -225Q221 -64 221 250T66 725Q56 737 55 738Q55 746 60 749Z"></path></g><g data-mml-node="mo" transform="translate(9354.3,0) translate(0 -0.5)"><path data-c="7C" d="M139 -249H137Q125 -249 119 -235V251L120 737Q130 750 139 750Q152 750 159 735V-235Q151 -249 141 -249H139Z"></path></g><g data-mml-node="mi" transform="translate(9632.3,0)"><path data-c="1D6FD" d="M29 -194Q23 -188 23 -186Q23 -183 102 134T186 465Q208 533 243 584T309 658Q365 705 429 705H431Q493 705 533 667T573 570Q573 465 469 396L482 383Q533 332 533 252Q533 139 448 65T257 -10Q227 -10 203 -2T165 17T143 40T131 59T126 65L62 -188Q60 -194 42 -194H29ZM353 431Q392 431 427 419L432 422Q436 426 439 429T449 439T461 453T472 471T484 495T493 524T501 560Q503 569 503 593Q503 611 502 616Q487 667 426 667Q384 667 347 643T286 582T247 514T224 455Q219 439 186 308T152 168Q151 163 151 147Q151 99 173 68Q204 26 260 26Q302 26 349 51T425 137Q441 171 449 214T457 279Q457 337 422 372Q380 358 347 358H337Q258 358 258 389Q258 396 261 403Q275 431 353 431Z"></path></g><g data-mml-node="msubsup" transform="translate(10198.3,0)"><g data-mml-node="mo" transform="translate(0 -0.5)"><path data-c="7C" d="M139 -249H137Q125 -249 119 -235V251L120 737Q130 750 139 750Q152 750 159 735V-235Q151 -249 141 -249H139Z"></path></g><g data-mml-node="mn" transform="translate(311,413) scale(0.707)"><path data-c="32" d="M109 429Q82 429 66 447T50 491Q50 562 103 614T235 666Q326 666 387 610T449 465Q449 422 429 383T381 315T301 241Q265 210 201 149L142 93L218 92Q375 92 385 97Q392 99 409 186V189H449V186Q448 183 436 95T421 3V0H50V19V31Q50 38 56 46T86 81Q115 113 136 137Q145 147 170 174T204 211T233 244T261 278T284 308T305 340T320 369T333 401T340 431T343 464Q343 527 309 573T212 619Q179 619 154 602T119 569T109 550Q109 549 114 549Q132 549 151 535T170 489Q170 464 154 447T109 429Z"></path></g><g data-mml-node="mn" transform="translate(311,-247) scale(0.707)"><path data-c="32" d="M109 429Q82 429 66 447T50 491Q50 562 103 614T235 666Q326 666 387 610T449 465Q449 422 429 383T381 315T301 241Q265 210 201 149L142 93L218 92Q375 92 385 97Q392 99 409 186V189H449V186Q448 183 436 95T421 3V0H50V19V31Q50 38 56 46T86 81Q115 113 136 137Q145 147 170 174T204 211T233 244T261 278T284 308T305 340T320 369T333 401T340 431T343 464Q343 527 309 573T212 619Q179 619 154 602T119 569T109 550Q109 549 114 549Q132 549 151 535T170 489Q170 464 154 447T109 429Z"></path></g></g><g data-mml-node="mo" transform="translate(11135.1,0)"><path data-c="2B" d="M56 237T56 250T70 270H369V420L370 570Q380 583 389 583Q402 583 409 568V270H707Q722 262 722 250T707 230H409V-68Q401 -82 391 -82H389H387Q375 -82 369 -68V230H70Q56 237 56 250Z"></path></g><g data-mml-node="mi" transform="translate(12135.3,0)"><path data-c="1D6FC" d="M34 156Q34 270 120 356T309 442Q379 442 421 402T478 304Q484 275 485 237V208Q534 282 560 374Q564 388 566 390T582 393Q603 393 603 385Q603 376 594 346T558 261T497 161L486 147L487 123Q489 67 495 47T514 26Q528 28 540 37T557 60Q559 67 562 68T577 70Q597 70 597 62Q597 56 591 43Q579 19 556 5T512 -10H505Q438 -10 414 62L411 69L400 61Q390 53 370 41T325 18T267 -2T203 -11Q124 -11 79 39T34 156ZM208 26Q257 26 306 47T379 90L403 112Q401 255 396 290Q382 405 304 405Q235 405 183 332Q156 292 139 224T121 120Q121 71 146 49T208 26Z"></path></g><g data-mml-node="mo" transform="translate(12775.3,0) translate(0 -0.5)"><path data-c="7C" d="M139 -249H137Q125 -249 119 -235V251L120 737Q130 750 139 750Q152 750 159 735V-235Q151 -249 141 -249H139Z"></path></g><g data-mml-node="mi" transform="translate(13053.3,0)"><path data-c="1D6FD" d="M29 -194Q23 -188 23 -186Q23 -183 102 134T186 465Q208 533 243 584T309 658Q365 705 429 705H431Q493 705 533 667T573 570Q573 465 469 396L482 383Q533 332 533 252Q533 139 448 65T257 -10Q227 -10 203 -2T165 17T143 40T131 59T126 65L62 -188Q60 -194 42 -194H29ZM353 431Q392 431 427 419L432 422Q436 426 439 429T449 439T461 453T472 471T484 495T493 524T501 560Q503 569 503 593Q503 611 502 616Q487 667 426 667Q384 667 347 643T286 582T247 514T224 455Q219 439 186 308T152 168Q151 163 151 147Q151 99 173 68Q204 26 260 26Q302 26 349 51T425 137Q441 171 449 214T457 279Q457 337 422 372Q380 358 347 358H337Q258 358 258 389Q258 396 261 403Q275 431 353 431Z"></path></g><g data-mml-node="msub" transform="translate(13619.3,0)"><g data-mml-node="mo" transform="translate(0 -0.5)"><path data-c="7C" d="M139 -249H137Q125 -249 119 -235V251L120 737Q130 750 139 750Q152 750 159 735V-235Q151 -249 141 -249H139Z"></path></g><g data-mml-node="mn" transform="translate(311,-150) scale(0.707)"><path data-c="31" d="M213 578L200 573Q186 568 160 563T102 556H83V602H102Q149 604 189 617T245 641T273 663Q275 666 285 666Q294 666 302 660V361L303 61Q310 54 315 52T339 48T401 46H427V0H416Q395 3 257 3Q121 3 100 0H88V46H114Q136 46 152 46T177 47T193 50T201 52T207 57T213 61V578Z"></path></g></g><g data-mml-node="mo" transform="translate(14333.9,0)"><path data-c="5D" d="M22 710V750H159V-250H22V-210H119V710H22Z"></path></g></g></g></svg></mjx-container></p><p>the <mjx-container class="MathJax" jax="SVG"><svg style="vertical-align: -0.025ex;" xmlns="http://www.w3.org/2000/svg" width="1.448ex" height="1.025ex" role="img" focusable="false" viewBox="0 -442 640 453"><g stroke="currentColor" fill="currentColor" stroke-width="0" transform="scale(1,-1)"><g data-mml-node="math"><g data-mml-node="mi"><path data-c="1D6FC" d="M34 156Q34 270 120 356T309 442Q379 442 421 402T478 304Q484 275 485 237V208Q534 282 560 374Q564 388 566 390T582 393Q603 393 603 385Q603 376 594 346T558 261T497 161L486 147L487 123Q489 67 495 47T514 26Q528 28 540 37T557 60Q559 67 562 68T577 70Q597 70 597 62Q597 56 591 43Q579 19 556 5T512 -10H505Q438 -10 414 62L411 69L400 61Q390 53 370 41T325 18T267 -2T203 -11Q124 -11 79 39T34 156ZM208 26Q257 26 306 47T379 90L403 112Q401 255 396 290Q382 405 304 405Q235 405 183 332Q156 292 139 224T121 120Q121 71 146 49T208 26Z"></path></g></g></g></svg></mjx-container> hyper-parameter is between 0 and 1 and controls how much L2 or L1 penalization is used (0 is ridge, 1 is lasso).</p><p>The usual approach to optimizing the lambda hyper-parameter is through cross-validation—by minimizing the cross-validated mean squared prediction error—but in elastic net regression, the optimal lambda hyper-parameter also depends upon and is heavily dependent on the alpha hyper-parameter (hyper-hyper-parameter?).</p><p>This blog post takes a cross-validated approach that uses grid search to find the optimal alpha hyper-parameter while also optimizing the lambda hyper-parameter for three different data sets. I also compare the performances against the stepwise regression and showcase some of the dangers of using stepwise feature selection.</p><h3 id="mtcars"><a href="#mtcars" class="headerlink" title="mtcars"></a>mtcars</h3><p>In this example, I try to predict “miles per gallon” from the other available attributes. The design matrix has 32 observations and 10 predictors and there is a high degree of collinearity (as measured by the <a href="https://en.wikipedia.org/wiki/Variance_inflation_factor">variance inflation factors</a>).</p><p><a href="/images/mtcarsplot.png"><img src="/images/mtcarsplot.png" alt="mtcars and elastic net regression"></a> The left panel above shows the leave-one-out cross validation (LOOCV) mean squared error of the model with the optimal lambda (as determined again by LOOCV) for each alpha parameter from 0 to 1. This panel indicates that if our objective is to purely minimize MSE (with no regard for model complexity) than pure ridge regression outperforms any blended elastic-net model. This is probably because of the substantial collinearity. Interestingly, the lasso outperforms blended elastic net models that weight the lasso heavily.</p><p>The right panel puts things in perspective by plotting the LOOCV MSEs along with the MSE of the “kitchen sink” regression (the blue line) that includes all features in the model. As you can see, any degree of regularization offers a substantial improvement in model generalizability.</p><p>It is also plotted with two estimates of the MSE for models that blindly use the coefficients from automated bi-directional stepwise regression. The first uses the features selected by performing the stepwise procedure on <em>the whole dataset</em> and then assesses the model performance (the red line). The second estimate uses the step procedure and resulting features on <em>only the training set for each fold of the cross validations</em>. This is the estimate without the subtle but treacherous “knowledge leaking” eloquently described in <a href="http://www.alfredo.motta.name/cross-validation-done-wrong/">this plot post</a>. This should be considered the more correct assessment of the model. As you can see, if we weren’t careful about interpreting the stepwise regression, we would have gotten an incredibly inflated and inaccurate view of the model performance.</p><h3 id="Forest-Fires"><a href="#Forest-Fires" class="headerlink" title="Forest Fires"></a>Forest Fires</h3><p>The second example uses a very-difficult-to-model dataset from <a href="http://archive.ics.uci.edu/ml/datasets/Forest+Fires">University of California, Irvine machine learning repository</a>. The task is to predict the burnt area from a forest fire given 11 predictors. It has 517 observations. Further, there is a relatively low degree of collinearity between predictors.</p><p><a href="/images/fireplot.png"><img src="/images/fireplot.png" alt="fireplot"></a></p><p>Again, highest performing model is the pure ridge regression. This time, the performance asymptotes as the alpha hyper-parameter increases. The variability in the MSE estimates is due to the fact that I didn’t use LOOCV and used 400-k CV instead because I’m impatient.</p><p>As with the last example, the properly measured stepwise regression performance isn’t so great, and the kitchen sink model outperforms it. However, in contrast to the previous example, there was a lot less variability in the selected features across folds—this is probably because of the significantly larger number of observations.</p><h3 id="“QuickStartExample”"><a href="#“QuickStartExample”" class="headerlink" title="“QuickStartExample”"></a>“QuickStartExample”</h3><p>This dataset is a contrived one that is included with the excellent <a href="http://web.stanford.edu/~hastie/glmnet/glmnet_alpha.html">glmnet package</a> (the one I’m using for the elastic net regression). This dataset has a relatively low degree of collinearity, has 20 features and 100 observations. I have no idea how the package authors created this dataset.</p><p><a href="/images/quickstartplot.png"><img src="/images/quickstartplot.png" alt="quickstartplot"></a></p><p>Finally, an example where the lasso outperforms ridge regression! I think this is because the dataset was specifically manufactured to have a small number of genuine predictors with large effects (as opposed to many weak predictors).</p><p>Interestingly, stepwise progression far outperforms both—probably for the very same reason. From fold to fold, there was virtually no variation in the features that the stepwise method automatically chose.</p><h3 id="Conclusion"><a href="#Conclusion" class="headerlink" title="Conclusion"></a>Conclusion</h3><p>So, there you have it. Elastic net regression is awesome because it can perform <em>at worst as good</em> as the lasso or ridge and—though it didn’t on these examples—can sometimes substantially outperform both.</p><p>Also, be careful with step-wise feature selection!</p><p>PS: If, for some reason, you are interested in the R code I used to run these simulations, you can find it <a href="https://gist.github.com/tonyfischetti/b4fd17a94aa0a6e07fdb">on this GitHub Gist</a>.</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;With the kind of data that I usually work with, overfitting regression models can be a huge problem if I’m not careful. Ridge regression is a really effective technique for thwarting overfitting. It does this by penalizing the L2 norm (euclidean distance) of the coefficient vector which results in “shrinking” the beta coefficients. The aggressiveness of the penalty is controlled by a parameter &lt;mjx-container class=&quot;MathJax&quot; jax=&quot;SVG&quot;&gt;&lt;svg style=&quot;vertical-align: -0.027ex;&quot; xmlns=&quot;http://www.w3.org/2000/svg&quot; width=&quot;1.319ex&quot; height=&quot;1.597ex&quot; role=&quot;img&quot; focusable=&quot;false&quot; viewBox=&quot;0 -694 583 706&quot;&gt;&lt;g stroke=&quot;currentColor&quot; fill=&quot;currentColor&quot; stroke-width=&quot;0&quot; transform=&quot;scale(1,-1)&quot;&gt;&lt;g data-mml-node=&quot;math&quot;&gt;&lt;g data-mml-node=&quot;mi&quot;&gt;&lt;path data-c=&quot;1D706&quot; d=&quot;M166 673Q166 685 183 694H202Q292 691 316 644Q322 629 373 486T474 207T524 67Q531 47 537 34T546 15T551 6T555 2T556 -2T550 -11H482Q457 3 450 18T399 152L354 277L340 262Q327 246 293 207T236 141Q211 112 174 69Q123 9 111 -1T83 -12Q47 -12 47 20Q47 37 61 52T199 187Q229 216 266 252T321 306L338 322Q338 323 288 462T234 612Q214 657 183 657Q166 657 166 673Z&quot;&gt;&lt;/path&gt;&lt;/g&gt;&lt;/g&gt;&lt;/g&gt;&lt;/svg&gt;&lt;/mjx-container&gt;.&lt;/p&gt;</summary>
    
    
    
    <category term="R" scheme="https://onthelambda.com/categories/R/"/>
    
    <category term="machine learning" scheme="https://onthelambda.com/categories/machine-learning/"/>
    
    
    <category term="R" scheme="https://onthelambda.com/tags/R/"/>
    
    <category term="research" scheme="https://onthelambda.com/tags/research/"/>
    
    <category term="statistics" scheme="https://onthelambda.com/tags/statistics/"/>
    
    <category term="data mining" scheme="https://onthelambda.com/tags/data-mining/"/>
    
    <category term="ggplot2" scheme="https://onthelambda.com/tags/ggplot2/"/>
    
    <category term="machine learning" scheme="https://onthelambda.com/tags/machine-learning/"/>
    
  </entry>
  
  <entry>
    <title>Lessons learned in high-performance R</title>
    <link href="https://onthelambda.com/2015/05/31/lessons-learned-in-high-performance-r/"/>
    <id>https://onthelambda.com/2015/05/31/lessons-learned-in-high-performance-r/</id>
    <published>2015-05-31T04:37:44.000Z</published>
    <updated>2022-07-23T00:39:55.082Z</updated>
    
    <content type="html"><![CDATA[<p>On this blog, I’ve had a long running investigation&#x2F;demonstration of how to make a “embarrassingly-parallel” but computationally intractable (on commodity hardware, at least) R problem more performant by using parallel computation and Rcpp.</p><span id="more"></span><p>The example problem is to find the mean distance between every airport in the United States. This silly example was chosen because it exhibits polynomial growth in running time as a function of the number of airports and, thus, quickly becomes intractable without sampling. It is also easy to parallelize.</p><p><a href="http://www.onthelambda.com/2013/11/13/parallel-r-and-air-travel/">TODO The first post</a> used the (now-deprecated in favor of ‘parallel’) multicore package to achieve a substantial speedup. <a href="http://www.onthelambda.com/2014/06/27/squeezing-more-speed-from-r-for-nothing-rcpp-style/">The second post</a> used Rcpp to achieve a statistically significant but, functionally, trivial speedup by replacing the inner loop (the distance calculation using the <a href="http://en.wikipedia.org/wiki/Haversine_formula">Haversine formula</a>) with a version written in C++ using Rcpp. Though I was disappointed in the results, it should be noted that porting the function to C++ took virtually no extra work.</p><p>By necessity, I’ve learned a lot more about high-performance R since writing those two posts (part of this is by trying to make <a href="https://github.com/tonyfischetti/assertr">my own R package</a> as performant as possible). In particular, I did the Rcpp version all wrong, and I’d like to rectify that in this post. I also compare the running times of approaches that use both parallelism and Rcpp.</p><p>###Lesson 1: use Rcpp correctly</p><p>The biggest lesson I learned, is that it isn’t sufficient to just replace inner loops with C++ code; the repeated transferring of data from R to C++ comes with a lot of overhead. By actually coding the loop in C++, the speedups to be had are often astounding.</p><p>In this example, the pure R version, that takes a matrix of longitude&#x2F;latitude pairs and computed the mean distance between all combinations, looks like this…</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br></pre></td><td class="code"><pre><span class="line">just.R <span class="operator">&lt;-</span> <span class="keyword">function</span><span class="punctuation">(</span>dframe<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">  numrows <span class="operator">&lt;-</span> nrow<span class="punctuation">(</span>dframe<span class="punctuation">)</span></span><br><span class="line">  combns <span class="operator">&lt;-</span> combn<span class="punctuation">(</span><span class="number">1</span><span class="operator">:</span>nrow<span class="punctuation">(</span>dframe<span class="punctuation">)</span><span class="punctuation">,</span> <span class="number">2</span><span class="punctuation">)</span></span><br><span class="line">  numcombs <span class="operator">&lt;-</span> ncol<span class="punctuation">(</span>combns<span class="punctuation">)</span></span><br><span class="line">  combns <span class="operator">%&gt;%</span></span><br><span class="line">    <span class="punctuation">&#123;</span> mapply<span class="punctuation">(</span><span class="keyword">function</span><span class="punctuation">(</span>x<span class="punctuation">,</span>y<span class="punctuation">)</span><span class="punctuation">&#123;</span></span><br><span class="line">        haversine<span class="punctuation">(</span>dframe<span class="punctuation">[</span>x<span class="punctuation">,</span><span class="number">1</span><span class="punctuation">]</span><span class="punctuation">,</span> dframe<span class="punctuation">[</span>x<span class="punctuation">,</span><span class="number">2</span><span class="punctuation">]</span><span class="punctuation">,</span></span><br><span class="line">                  dframe<span class="punctuation">[</span>y<span class="punctuation">,</span><span class="number">1</span><span class="punctuation">]</span><span class="punctuation">,</span> dframe<span class="punctuation">[</span>y<span class="punctuation">,</span><span class="number">2</span><span class="punctuation">]</span><span class="punctuation">)</span> <span class="punctuation">&#125;</span><span class="punctuation">,</span></span><br><span class="line">        .<span class="punctuation">[</span><span class="number">1</span><span class="punctuation">,</span><span class="punctuation">]</span><span class="punctuation">,</span> .<span class="punctuation">[</span><span class="number">2</span><span class="punctuation">,</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">&#125;</span> <span class="operator">%&gt;%</span></span><br><span class="line">    <span class="built_in">sum</span> <span class="operator">%&gt;%</span></span><br><span class="line">    <span class="punctuation">(</span><span class="keyword">function</span><span class="punctuation">(</span>x<span class="punctuation">)</span> x<span class="operator">/</span><span class="punctuation">(</span>numrows<span class="punctuation">\</span><span class="operator">*</span><span class="punctuation">(</span>numrows<span class="operator">-</span><span class="number">1</span><span class="punctuation">)</span><span class="operator">/</span><span class="number">2</span><span class="punctuation">)</span><span class="punctuation">)</span></span><br><span class="line"><span class="punctuation">&#125;</span></span><br></pre></td></tr></table></figure><p>The naīve usage of Rcpp (and the one I used in the second blog post on this topic) simply replaces the call to “haversine” with a call to “haversine_cpp”, which is written in C++. Again, a small speedup was obtained, but it was functionally trivial.</p><p>The better solution is to completely replace the combinations&#x2F;“mapply” construct with a C++ version. Mine looks like this…</p><figure class="highlight c"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br><span class="line">11</span><br><span class="line">12</span><br></pre></td><td class="code"><pre><span class="line"><span class="type">double</span> <span class="title function_">all_cpp</span><span class="params">(Rcpp::NumericMatrix&amp; mat)</span>&#123;</span><br><span class="line">    <span class="type">int</span> nrow = mat.nrow();</span><br><span class="line">    <span class="type">int</span> numcomps = nrow*(nrow<span class="number">-1</span>)/<span class="number">2</span>;</span><br><span class="line">    <span class="type">double</span> running_sum = <span class="number">0</span>;</span><br><span class="line">    <span class="keyword">for</span>( <span class="type">int</span> i = <span class="number">0</span>; i &lt; nrow; i++ )&#123;</span><br><span class="line">        <span class="keyword">for</span>( <span class="type">int</span> j = i+<span class="number">1</span>; j &lt; nrow; j++)&#123;</span><br><span class="line">            running_sum += haversine_cpp(mat(i,<span class="number">0</span>), mat(i,<span class="number">1</span>),</span><br><span class="line">                                         mat(j,<span class="number">0</span>), mat(j,<span class="number">1</span>));</span><br><span class="line">        &#125;</span><br><span class="line">    &#125;</span><br><span class="line">    <span class="keyword">return</span> running_sum / numcomps;</span><br><span class="line">&#125;</span><br></pre></td></tr></table></figure><p>The difference is incredible…</p><figure class="highlight r"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br><span class="line">8</span><br><span class="line">9</span><br><span class="line">10</span><br></pre></td><td class="code"><pre><span class="line">res <span class="operator">&lt;-</span> benchmark<span class="punctuation">(</span>R.calling.cpp.naive<span class="punctuation">(</span>air.locs<span class="punctuation">[</span><span class="punctuation">,</span><span class="operator">-</span><span class="number">1</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">,</span></span><br><span class="line">                 just.R<span class="punctuation">(</span>air.locs<span class="punctuation">[</span><span class="punctuation">,</span><span class="operator">-</span><span class="number">1</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">,</span></span><br><span class="line">                 all_cpp<span class="punctuation">(</span>as.matrix<span class="punctuation">(</span>air.locs<span class="punctuation">[</span><span class="punctuation">,</span><span class="operator">-</span><span class="number">1</span><span class="punctuation">]</span><span class="punctuation">)</span><span class="punctuation">)</span><span class="punctuation">,</span></span><br><span class="line">                 columns <span class="operator">=</span> <span class="built_in">c</span><span class="punctuation">(</span><span class="string">&quot;test&quot;</span><span class="punctuation">,</span> <span class="string">&quot;replications&quot;</span><span class="punctuation">,</span> <span class="string">&quot;elapsed&quot;</span><span class="punctuation">,</span> <span class="string">&quot;relative&quot;</span><span class="punctuation">)</span><span class="punctuation">,</span></span><br><span class="line">                                  order<span class="operator">=</span><span class="string">&quot;relative&quot;</span><span class="punctuation">,</span> replications<span class="operator">=</span><span class="number">10</span><span class="punctuation">)</span></span><br><span class="line">res</span><br><span class="line"><span class="comment">#                                   test replications elapsed relative</span></span><br><span class="line"><span class="comment"># 3  all_cpp(as.matrix(air.locs[, -1]))           10   0.021    1.000</span></span><br><span class="line"><span class="comment"># 1 R.calling.cpp.naive(air.locs[, -1])           10  14.419  686.619</span></span><br><span class="line"><span class="comment"># 2              just.R(air.locs[, -1])           10  15.068  717.524</span></span><br></pre></td></tr></table></figure><p>The properly written solution in Rcpp is 718 times faster than the native R version and 687 times faster than the naive Rcpp solution (using 200 airports).</p><h3 id="Lesson-2-Use-mclapply-x2F-mcmapply"><a href="#Lesson-2-Use-mclapply-x2F-mcmapply" class="headerlink" title="Lesson 2: Use mclapply&#x2F;mcmapply"></a>Lesson 2: Use mclapply&#x2F;mcmapply</h3><p>In the first blog post, I used a messy solution that explicitly called two parallel processes. I’ve learned that using mclapply&#x2F;mcmapply is a lot cleaner and easier to intregrate into idiomatic&#x2F;functional R routines. In order to parallelize the native R version above, all I had to do is replace the call to “mapply” to “mcmapply” and set the number of cores (now I have a 4-core machine!).</p><p>Here are the benchmarks:</p><figure class="highlight plaintext"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br></pre></td><td class="code"><pre><span class="line">                                           test replications elapsed relative</span><br><span class="line">2 R.calling.cpp.naive.parallel(air.locs[, -1])           10  10.433    1.000</span><br><span class="line">4              just.R.parallel(air.locs[, -1])           10  11.809    1.132</span><br><span class="line">1          R.calling.cpp.naive(air.locs[, -1])           10  15.855    1.520</span><br><span class="line">3                       just.R(air.locs[, -1])           10  17.221    1.651</span><br></pre></td></tr></table></figure><h3 id="Lesson-3-Smelly-combinations-of-Rcpp-and-parallelism-are-sometimes-counterproductive"><a href="#Lesson-3-Smelly-combinations-of-Rcpp-and-parallelism-are-sometimes-counterproductive" class="headerlink" title="Lesson 3: Smelly combinations of Rcpp and parallelism are sometimes counterproductive"></a>Lesson 3: Smelly combinations of Rcpp and parallelism are sometimes counterproductive</h3><p>Because of the nature of the problem and the way I chose to solve it, the solution that uses Rcpp correctly is not easily parallelize-able. I wrote some *extremely* smelly code that used explicit parallelism to use the proper Rcpp solution in a parallel fashion; the results were interesting:</p><figure class="highlight plaintext"><table><tr><td class="gutter"><pre><span class="line">1</span><br><span class="line">2</span><br><span class="line">3</span><br><span class="line">4</span><br><span class="line">5</span><br><span class="line">6</span><br><span class="line">7</span><br></pre></td><td class="code"><pre><span class="line">                                          test replications elapsed relative</span><br><span class="line">5           all_cpp(as.matrix(air.locs[, -1]))           10   0.023    1.000</span><br><span class="line">4              just.R.parallel(air.locs[, -1])           10  11.515  500.652</span><br><span class="line">6             all.cpp.parallel(air.locs[, -1])           10  14.027  609.870</span><br><span class="line">2 R.calling.cpp.naive.parallel(air.locs[, -1])           10  17.580  764.348</span><br><span class="line">1          R.calling.cpp.naive(air.locs[, -1])           10  21.215  922.391</span><br><span class="line">3                       just.R(air.locs[, -1])           10  32.907 1430.739</span><br></pre></td></tr></table></figure><p>The parallelized proper Rcpp solution (all.cpp.parallel) was outperformed by the parallelized native R version. Further the parallelized native R version was much easier to write and was idiomatic R.</p><h3 id="How-does-it-scale"><a href="#How-does-it-scale" class="headerlink" title="How does it scale?"></a>How does it scale?</h3><p><a href="/images/hpc-methods.png"><img src="/images/hpc-methods.png" alt="Comparing performance of different HP methods"></a></p><p>Two quick things…</p><ul><li><p>The “all_cpp” solution doesn’t appear to exhibit polynomial growth; it does, it’s just so much faster than the rest that it looks completely flat</p></li><li><p>It’s hard to tell, but that’s “just.R.parallel” that is tied with “R.calling.cpp.naive.parallel”</p></li></ul><p><strong>Too long, didn’t read:</strong> If you know C++, try using Rcpp (but correctly). If you don’t, try multicore versions of lapply and mapply, if applicable, for great good. If it’s fast enough, leave well enough alone.</p><p>PS: I way overstated how “intractable” this problem is. According to my curve fitting, the vanilla R solution would take somewhere between 2.5 and 3.5 hours. The fastest version of these methods, the non-parallelized proper Rcpp one, took 9 seconds to run. In case you were wondering, the answer is 1,869.7 km (1,161 miles). The geometric mean might have been more meaningful in this case, though.</p>]]></content>
    
    
    <summary type="html">&lt;p&gt;On this blog, I’ve had a long running investigation&amp;#x2F;demonstration of how to make a “embarrassingly-parallel” but computationally intractable (on commodity hardware, at least) R problem more performant by using parallel computation and Rcpp.&lt;/p&gt;</summary>
    
    
    
    <category term="R" scheme="https://onthelambda.com/categories/R/"/>
    
    
    <category term="high performance computing" scheme="https://onthelambda.com/tags/high-performance-computing/"/>
    
    <category term="R" scheme="https://onthelambda.com/tags/R/"/>
    
    <category term="magrittr" scheme="https://onthelambda.com/tags/magrittr/"/>
    
  </entry>
  
</feed>
