<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>davidcel.is</title>
  <subtitle>Articles</subtitle>
  <author>
    <name>David Celis</name>
    <email>me@davidcel.is</email>
    <uri>https://davidcel.is/</uri>
  </author>
  <icon>https://davidcel.is/assets/me-46c38f84.jpg</icon>
  <logo>https://davidcel.is/assets/me-46c38f84.jpg</logo>
  <link rel="alternate" type="text/html" href="https://davidcel.is/articles"/>
  <link rel="self" type="application/atom+xml" href="https://davidcel.is/feeds/articles"/>
  <id>https://davidcel.is/feeds/articles</id>
  <updated>2025-12-19T20:26:36Z</updated>
  <rights>© 2026 David Celis</rights>
  <entry>
    <id>https://davidcel.is/posts/2002113355361636900</id>
    <title>Writing Code Is Fun</title>
    <content type="html" xml:lang="en">
      <![CDATA[<img src="https://cdn.davidcel.is/blog/writing-code-is-fun.png" alt="A snippet of Ruby code that powers my blog">
<p>I became a software engineer because writing code is fun. Thinking through hard problems, designing elegant solutions, seeing the things you’ve built working for the first time… these moments are all deeply satisfying, so why in the world would I ever surrender them to AI? <!--more--></p>
<p>I know the arguments for using AI to write code; I hear them constantly from all levels of the tech industry. I’m told that it’ll make me more productive. I could use Copilot to write boilerplate or unit tests for me so I can focus on more creative work. I could use Claude Code as a sounding board or planning tool for complex features, refactors, or other projects. I could run and coordinate multiple concurrent agents and have them write and maintain an entire application for me from scratch. I could do all of these things, but I won’t. Writing code is just too much fun!</p>
<p>When people use generative AI to produce images, videos, music, etc., they like to style themselves as “AI artists”. But they’re not artists. Even if you decide to be generous and call the final piece “art”, they’re <em>commissioning</em> that art, no matter how detailed or iterative they are with their prompts. They are, at best, art directors.</p>
<p>Likewise, the more you rely on AI to produce code for you, the less of a software engineer you become. Instead of spending your time solving problems and writing code, you spend your time reviewing AI-generated code. Maybe you’re relying on AI for code review, too, in which case you’re mostly just writing requirements, determining how best to format them into a prompt, and maybe managing AI agents. When someone’s primary job is to figure out and write requirements or manage the entities who are actually producing the code, we don’t usually call that person a software engineer. We call them a product or project manager. This isn’t to say that it’s bad to be a product or project manager! We sometimes need product and project managers just like the art world sometimes needs people to commission and direct art.</p>
<p>I’ve reluctantly tried various AI coding tools over the last few years and, while they <em>have</em> become more impressive, they remain deeply unfun. Writing code is, of course, only one duty of being a software engineer, but where code <em>is</em> concerned, there are two primary responsibilities: writing it and reviewing it. For people who enjoy being a software engineer, I think it’s safe to say that most of them would agree that writing the code is more fun than reviewing it, so I ask again: why cede the fun parts of your job to AI? Just to chase the increasingly demanding productivity requirements of the ruling class? Just to produce a bit more in a shorter amount of time?</p>
<p>The thing is, in all my attempts to use AI coding tools, they’ve never actually enabled me to move faster. They initially <em>felt</em> faster because the tangible output came more quickly but, in almost every case, I realized that I spent at least the same amount of time that I would have spent if I had just written the code myself (even just for boilerplate). AI produced the code more quickly than I would have, but between writing prompts, reviewing/scrutinizing code, and tweaking follow-up prompts to fix issues, I saved no time at all. I traded happiness for the illusion of speed (and <a href="https://www.coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report" target="_blank" rel="nofollow noopener noreferrer">more problems</a>). I could surely spend a lot of time learning to use these AI tools more effectively, but why do something that’s worse <em>and</em> less enjoyable? I’d rather spend that time learning other things.</p>
<p>Writing code is fun and helps me improve my skills. Solving problems keeps my mind sharp and helps me learn even more. Using AI doesn’t help me do any of this; it only sucks the joy out of my work.</p>]]>
    </content>
    <published>2025-12-19T20:26:36Z</published>
    <updated>2025-12-19T20:26:36Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/writing-code-is-fun"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/1969109987865000737</id>
    <title>The DHH Problem</title>
    <content type="html" xml:lang="en">
      <![CDATA[<figure>
  <img src="https://cdn.davidcel.is/blog/the-dhh-problem.webp" alt="The title card of Tom Stuart’s presentation, The DHH Problem.">
  <figcaption>The title card of Tom Stuart’s presentation, The DHH Problem.</figcaption>
</figure>
<p>Ages ago, when I was still a student, I taught myself Ruby on Rails for my senior thesis and fell in love. Fifteen years later, and I’ve used Rails at every job I’ve ever held in the tech industry. Fifteen years, and I still love Rails! But there’s something rotten at its core, and we share a name. <!--more--></p>
<p>Back in 2014, Tom Stuart delivered a pithy yet salient lightning talk at the Scottish Ruby Conference titled <a href="https://tomstu.art/the-dhh-problem" target="_blank" rel="nofollow noopener noreferrer">The DHH Problem</a>, in which he succinctly describes the character of David Heinemeier Hansson, the creator of Rails. I recommend watching it because, in just three and a half minutes, he effectively provides context on who DHH was, who he still is, and why his future shift to the right would make so much sense and be so unsurprising. But, if you’d rather not, I’ll include a transcript below:</p>
<blockquote>
<p>DHH is an intelligent and successful person. DHH invented Rails, and now I can get paid for writing Ruby, even though I enjoy it. So: thanks, DHH (sincerely). But, there’s a problem. DHH is “Ruby Famous”, which means that DHH is extremely visible (some might say disproportionately visible) in the Ruby community. Why is that a problem? Because DHH is the Fox News of Ruby.</p>
<p>He’s noisy, he’s reactionary, he’s anti-intellectual, he’s very sure that he is right, and he enjoys being rude. <em>[Tom cuts to a photo of DHH delivering a conference talk in 2006, standing in front of a slide that just says “Fuck You”]</em> That was eight years ago, but things haven’t changed much. Just like Fox News, DHH appeals to “common sense” and makes a show of being “fair and balanced” but, in reality, his arguments use aggressive rhetoric and rely on a fixed viewpoint.</p>
<p>To paint a topical example, is TDD hard in Rails apps because TDD is dead, or because Rails makes TDD hard? Is TDD not worth the effort because TDD is dead, or because the complexity of software can be managed more effectively if you only work on one product for which you control the requirements? If we only listen to DHH, then we’ll never know, because DHH is just one person and he only has DHH’s experiences.</p>
<p>All I’m saying is: the Ruby community is large and diverse and thoughtful, and that is why I love it. Please listen to DHH: his experiences are valuable. But DHH does not speak for me, and he probably doesn’t speak for you. My personal preference is for a bit less of this <em>[Tom cuts back to the photo of DHH’s “Fuck You” slide]</em> and a bit more of this <em>[he then cuts to a different man delivering a conference talk, standing in front of a slide that says “What others do may be the stimulus of our feelings, but never the cause.”]</em></p>
<p>Please give DHH’s opinions the weight they deserve: they are what one man thinks. And if you disagree with him, please speak up—write a blog post, give a talk, create a web framework—so that we can all learn what the world looks like to you.</p>
</blockquote>
<p>That was eleven years ago, but things haven’t changed much. DHH is still noisy, still reactionary, still anti-intellectual, still very sure that he is right, and still enjoys being rude. You need only look at his <a href="https://world.hey.com/dhh" target="_blank" rel="nofollow noopener noreferrer">personal blog</a> to witness this yourself, but doing so isn’t for the faint of heart; this iteration of DHH’s blog on Hey World started a mere four years ago but, in that time, he’s written 476 posts on a variety of topics. That’s too many posts for a deep dive, but the broad trends show a reactionary who has moved, at times steadily and at times quickly, to the far right.</p>
<p>Early on, DHH comes off as a man who held a <a href="https://world.hey.com/dhh/mosaics-of-positions-ae6d4d9e" target="_blank" rel="nofollow noopener noreferrer">variety of political beliefs</a>. He would write about <a href="https://world.hey.com/dhh/you-gotta-read-less-is-more-88a4f37f" target="_blank" rel="nofollow noopener noreferrer">the dire state of our climate</a> and then, a couple weeks later, drop hints of an increasingly present narrative around <a href="https://world.hey.com/dhh/you-gotta-read-less-is-more-88a4f37f" target="_blank" rel="nofollow noopener noreferrer">cancel culture</a>. Early posts on this blog tended to follow a theme of his “natural political stance” ostensibly being left-leaning, while being open to right-wing views as a way to <a href="https://world.hey.com/dhh/books-that-bust-bubbles-35c46be2" target="_blank" rel="nofollow noopener noreferrer">“find common ground”</a>. I’ll be honest and say that I’m not sure I believe that he was ever truly left-leaning. My own observations show someone who is fairly centrist or even slightly right of center, which tracks with what standard Liberalism has become in America.</p>
<p>Over the next few years, we’d see posts from DHH covering a range of far-right talking points: <a href="https://world.hey.com/dhh/are-we-past-peak-woke-c313b7d1" target="_blank" rel="nofollow noopener noreferrer">anti-wokeness</a>, <a href="https://world.hey.com/dhh/the-waning-days-of-dei-s-dominance-9a5b656c" target="_blank" rel="nofollow noopener noreferrer">anti-DEI</a>, <a href="https://world.hey.com/dhh/where-at-least-i-know-i-m-free-da04f873" target="_blank" rel="nofollow noopener noreferrer">free speech absolutism</a>, support of figures who range from right-leaning (like <a href="https://world.hey.com/dhh/everything-popular-is-problematic-eb039e2b" target="_blank" rel="nofollow noopener noreferrer">Joe Rogan</a>) to far-right (like <a href="https://world.hey.com/dhh/the-faith-of-andrew-tate-a8a4d448" target="_blank" rel="nofollow noopener noreferrer">Andrew Tate</a> or <a href="https://world.hey.com/dhh/the-parental-dead-end-of-consent-morality-e4e8a8ee" target="_blank" rel="nofollow noopener noreferrer">Jordan Peterson</a>), to <a href="https://world.hey.com/dhh/bad-therapy-08849dc9" target="_blank" rel="nofollow noopener noreferrer">anti-trans</a> <a href="https://world.hey.com/dhh/gender-and-sexuality-alliances-in-primary-school-at-cis-97f66c06" target="_blank" rel="nofollow noopener noreferrer">fearmongering</a>, all the way to <a href="https://world.hey.com/dhh/national-pride-f7aa1e92" target="_blank" rel="nofollow noopener noreferrer">outright nationalism</a> (this post may seem innocuous on its face, but hints at <a href="https://world.hey.com/dhh/words-are-not-violence-c751f14f" target="_blank" rel="nofollow noopener noreferrer">his inability to recognize hate speech as violence</a> and includes the common dog whistle of countries having a “strong national identity”). He <a href="https://world.hey.com/dhh/the-endangered-state-of-normality-d632a7fe" target="_blank" rel="nofollow noopener noreferrer">rails against finding beauty in diversity rather than what’s “normal”</a> or in <a href="https://world.hey.com/dhh/the-beauty-of-ideals-b3dccf72" target="_blank" rel="nofollow noopener noreferrer">people of all sizes</a>, then <a href="https://world.hey.com/dhh/you-expect-principles-but-should-wish-for-none-531988ec" target="_blank" rel="nofollow noopener noreferrer">cheers</a> when companies cave to the conservative backlash against DEI and stop celebrating that beauty. But what I find most indicative of DHH’s shift (whether towards the right, or towards honesty that he’s always been on the right) comes from a quote in one of his earliest posts on this blog, <a href="https://world.hey.com/dhh/legacy-without-nostalgia-b19708c9" target="_blank" rel="nofollow noopener noreferrer">Legacy without nostalgia</a>:</p>
<blockquote>
<p>I’m just not a nostalgic person. In fact, I’m deeply skeptical of nostalgia. It too often feels like a trap to romanticize the past at the expense of the future. If you fall in love with who you once were, it’s too easy to forget to keep going. I want to keep going.</p>
</blockquote>
<p>This is <em>deeply</em> ironic given just how much DHH thirsts for nostalgia in his later posts. He frequently yearns for the world of the past, especially <a href="https://world.hey.com/dhh/the-80s-are-still-alive-in-denmark-54e7a404" target="_blank" rel="nofollow noopener noreferrer">the 80s</a> (but not the 90s, which he <a href="https://world.hey.com/dhh/it-s-beginning-to-feel-like-the-80s-in-america-again-68c2708e" target="_blank" rel="nofollow noopener noreferrer">hates</a>), before all of this pesky talk about DEI, before cancel culture, <a href="https://world.hey.com/dhh/pessimism-is-on-the-retreat-e67dbd7e" target="_blank" rel="nofollow noopener noreferrer">before “pessimism”</a>. But no post encompasses this hypocritical desire for nostalgia more than his essay from a few days ago, <a href="https://world.hey.com/dhh/as-i-remember-london-e7d38e64" target="_blank" rel="nofollow noopener noreferrer">As I remember London</a>, a post which is quite simply fascism on full display.</p>
<h3>The fascism of DHH</h3>
<p>Some people <em>really</em> don’t like it when you throw the word “fascism” around or call people a fascist, and there are many cases where its usage is undue. This is not one of those cases. There is a <a href="https://en.wikipedia.org/wiki/Definitions_of_fascism" target="_blank" rel="nofollow noopener noreferrer">long history of attempts to coalesce definitions of fascism</a> (many of which delve into extreme nationalism), but I like Roger Griffin’s definition, which he argues can be condensed into one sentence:</p>
<blockquote>
<p>Fascism is a political ideology whose mythic core in its various permutations is a palingenetic form of populist ultra-nationalism.</p>
</blockquote>
<p>That’s admittedly a dense sentence, so why don’t we unpack it a little bit, starting with <a href="https://en.wikipedia.org/wiki/Palingenesis" target="_blank" rel="nofollow noopener noreferrer">Palingenesis</a> being the concept of a rebirth or recreation. Fascism has an obsession with a cult of national purity and the idea of a mythic past to which we must return through some form of national rebirth (<a href="https://en.wikipedia.org/wiki/Palingenetic_ultranationalism" target="_blank" rel="nofollow noopener noreferrer">palingenetic ultranationalism</a>). Fascists will attempt to paint their nation as humiliated and threatened on the global stage and that it must be reborn through a populist, revolutionary, nationalist movement. This is a <em>style</em> of politics, not a form of government; a process, not an end result (even if the end results tend to be similar); and “As I remember London” fits this bill.</p>
<p>This post echoes that myth of a golden past, of London in the late 90s, and <em>very quickly</em> claims that London has strayed from it because “it’s no longer full of native Brits” in his view, citing statistical changes in London’s <em>ethnic groups</em>. This is an obvious and purposeful feint that equates “native Brits” with “white Brits” and falls into the fascist rhetoric of their national purity suffering as a result. DHH’s entire post is a subtextual frame for the hope that Britain can experience a national rebirth and return to a time of white purity.</p>
<p>After citing those statistics, DHH lionizes <a href="https://en.wikipedia.org/wiki/Tommy_Robinson" target="_blank" rel="nofollow noopener noreferrer">Tommy Robinson</a>, calling his recent march “heartwarming”. Robinson is one of Britain’s most prominent far-right activists (with a string of <a href="https://en.m.wikipedia.org/wiki/Tommy_Robinson#Criminal_offences" target="_blank" rel="nofollow noopener noreferrer">multiple arrests and convictions</a>) and has been involved in several hate-based organizations like the <a href="https://en.m.wikipedia.org/wiki/English_Defence_League" target="_blank" rel="nofollow noopener noreferrer">English Defence League</a>, <a href="https://en.m.wikipedia.org/wiki/Pegida_UK" target="_blank" rel="nofollow noopener noreferrer">Pegida UK</a>, and the <a href="https://en.wikipedia.org/wiki/British_National_Party" target="_blank" rel="nofollow noopener noreferrer">British Nationalist Party</a>. He cites <a href="https://www.bbc.com/news/uk-politics-66960890" target="_blank" rel="nofollow noopener noreferrer">misleading reports</a> of “Pakistani Rape Gangs” despite later <a href="https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/944206/Group-based_CSE_Paper.pdf" target="_blank" rel="nofollow noopener noreferrer">research</a> making it clear that the majority of group-based child sex offenders are white. DHH’s quick reinforcement of his support for <a href="https://en.wikipedia.org/wiki/Graham_Linehan" target="_blank" rel="nofollow noopener noreferrer">Graham Linehan</a> (the notably anti-trans Irish comedian) is disgusting icing on an already disgusting cake.</p>
<p>DHH takes this growing multiculturalism in Britain and uses it to stoke fear of the same thing happening in Denmark, and Steve Klabnik was quick to <a href="https://bsky.app/profile/steveklabnik.com/post/3lyvyzftf5c2i" target="_blank" rel="nofollow noopener noreferrer">notice</a> the ridiculousness of saying that there’s “absolutely nothing racist or xenophobic in saying that Denmark is primarily a country for the Danes, Britain primarily a united kingdom for the Brits, and Japan primarily a set of islands for the Japanese” and, worse, how that verbiage disturbingly echoes the sentiments of the <a href="https://www.wisconsinhistory.org/Records/Image/IM58857" target="_blank" rel="nofollow noopener noreferrer">Ku Klux Kreed</a> (America for Americans)</p>
<h3>Something needs to change</h3>
<p>This post was my breaking point. This post is why I’m here right now, writing something much longer than the social media quips I’ve written over the years in response to his ongoing descent into far-right white nationalism and fascist rhetoric (again, <a href="https://theconversation.com/is-every-nationalist-a-potential-fascist-a-historian-weighs-in-256826" target="_blank" rel="nofollow noopener noreferrer">nationalism is the bedrock of fascism</a>). Every time DHH spews a new piece of hatred, he <a href="https://tekin.co.uk/2025/09/the-ruby-community-has-a-dhh-problem" target="_blank" rel="nofollow noopener noreferrer">alienates</a> people that slowly decide to leave the Rails or greater Ruby communities. But I’ve noticed another common response over the years—often from multiple people—every time there’s a new controversy: “Wow, I had no idea DHH went fully mask off.” There’s a clear lack of awareness. Or, worse: people simply forget.</p>
<p>“As I remember London” should be a wake-up call for everyone in the Ruby and Rails communities. This is not a man who wants to keep going. This is a man who romanticizes a past that was predominantly good for white men. This is a man who has spent years railing against diversity, equity and inclusion and who spreads anti-trans rhetoric. This is a man who is deeply afraid of immigration changing countries’ cultures and “national identities”, despite this kind of change being a constant for the whole of human history. This is a man who is a white nationalist. And he is in sole control of Rails, this framework we all love.</p>
<p>There’s a growing discontent that I see on social media over leadership in Rails, most of which comes down to DHH’s continued involvement, waning trust in the  organizations who continue to platform him (like Ruby Central), and the problematic reality of how much Ruby and Rails is beholden to Shopify money. But what’s the answer? What can we reasonably change? I don’t know! It doesn’t help that, as I was nearing the end of this post, yet another Ruby Drama unfolded live on social media. This latest drama (in a community that has a long history of drama the likes of which I’ve not seen in any other programming community) is that, over the course of the last week and a half, <a href="https://joel.drapper.me/p/rubygems-takeover/" target="_blank" rel="nofollow noopener noreferrer">Ruby Central forcibly took control of the rubygems.org codebase on GitHub</a> and revoked access from all of its maintainers in a massive shift of governance that came without advanced warning or communication of any kind. They <a href="https://rubycentral.org/news/strengthening-the-stewardship-of-rubygems-and-bundler/" target="_blank" rel="nofollow noopener noreferrer">responded</a> several hours later with some very hand-wavey corpospeak describing their motive of “a fiduciary duty to safeguard the supply chain and protect the long-term stability of the ecosystem,” but Joel Drapper’s summary makes it clear that this came down to Shopify taking advantage of funding being pulled from Ruby Central due to their continued relationship with DHH. The whole situation has reeked of incompetence and I’ve lost any trust I had left in the organization, which had admittedly become very little. I’m tired of a lack of transparency from our community’s unelected leaders and board members. I’m tired of the vagueposting that makes it impossible to decipher what events like these are truly about. I’m just so tired in general, and now I’m thinking not only of Rails’ governance, but governance in the greater Ruby community as a whole.</p>
<p>I appreciate a lot of the work Ruby Central has done over the years in stewarding important projects and infrastructure that are central to Ruby and Rails, but we need something that is beholden to the community rather than a board of appointed officials who can become beholden to corporate sponsors. What I’d prefer to see is something more akin to a co-operative. What if we had a public organization with open membership (whether via dues or some other mechanism) and democratically elected leadership? I would feel much more comfortable with that kind of organization maintaining and stewarding important projects like RubyGems, Rails, Rack et al. and a democratic process would abate a lot of my negative feelings around our current state of corporate sponsorship in Rails.</p>
<p>At the same time, I recognize that this kind of change would be incredibly difficult. I don’t think there’s any chance of removing DHH from Rails, and I can’t see him stepping down of his own accord. He owns the copyrights to Rails, he’s the founder and chair of Rails’ current governing body, the <a href="https://rubyonrails.org/foundation" target="_blank" rel="nofollow noopener noreferrer">Rails Foundation</a>, and—despite his open politics—he continues to be platformed. He’s invited to speak on podcasts. He’s invited to speak or keynote at conferences. Even Ruby Central <a href="https://rubyonrails.org/2025/5/29/final-railsconf" target="_blank" rel="nofollow noopener noreferrer">invited him back for the last RailsConf</a> earlier this year after he threw a <a href="https://world.hey.com/dhh/no-railsconf-faa7935e" target="_blank" rel="nofollow noopener noreferrer">hissy fit</a> over being asked to cede the keynote stage in 2022. Rails core members and organizations like Ruby Central continue to work with and collaborate with him closely and, though many have walked away in recent years, I wish more would follow. I get that doing so is an incredibly difficult decision to make, but people have had years to distance themselves from this man. But they don’t, so Ruby and Rails continue to be directed by far-right sympathizers like DHH and <a href="https://techwontsave.us/episode/232_shopifys_right_wing_inner_circle_w_luke_lebrun__rachel_gilmore" target="_blank" rel="nofollow noopener noreferrer">Tobi Lütke</a> (and their money). And all of this is allowed because Ruby and Rails have a <a href="https://rubyonrails.org/conduct" target="_blank" rel="nofollow noopener noreferrer">code of conduct</a> that purposefully falls prey to the <a href="https://en.wikipedia.org/wiki/Paradox_of_tolerance" target="_blank" rel="nofollow noopener noreferrer">paradox of tolerance</a>.</p>
<p>But is difficulty really an excuse not to try? Maybe, just maybe, we <em>can</em> do something about this. As I mentioned before, many prominent maintainers of projects like RubyGems and Rails have already stepped away due to ideological issues with DHH and RubyCentral. What if they were be open to returning with new governance? Of course, the unfortunate nature of copyright and current ownership means we would likely be facing a number of community forks (at the very least for Rails, and quite possibly for RubyGems as well), but <a href="https://www.softsculptor.com/why-open-source-forking-is-a-hot-button-issue/" target="_blank" rel="nofollow noopener noreferrer">community forks need not be doomed to fail</a>! Starting the right conversations with the right people early on can do wonders for momentum. So, why collectively settle for this? I have to believe that we can do better than the leadership we currently have. We can do better than the Fox News of Ruby.</p>]]>
    </content>
    <published>2025-09-19T18:43:00Z</published>
    <updated>2026-08-22T19:54:30Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/the-dhh-problem"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/1738384198135409427</id>
    <title>Fraying Threads</title>
    <content type="html" xml:lang="en">
      <![CDATA[<figure>
  <img src="https://cdn.davidcel.is/blog/fraying-threads.webp" alt="The App Store listing for “Threads, an Instagram app”, showing a two star rating.">
  <figcaption>The app store listing for Threads. Photo by <a href="https://unsplash.com/@aussiedave" target="_blank" rel="nofollow noopener noreferrer">Dave Adamson</a> on <a href="https://unsplash.com/" target="_blank" rel="nofollow noopener noreferrer">Unsplash</a>.</figcaption>
</figure>
<p>At <a href="https://xoxo.zone" target="_blank" rel="nofollow noopener noreferrer">xoxo.zone</a>, we chose to block federation with Threads as soon as they’d publicized the <a href="https://threads.net" target="_blank" rel="nofollow noopener noreferrer">threads.net</a> domain they intended to use. There were a few questions but, thankfully, that decision wasn’t particularly controversial in our community. <em>Outside</em> of our community, though, I’ve seen a lot of skepticism about preemptively blocking Threads.</p>
<p>One of the first and most widely-shared takes I saw was from John Gruber of Daring Fireball, “<em><a href="https://daringfireball.net/linked/2023/06/19/not-that-kind-of-open" target="_blank" rel="nofollow noopener noreferrer">Not That Kind of ‘Open’</a></em>”, in which he linked to the <a href="https://fedipact.online" target="_blank" rel="nofollow noopener noreferrer">Anti-Meta Fedi Pact</a> and wrote… <!--more--></p>
<blockquote>
<p>The whole point of ActivityPub as an open protocol is to turn Twitter/Instagram-like social networking into something more akin to email: truly open. If Facebook were on the cusp of launching a Gmail-like email service, would you preemptively declare that your email server would block them? To me that’s what this “Anti-Meta Fedi Pact” is arguing for.</p>
<p>Maybe I’m wrong! I certainly don’t think the “let’s pledge to block Facebook before their Fediverse thing even starts” people are nuts. But to me this feels like convicting Facebook of a pre-crime. Is the goal of the Fediverse to be anti-corporate/anti-commercial, or to be pro-openness? I think openness is the answer. Others clearly disagree.</p>
</blockquote>
<p>Gruber misses the point entirely. ActivityPub isn’t some critical communication protocol akin to email and, while openness is <em>one</em> goal of ActivityPub, the freedom of server admins to choose how they federate is another. More importantly, blocking Meta isn’t about being anti-corporate/anti-commercial (though I’m sure that’s partially at play for some), and the idea that it’s akin to “convicting Facebook of a pre-crime” is a naïve conclusion to draw. Our problem with Meta isn’t that we’re worried that they’ll be bad for the fediverse; our problem is that Meta is <em>already</em> bad for the <em>entire internet and society at large</em>.</p>
<p>Meta’s <em>existing</em> atrocities form a long list, but <a href="https://mas.to/@kissane" target="_blank" rel="nofollow noopener noreferrer">Erin Kissane</a> provides <a href="https://erinkissane.com/untangling-threads" target="_blank" rel="nofollow noopener noreferrer">a thoughtful and detailed exploration of Threads and federation</a>. She approaches the argument from all angles while doing her best to check her own biases and, if you don’t understand why federating with Threads is a bad idea, you should stop here, read her article, and then come back. If you have even <em>more</em> time, you should read her entire <a href="https://erinkissane.com/meta-in-myanmar-full-series" target="_blank" rel="nofollow noopener noreferrer">Meta in Myanmar series</a>. Either way:</p>
<blockquote>
<p>Threads isn’t yet running ads-qua-ads, but it launched with a preloaded fleet of “brands” and the promise of being a <em>nice</em>, un-heated space for conversation—which to say, an explicitly brand-friendly environment. (So far, this has meant <em>no to butts</em> and <em>no to searching for long covid info</em> and <em>yes to accounts devoted to stochastic anti-LGBT terrorism for profit</em>, so perhaps that’s a useful measure of what brands consider safe and neutral.) Perhaps there’s a world in which Threads doesn’t accept ads, but I have difficulty seeing it.</p>
</blockquote>
<p>That “preloaded fleet of brands” made Threads feel particularly soulless when I first joined, and a lot of the content I get served on Threads still falls into the category of “brands and blue checks post inane, thought-leadering engagement bait.” Most of it is just tedious to read and boring to interact with, but the <em>worst</em> kind of content on Threads is <em>actually</em> sinister and (<em>surprise!</em>) Meta does practically nothing in the way of content moderation to prevent it. This brings me to a particularly sobering excerpt from Kissane’s article:</p>
<blockquote>
<p>I <em>would</em> note that most mainstream fedi servers maintain policies that at least claim to ban (open) harassment or hateful content based on gender, gender identity, race or ethnicity, or sexual orientation. On this count, I’d argue that Threads already fails the first principle of the Mastodon Server Covenant:</p>
<blockquote>
<p><strong>Active moderation against racism, sexism, homophobia and transphobia</strong> Users must have the confidence that they are joining a safe space, free from white supremacy, anti-semitism and transphobia of other platforms.</p>
</blockquote>
<p>Don’t take my word for this failure. Twenty-four civil rights, digital justice and pro-democracy organizations <a href="https://accountabletech.org/wp-content/uploads/Letter-to-Meta.pdf" target="_blank" rel="nofollow noopener noreferrer">delivered an open letter last summer</a> on Threads’ immediate content moderation…challenges:</p>
<blockquote>
<p>…we are observing neo-Nazi rhetoric, election lies, COVID and climate change denialism, and more toxicity. They posted bigoted slurs, election denial, COVID-19 conspiracies, targeted harassment of and denial of trans individuals’ existence, misogyny, and more. Much of the content remains on Threads indicating both gaps in Meta’s Terms of Service and in its enforcement, unsurprising given your long history of inadequate rules and inconsistent enforcement across other Meta properties.</p>
<p>Rather than strengthen your policies, Threads has taken actions doing the opposite, by purposefully not extending Instagram’s fact-checking program to the platform and capitulating to bad actors, and by removing a policy to warn users when they are attempting to follow a serial misinformer. Without clear guardrails against future incitement of violence, it is unclear if Meta is prepared to protect users from high-profile purveyors of election disinformation who violate the platform’s written policies.</p>
</blockquote>
</blockquote>
<p>That open letter was back in <em>July</em>, well before the ongoing humanitarian crisis and genocide in Gaza. Every time I’ve decided to check in on Threads recently, I’ve seen (in addition to everything observed above) a virulent mix of antisemitism, Islamophobia, and propaganda, all courtesy of Threads’ algorithm.</p>
<p>Like Kissane, I try to be pragmatic. I understand why people would choose to federate with Threads. I spent a lot of my life on Twitter and I still struggle with the fact that my community has scattered to the wind, with only a small percentage finding their footholds elsewhere. Some people think that Threads is best positioned to be the next large, text-based social network now that so many have fled Twitter. Personally, I don’t believe there <em>will</em> be another huge, centralized social network like Twitter (at least not any time soon).</p>
<p>For me (and, thankfully, the community I help moderate), federating with Threads isn’t an option. Meta’s long track record should make it painfully obvious that they care more about their business metrics than the community they profit from. They continue to avoid content moderation as much as possible and have proven time and time again that not only will they <em>allow</em> vile and hateful content in their networks; they’ll visibly <em>proliferate</em> it. Meta is not a good steward of their communities, and we’re under no obligation to allow them unfettered access to any of ours.</p>]]>
    </content>
    <published>2023-12-23T02:21:06Z</published>
    <updated>2023-12-23T16:11:21Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/fraying-threads"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/1610334269207479338</id>
    <title>Fifteen Years of Ru...ewriting My Website</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>A couple weeks ago, an internet friend/mutual of mine, <a href="https://steveklabnik.com/" target="_blank" rel="nofollow noopener noreferrer">Steve Klabnik</a>, published a very timely blog post titled <a href="https://steveklabnik.com/writing/ten-years-of-ru---ewriting-my-website" target="_blank" rel="nofollow noopener noreferrer"><em>Ten Years of Ru…ewriting my website</em></a> and, well, I was inspired. His post was timely for me because I was in the middle of rewriting my website too. Now that it’s live, I thought a similar post would be really fitting.</p>
<!--more-->
<p>Steve’s post is cleverly titled to play on the fact that it was published on the tenth anniversary of his foray into Rust, but he mainly talks about how he rewrote his website to learn a new framework (<a href="https://astro.build/" target="_blank" rel="nofollow noopener noreferrer">Astro</a>) and to use a language that he doesn’t frequently employ (TypeScript). Me, on the other hand? I’m mostly staying comfortable. You see, fifteen years ago, I was introduced to Ruby and Rails, and I’ve been working with them ever since. Fifteen years ago is also when I registered my first domain name (davidcelis dot com) so I could start hosting a personal site. All that to say: I’ve been rewriting this dang site for over fifteen years.</p>
<p>My first website was a WordPress blog hosted on <a href="https://www.nearlyfreespeech.net/" target="_blank" rel="nofollow noopener noreferrer">NearlyFreeSpeech.NET</a> so I could share some very amateur and very infrequent photography. During the next few years, as I got more involved in programming and was nearing my college graduation, I started to care more about my internet presence and having a more professional website I could share on my résumé. I played around with a second wordpress site centered more on myself than my photography, then a <code>CNAME</code> to my tumblr (easier to maintain), then a Jekyll blog powered by Octopress and, eventually, a simple, vanilla Jekyll blog.</p>
<p>Since then, I’ve rewritten and redesigned my Jekyll website every few years, and even switched from davidcelis dot com to my current Icelandic domain “hack”. My website has never blown me away (I’m not a designer or even a very good frontend developer), but each time did get a little better. Today marks a brand new era for my website. After fifteen years, I’m right back where I started in my programming journey: Ruby on Rails.</p>
<p>Why Rails? Why not keep it a static site? Or why not, like Steve, take this as an opportunity to stretch my legs, leave my comfort zone, and learn something new? These are reasonable questions. Static sites perform better and are generally easier to maintain and update. You don’t need to worry about a server or PaaS for running a Ruby process indefinitely. You can just build a static site, push it to GitHub, and add a <code>CNAME</code>. And this is exactly what I had already been doing for years. So why change this and replace it with a Rails app? Steve rewrote his website to learn something new and, seemingly in the process, to reinvigorate his desire to write. I want to reinvigorate my desire to write too! But the biggest reason I’m rewriting my website is that Twitter is dying.</p>
<p>I know Twitter has completely dominated news cycles and online discussions for months now, so I won’t spend a ton of time on that. But Twitter has been an incredibly important part of my life for a long time. I signed up in early 2007 and I’ve used it ever since. Sure, I’ve used it to shitpost and spew random thoughts into a void for meaningless internet points, but I’ve also used it to make friends, build communities, and even find multiple jobs that strengthened my career. I’ve maintained and rewritten and redesigned this website a lot, but Twitter has always been my home on the internet. I loved Twitter, but I don’t love it anymore, and most of my friends are leaving or have already left.</p>
<p>At this point, I’ve made <a href="https://xoxo.zone/@davidcelis" target="_blank" rel="nofollow noopener noreferrer">Mastodon</a> the place where I post what I would’ve posted on Twitter, but it still doesn’t scratch the same itch. Federation is messy, and Mastodon lacks a lot of the exploration features that Twitter had. Then again, with how addictive social media can be, maybe it’s a good thing that Mastodon can’t scratch the same itch Twitter did. Maybe that itch was bad for me. I say I loved Twitter but, like a lot of people, I also really hated Twitter. It was a terrible place with terrible people and it made me care too much about things like likes and retweets.</p>
<p>Twitter’s collapse has made me realize that what I really want is to reclaim some control. Right now, Mastodon seems to be the place where most of my network is migrating. Maybe that will stay true for a while, and maybe it will change. Maybe some other centralized platform will grab wide-spread attention. Maybe people, including my own disparate groups of online friends, will fragment and coalesce across multiple networks. The future is uncertain, but I’m trying to carve out a space for myself that I can rely on, and where friends can find me regardless of where they themselves end up: my own website. And, after all this time, this is what brings me back to Rails.</p>
<p>You see, while it can be really easy to maintain and deploy a static site, it’s not that easy to publish new content to it. I was never a frequent blogger, but when I had some motivation for longform writing, I used to write at least a couple times a year. That turned into once every couple years. Then, after my last post in 2017, I lost steam entirely. Meanwhile, on Twitter, I’d post almost every single day, often multiple times a day. With Twitter dying, and with an uncertain future for that kind of social media, I wanted to refocus on my own space and refocus on ways to reclaim my own content. But I had a few important critera:</p>
<ol>
<li>It needs to be easy for me to post. With Twitter, it was so easy to have a thought, pull out my phone, instantly share it with my friends and followers, and feel like I was connecting with people. That’s a lot harder to do with a static site. I didn’t want to be fiddling around with complicated ways to publish short notes to a GitHub repository from my phone. Sharing photos would be even worse.</li>
<li>I want to be able share all sorts of things: text, yes, but also photos, videos, links, and maybe even things like location-based check-ins.</li>
<li>I want to reclaim control of my content while still being able to reach my friends. This means continuing to share my content on platforms like Mastodon (and, for now, Twitter) but treating my own website as the final source of truth, basically joining the <a href="https://indieweb.org/" target="_blank" rel="nofollow noopener noreferrer">IndieWeb</a>.</li>
</ol>
<p>For a bit, I tried out <a href="https://micro.blog" target="_blank" rel="nofollow noopener noreferrer">micro.blog</a>, which gives you a static website but with a Twitter-like posting interface that lets you publish new posts and photos really quickly. It’s admittedly a great service that tackles those three requirements, but I’m really picky, and there were enough little ways in which it didn’t work in the way I wanted it to work, so I gave up on it pretty quickly. I also thought about keeping my Jekyll site, but it’s too hard to update a Jekyll site from a mobile device. I really didn’t want to have to do anything like fiddle with iOS shortcuts or complicated workflows that take text or photos from my phone and somehow put them in my website’s git repository. I just want to be able to pull out my phone, write something in Markdown, maybe add a photo or two (or four), hit “Send”, and have it available to read on my website and a few other websites I’m using. And, of course, I want all of it to work in exactly the way that <em>I</em> want it to work.</p>
<p>For all those reasons, I ended up doing something I hadn’t done since running through my very first Rails tutorial fifteen years ago. I opened up my terminal, typed in <code>rails new blog</code>, and hit enter. And for the first time in recent memory, I feel excited about my website again. I feel excited about writing again.</p>
<p>Here’s to another fifteen years.</p>]]>
    </content>
    <published>2023-01-03T17:56:23Z</published>
    <updated>2023-01-03T17:56:23Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/fifteen-years-of-ru-ewriting-my-website"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/817043227407287337</id>
    <title>Give It a REST: Use GraphQL for Your APIs</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>In the world of API architecture, REST has been the reigning ruler for well over a decade. Mobile devices like smartphones and tablets have made it nearly impossible to avoid REST APIs. If you use technology at all, chances are that you use software built on a REST API constantly every day. Maybe you’ve even worked on a REST API or written one yourself! Despite REST’s popularity, however, it has a few glaring flaws.</p>
<!--more-->
<h2>What is REST?</h2>
<p>In REST APIs, the server defines a specific set of resources that a client can request, and these resources are defined by unique URLs. For example, in the API for a generic microblogging platform, the URL <code>/users/1</code> may denote the first user in the system, <code>/users/1/posts</code> could return a collection of all posts that user has written, and <code>/users/1/posts/327</code> could return a single post. REST has many nuances and a well-documented specification for behavior, but URL-based resources cover the basic idea. What is ultimately important is that the <em>server</em> defines the structure of the data that the client can request.</p>
<h2>What’s wrong with REST?</h2>
<p>Imagine you work for the aforementioned Generic Microblogginator™ company as a mobile app developer. You’re given the task of writing the mobile view for a user’s profile, which needs to show information about the user and list their posts. This isn’t too difficult; just hit the <code>/users/{id}</code> endpoint to get the former, and <code>/users/{id}/posts</code> to get the latter.</p>
<p>You ship the mobile view and wait to be ✨ dazzled ✨ by all of the customer feedback and app reviews. Next week, once all of the reviews have poured in, you get a new requirement. That Other Microblogger™ shows a couple of comments on each post in their profile view. Why don’t we do that, too? Luckily, your API already has an endpoint to get a blog post’s comments: <code>/users/{id}/posts/{id}/comments</code>. You change the view to hit that endpoint for each post you show on a user’s profile page, and you’re done.</p>
<p>But now your app is slow, and this leads us to one of the major problems with REST APIs:</p>
<h2>Too many HTTP requests</h2>
<p>Let’s face it: client applications rarely stay simple. More often than not, each client has a fairly specific set of requirements that reflect what data they need from your system. If you provide only one absolute way to request data, you’ll get clients trying to ram a rhomboid peg into a diamond-shaped hole.</p>
<p>In our previous example, our mobile app will become slower and slower with each post a user writes. If a user has twenty posts listed on their profile, we’re issuing <em>22 API requests</em>. One for information on the user, one for their list of posts, and then twenty requests to get each post’s comments.</p>
<p>As you add more components to your mobile app’s interface, this problem will get worse. With each new UI component comes a new API call or a new customization to existing API endpoints. You can nest objects within each other to avoid extra API calls, but as your view becomes more complex, you’ll inevitably start nesting irrelevant data. You’ll end up with endpoints that don’t describe a single resource but, instead, a view of multiple resources. Now your API doesn’t seem so RESTful anymore.</p>
<p>Even worse, you’ll need to support any old endpoints as long as there are old versions of clients in the wild lest you risk breaking those clients. This leads to another major problem with REST:</p>
<h2>“Versioning” REST APIs is a pain</h2>
<p>The structure of responses from REST APIs is important. Clients build themselves around the knowledge that each resource has a specific structure. When Generic Microblogginator™ first released their API, this is what the response for getting a single post looked like:</p>
<pre><code class="language-json">{
  &quot;author_id&quot;: 1,
  &quot;title&quot;: &quot;Give it a REST: Use GraphQL!&quot;,
  &quot;body&quot;: &quot;In the world of API architecture, REST has been the reigning ruler for a decade or more.&quot;,
  &quot;published_at&quot;: &quot;Thu Jan 05 2017 14:45:10 GMT-0800 (PST)&quot;
}
</code></pre>
<p>After some time has passed, you decide there are a couple things you want to improve about a post’s structure in the API. Posts are about to get categories, so you’ll need to add those as a new field. You’ve also received feedback that the format for <code>published_at</code> isn’t very friendly. JavaScript clients can parse it okay, but you’d rather any tool be able to parse your timestamps easily, so you decide to change it to an ISO-8601 format. When all is said and done, you want the new structure to look like this:</p>
<pre><code class="language-json">{
  &quot;author_id&quot;: 1,
  &quot;title&quot;: &quot;Give it a REST: Use GraphQL!&quot;,
  &quot;categories&quot;: [&quot;tech&quot;, &quot;programming&quot;, &quot;graphql&quot;, &quot;rest&quot;, &quot;api&quot;],
  &quot;body&quot;: &quot;In the world of API architecture, REST has been the reigning ruler for a decade or more.&quot;,
  &quot;published_at&quot;: &quot;2017-01-05T14:45:10-08:00&quot;
}
</code></pre>
<p>Looking good! Unfortunately, one of your changes will break all of your existing clients. Every client expects <code>published_at</code> to be the less-friendly format, so that’s how they’ll try to parse it. If you want to update a field or remove a field, you have to version your API (whether it’s via the URL or an HTTP header) and try to get clients to upgrade. It’s unlikely you’d get every client to upgrade, so you have two choices:</p>
<ol>
<li>Be okay with breaking old versions of clients (including your own app)</li>
<li>Support old versions of your API until the day your company decides to announce a new chapter in their incredible journey.</li>
</ol>
<p>The easiest thing to do is simply leave your old code alone, which means piling more and more versions of your API versions on top of the old ones.</p>
<h2>A challenger approaches</h2>
<p>Enter GraphQL, a technology written by Facebook. Facebook was facing major problems with the data pipeline for their mobile applications. Their mobile apps used to be wrappers around web views and, as the mobile apps increased in complexity, they began to suffer performance problems and frequent crashes. Facebook turned to writing native applications and found themselves needing a new API to retrieve data for their native views. They evaluated REST and other options but, given problems like those described above, ultimately took the opportunity to produce something truly new.</p>
<h2>What is GraphQL?</h2>
<p>GraphQL is, as the name might suggest, a query language. It’s also perfect for APIs. It allows you to define your data using a fully-fledged type system, forming a schema that is self-documenting. It also gives clients full control over the data they request.</p>
<h2>Too many HTTP requests? How about one HTTP request?</h2>
<p>With GraphQL, clients can get all of the data they need to render a view using only one request. With our previous profile page example, a client would need to issue one request to get a user’s information, one request to get that user’s posts, and then another request for each post to get a few comments. With GraphQL, that client could get all of the above data with one request:</p>
<pre><code class="language-graphql">query {
  user(id: 1) {
    username
    fullName
    avatarUrl

    posts(first: 25) {
      title
      body
      publishedAt
      categories

      comments(last: 2) {
        author {
          username
        }
        body
      }
    }
  }
}
</code></pre>
<p>Boom! 💥 There are other benefits to this aside from the fact that we went from 22 HTTP requests to one. For instance, your User may have other information attached to it. Maybe you expose the timestamp of when a user signed up. Maybe another client doesn’t care about a post’s categories. If a client doesn’t need to query for a piece of data, <em>neither does your server</em>. So when a client saves, you can save too by simplifying your own database queries.</p>
<h2>Versioning? Just deprecate!</h2>
<p>As with (most) REST APIs, you can add fields to GraphQL types without fear. To remove functionality, GraphQL includes deprecation as a feature. Instead of fully removing a field and breaking clients, you can declare a field as deprecated and hide it from tools as it ages.</p>
<h2>Documentation: you’ll barely need to worry about it</h2>
<p>Let me be real for a second here: I can count the number of times I’ve used a well-documented API on one hand. Many times, APIs remain undocumented or poorly documented. With GraphQL, your schema is practically self-documenting. All you have to do is give your types and fields descriptions when necessary, and this happens in the code itself. Clients can issue special GraphQL queries to introspect on your application’s schema and know, in one query, all of the data they can request, what it’s called, and what it describes. Developers can also use tools that are built on this introspection like <a href="https://github.com/graphql/graphiql" target="_blank" rel="nofollow noopener noreferrer">GraphiQL</a>, which allows clients to test their queries with live syntax highlighting and error detection.</p>
<h2>Get started with GraphQL</h2>
<p>Are you sold enough to try out GraphQL? There are plenty of resources to help you get started on your journey:</p>
<ul>
<li>Check out <a href="http://graphql.org/" target="_blank" rel="nofollow noopener noreferrer">GraphQL’s official website</a> for documentation and examples</li>
<li>Play around with a working example, like the <a href="https://acnh.apps.davidcel.is/">Nook Stop API</a>, an <em>Animal Crossing: New Horizons</em> themed “GiraffeQL” API that I built for instructive purposes (you can also <a href="https://github.com/davidcelis/nook_stop_api" target="_blank" rel="nofollow noopener noreferrer">check out the code that powers it on GitHub</a>)</li>
<li>If you’re into the nitty gritty, you can read the <a href="http://facebook.github.io/graphql/" target="_blank" rel="nofollow noopener noreferrer">GraphQL Specification</a> itself.</li>
</ul>]]>
    </content>
    <published>2017-01-05T16:21:00Z</published>
    <updated>2022-12-29T23:06:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/give-it-a-rest-use-graphql-for-your-apis"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/676205209973687336</id>
    <title>Easily Publish Your Site to S3 and CloudFront</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p><em>Note: you’ll never truly be done rewriting your site so, true to fashion, my website is now a Rails app. Also, it’s way easier to put your static site behind something like Cloudflare to get automatic SSL, a CDN with caching, and plenty more. Because of that, I don’t really recommend using this post as a guide; many parts are out of date and some are no longer true. Like other defunct posts, I’ll leave it in my garden, but I’m unlikely to tend to it.</em></p>
<!--more-->
<p>Until recently, my site was hosted on GitHub Pages. It’s a static site built using Jekyll so, because Pages is free, supports custom domains, and automatically builds Jekyll sites for you, it made sense. Unfortunately, ease often comes with caveats. In the case of Pages, there were three main caveats for me:</p>
<h3>No SSL on custom domains</h3>
<p>In the age of <a href="https://www.eff.org/https-Everywhere" target="_blank" rel="nofollow noopener noreferrer">HTTPS Everywhere</a>, it’s a major drawback not to be able to support SSL on your website. Even Google has said that <a href="http://www.newsledge.com/seo-google-encryption-let-rush-https-begin-8485" target="_blank" rel="nofollow noopener noreferrer">HTTPS sites will get preferential placement in search results</a>. I know that my site is just a static application with no forms or POSTs, but even having an SSL certificate that shows this domain is really owned by me is nice for verification purposes.</p>
<h3>No custom caching headers</h3>
<p>Now, to be fair, GitHub <em>does</em> use Fastly as their CDN. However, the speed benefits are apparently somewhat lost if you use a naked subdomain instead of something like “www”. Additionally, they set their max-age value for assets to only 600 seconds. The font files on my website will never change, and the CSS (and even existing pages) very rarely change. I wanted to be able to tell web browsers to fetch these files from the browser cache for up to one year and prevent retrieving them over the network.</p>
<h3>Jekyll is forced into safe mode</h3>
<p>GitHub runs Jekyll in safe mode by default, for obvious reasons. This means that you are unable to use <a href="http://jekyllrb.com/docs/plugins/" target="_blank" rel="nofollow noopener noreferrer">Jekyll Plugins</a> on GitHub Pages. This wasn’t actually a <em>huge</em> deal for me, as I was able to get by somewhat easily without introducing custom plugins. But there’s some nice stuff out there in the Jekll ecosystem.</p>
<h2>Onward to AWS!</h2>
<p>With these issues in mind, I decided to try moving my site to AWS. I knew that this would involve hosting my built site in an Amazon S3 bucket and, in order to have effective caching and serve my site over a CDN, I’d need to put a CloudFront distribution in front of it. I didn’t think this process would be as tricky to get right, but it ended up being a bit arduous. However, by the end, I was left with mostly what I wanted:</p>
<ul>
<li>My site now runs on SSL at <a href="https://davidcel.is/">https://davidcel.is/</a></li>
<li>If a user visits my site on HTTP, they are redirected to HTTPS.</li>
<li>If a user visits my site with a www subdomain, they are redirected to a naked subdomain.</li>
<li>CSS files are versioned with a query parameter (I wanted fingerprints but couldn’t get existing third party plugins to work for my purposes).</li>
<li>After the first visit, browser caching has most pages typically loading and fully rendered in less than 100ms (blog posts with embedded media are still a bit slower, but what are you gonna do).</li>
</ul>
<p>Want the same sort of setup for your own site? Here’s how I did it:</p>
<h3>s3_website</h3>
<p><a href="https://github.com/laurilehmijoki/s3_website" target="_blank" rel="nofollow noopener noreferrer">s3_website</a> is a Ruby gem that facilitates the publishing of a static site to S3 and CloudFront. It supports Jekyll out of the box and automatically detects a site built and placed in the <code>_site</code> directory of a project. It also has a fairly sane default configuration and has decent usage documentation. I thought that my search was over when I found this project, but I hit a few snags and had to do too much reading in CloudFront’s documentation to let anybody else do the same. Anyway, start by installing s3_website and generating a configuration in your project:</p>
<pre><code class="language-sh">gem install s3_website
s3_website cfg create
</code></pre>
<p>This will give you a YAML configuration file in the root of your project. Read through it to see various configuration options you might want to set, but I’ll go ahead and give you my final configuration in a bit. In the meantime, head over to the <a href="http://aws.amazon.com" target="_blank" rel="nofollow noopener noreferrer">AWS Console</a>. You’ll need to generate an access key and secret if you haven’t already. Under the “Security &amp; Identity” group of apps, enter the “Identity &amp; Access Management” (IAM) app. Click “Users” and create one. Then, click through to your new user and attach a policy. The one you’ll want is called “AdministratorAccess”. Or you can just give it access to S3 and CloudFront. Finally, visit your new user’s Security Credentials tab and create an access key. This is the only time that the access key’s secret will be shown, so make sure to grab them both before you leave.</p>
<p>Alright. Once you’ve gotten your credentials and read through the default configuration, you can take a look at what I’ve got:</p>
<pre><code class="language-yaml">s3_id: &lt;%= ENV['DAVIDCELIS_S3_ID'] %&gt;
s3_secret: &lt;%= ENV['DAVIDCELIS_S3_SECRET'] %&gt;
s3_bucket: &lt;%= ENV['DAVIDCELIS_S3_BUCKET'] %&gt; # Your bucket's name (e.g: example.com)
s3_endpoint: &lt;%= ENV['DAVIDCELIS_S3_ENDPOINT'] %&gt; # The S3 endpoint (e.g: us-east-1)

cloudfront_distribution_id: &lt;%= ENV['DAVIDCELIS_CLOUDFRONT_DISTRIBUTION_ID'] %&gt;
cloudfront_invalidate_root: true
cloudfront_distribution_config:
  default_cache_behavior:
    min_TTL: &lt;%= 60 * 60 * 24 %&gt;
    viewer_protocol_policy: redirect-to-https
  aliases:
    quantity: 1
    items:
      CNAME: davidcel.is

index_document: index.html
error_document: 404.html

max_age:
  &quot;css/*&quot;: &lt;%= 60 * 60 * 24 * 365 %&gt;
  &quot;fonts/*&quot;: &lt;%= 60 * 60 * 24 * 365 %&gt;
  &quot;images/*&quot;: &lt;%= 60 * 60 * 24 * 365 %&gt;
  &quot;*&quot;: &lt;%= 60 * 60 * 24 %&gt;

gzip:
  - .html
  - .css
</code></pre>
<p>I placed any sensitive configuration options in environment variables so that anybody viewing my site’s <a href="https://github.com/davidcelis/davidcel.is/" target="_blank" rel="nofollow noopener noreferrer">repository on GitHub</a> could see what values I have set. <code>s3_id</code> and <code>s3_secret</code> are the Access Key/Secret that you created for your new AWS user. <code>s3_bucket</code> and <code>s3_endpoint</code> configure your bucket’s name and which of Amazon’s data centers it will be in (there’s a <a href="http://docs.aws.amazon.com/general/latest/gr/rande.html#s3_region" target="_blank" rel="nofollow noopener noreferrer">list of these endpoints</a> available in Amazon’s documentation). The bucket doesn’t have to exist yet; <code>s3_website</code> will create it for you! In fact, you can leave <code>cloudfront_distribution_id</code> out too as <code>s3_website</code> will also create your CloudFront distribution for you. In the meantime, let’s walk through the rest of these values to get a sense of what’ll be going on.</p>
<p>This <code>cloudfront_invalidate_root</code> option is necessary if you’re using pretty URLs that don’t involve an <code>index.html</code> file (e.g. https://example.com/about/ instead of https://example.com/about/index.html). This makes sure that when the, say, <code>/about/index.html</code> file changes, CloudFront invalidates the cache for the resource located at <code>/about/</code>. Handy!</p>
<p>The <code>cloudfront_distribution_config</code> section contains settings for the CloudFront distribution and its behaviors. I’ve set the minimum caching TTL to one day, and stated that HTTP requests should be redirected to HTTPS. Then, I’ve provided one CNAME alias: my domain, davidcel.is.</p>
<p>Then I tell S3 that my site’s root document (located at https://davidcel.is/) is the top-level <code>index.html</code> file, and that any errors experienced should render the top-level <code>404.html</code> file.</p>
<p>Finally, I’ve set some additional caching rules. CSS, fonts, and images will all get a max-age of a whole whopping year, so browsers are instructed to retrieve them from their local cache until that age has expired. This is great for font files and images which should never change. CSS changes do occur, so I end up versioning those files using query parameters (more on that later). Last but not least, I tell CloudFront to deliver my HTML and CSS files gzipped. Normally you’d include <code>.js</code> in that list, but the only JavaScript I use is for analytics.</p>
<p>Once you’ve got the configuration you want, you can run <code>s3_website cfg apply</code>. This command will use your AWS credentials on Amazon’s API to create your S3 bucket and CloudFront distribution with the configuration we just talked about.</p>
<p>Finally, you can build and publish your site:</p>
<pre><code class="language-sh">jekyll build
s3_website push
</code></pre>
<p>Any time you update your website and re-build it, just use <code>s3_website push</code> to update it in CloudFront. The command will calculate the diff of your new site with the old site and update only the files that need to be invalidated. If you want to invalidate everything, you can add the <code>--force</code> option.</p>
<h3>HTTPS</h3>
<p>In order to serve HTTPS from your own domain name, you’ll need to get an SSL certificate. I won’t detail the process of obtaining one, but I will recommend two sources:</p>
<ol>
<li><a href="http://startssl.com" target="_blank" rel="nofollow noopener noreferrer">StartSSL</a> offers free SSL certificates that last for a year. Sign-up is a bit time consuming and involves creating and saving an SSL certificate just to identify you and authenticate on their website, but I’d rather go through their process than pay exorbitantly for a certificate. Their certificates require a subdomain, but will also work on a naked domain. I’d suggest getting a certificate with the “www” subdomain which would, for example, work for both “www.example.com” and “example.com”.</li>
<li><a href="http://letsencrypt.org" target="_blank" rel="nofollow noopener noreferrer">Let’s Encrypt</a> has been making the rounds lately, promising free, automated SSL certificates. They’ll even be issuing wildcard certificates! The caveat here is that it’s still a beta product, and certificates currently only last for 90 days.</li>
</ol>
<p>Once you have your SSL certificate, you have to upload it to Amazon. I spent ages trying to figure out how to do this in their web console, but it’s much easier to do it from the command line. You’ll need a few files, the names of which I’ll assume:</p>
<ul>
<li>Your newly generated SSL certificate (<code>ssl.crt</code>)</li>
<li>The SSL certificate’s private key (<code>ssl.key</code>)</li>
<li>Your provider’s Certificate Authority (CA) bundle (<code>ca-bundle.pem</code>)</li>
</ul>
<p>When you’ve got these files ready, install the AWS Command Line Interface and go to town (replacing <code>example.com</code> with your domain):</p>
<pre><code class="language-sh">brew install awscli
export AWS_ACCESS_KEY_ID= # same value as s3_id from earlier
export AWS_ACCESS_KEY_SECRET= # same value as s3_secret from earlier
aws iam upload-server-certificate \
  --server-certificate-name example.com \
  --certificate-body file:///full/path/to/ssl.crt \
  --private-key file:///full/path/to/ssl.key \
  --certificate-chain file:///full/path/to/ca-bundle.pem \
  --path /cloudfront/example.com/ # /cloudfront/ is necessary for detection
</code></pre>
<p>Once the certificate is uploaded, head back to the <a href="http://aws.amazon.com" target="_blank" rel="nofollow noopener noreferrer">AWS Console</a> and navigate to the CloudFront app. Click into the distribution that was created for your site,  click “Edit”, and select “Custom SSL Certificate (stored in AWS IAM)”. Make sure the certificate in the dropdown is selected, and save.</p>
<h3>Hook up your DNS</h3>
<p>The last step is pointing your domain’s DNS at your newly created CloudFront distribution. In the console, grab your distribution’s “Domain Name” value. It should look something like “a2ebi9s43z9v9o.cloudfront.net”. You’ll want the following two records, replacing <code>example.com</code> with your own domain and <code>a2ebi9s43z9v9o.cloudfront.net</code> with your actual CloudFront distribution’s domain name:</p>
<ul>
<li>An ALIAS record for <code>example.com</code> pointing to <code>a2ebi9s43z9v9o.cloudfront.net</code>. If your DNS provider doesn’t support ALIAS records, your only choice is a CNAME. Unfortunately, CNAMEs can’t be used to point naked domains at another domain, so you’ll need something like <code>www.example.com</code> or <code>blog.example.com</code>.</li>
<li>A URL record for <code>www.example.com</code> pointing to <code>example.com</code>. This is only if you want to redirect a <code>www</code> subdomain to a naked domain. If your DNS provider doesn’t have this record type, the only solution I can think of is routing <code>www.example.com</code> to a lightweight VPS running nginx that just redirects to <code>example.com</code>.</li>
</ul>
<p>There’s also a big caveat with that URL record. These records can’t redirect over SSL, so if somebody were to visit <code>https://www.example.com</code>, it’ll hang in the web browser. I personally want the redirection and, because my site has never been on SSL before, I doubt there are any links out there in the wild to the https://www version of my website. I’m willing to risk it, but you might not be. I also tried to set up a second origin in CloudFront for the <code>www</code> subdomain that would read out of a second S3 bucket named <code>www.davidcel.is</code> (merely a mirror for my original bucket) but I couldn’t get it to quite work right. If anybody else has, please email me with steps so I can put them here! Otherwise, I’d recommend the VPS solution if you don’t mind paying or if you already have a box. That Nginx configuration could be as simple as this:</p>
<pre><code class="language-nginx">server {
  listen 80;
  server_name www.example.com;
  return 301 https://example.com$request_uri;
}

server {
  listen 443;
  server_name www.example.com;
  return 301 https://example.com$request_uri;

  ssl on;
  ssl_certificate /path/to/ssl.crt;
  ssl_certificate_key /path/to/ssl.key;
}
</code></pre>
<h3>Wait a bit…</h3>
<p>Once CloudFront finishes deploying your cached distribution and DNS kicks in, you should have a nice, fast site running low-cost on AWS. Rejoice!</p>]]>
    </content>
    <published>2015-12-14T01:01:00Z</published>
    <updated>2022-12-29T22:47:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/easily-publish-your-site-to-s3-and-cloudfront"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/560991789511607335</id>
    <title>Distance Constraints with PostgreSQL and PostGIS</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>So you’re a squirrel, and winter is coming. It’s time to start gathering nuts for… woah, wait! Come back! Remember last year? <em>Sigh</em>, of course you don’t… You squirrels have great spatial memory for their caches, but not the spots you’ve foraged in. Last year, you did a lot of searching for nuts in places you’d already cleaned out. What a waste of time! Don’t worry, though; I’m here to help! To increase your productivity, we’re going to keep a spatial database of all of your foraging spots.</p>
<!--more-->
<p>One of the most popular solutions for spatial databases is <a href="http://postgis.net/" target="_blank" rel="nofollow noopener noreferrer">PostGIS</a>. PostGIS is a <a href="https://www.postgresql.org/" target="_blank" rel="nofollow noopener noreferrer">PostgreSQL</a> extension that adds support for spatial data types and queries. This should suit our needs perfectly! Let’s start by opening up your command line and installing both PostgreSQL and PostGIS, and then initializing a new database! If you’re on macOS, you can install both using <a href="https://brew.sh/" target="_blank" rel="nofollow noopener noreferrer">Homebrew</a>; if you’re on Ubuntu, you can use <code>apt-get</code>.</p>
<pre><code class="language-sh"># Mac OS X
brew install postgresql postgis
createdb nutsy

# Ubuntu
sudo apt-get install postgresql-9.4-postgis-2.1
createdb nutsy
</code></pre>
<p>Alternatively, if you use Docker, there’s a community-maintained <a href="https://github.com/postgis/docker-postgis" target="_blank" rel="nofollow noopener noreferrer">PostGIS Docker Image</a>!</p>
<p>Anyway, now we’ve got a PostgreSQL database named <code>nutsy</code>. Normally you’d want to use some sort of migration utility to help you manage your database schema, but let’s just jump into our PostgreSQL console to play around. Just type <code>psql -d nutsy</code> in your command line. You should see something like:</p>
<pre><code>psql (14.6)
Type &quot;help&quot; for help.

nutsy=#
</code></pre>
<p>What we want to end up with is a table we can use to keep track of your past foraging spots. PostGIS gives us a few special data types to store things like points, lines, or polygons on a coordinate. The main two data types are geometries and geographies. Geometries are used to represent objects in a Euclidean coordinate system. Geographies, on the other hand, are used to represent objects in the round-earth coordinate system. While either geometries or geographies could make sense for the task at hand, our specific calculations will be made easier if we use geographies.</p>
<pre><code class="language-sql">CREATE EXTENSION &quot;postgis&quot;;

CREATE TABLE foraging_spots (
  id          SERIAL                 PRIMARY KEY,
  coordinates GEOGRAPHY(POINT, 4326) NOT NULL,
  nuts        INT
);
</code></pre>
<p>In the first line of this file, we’re enabling usage of the PostGIS extension in our database, allowing us to use its functions and types. Then we create a table named <code>foraging_spots</code> with an auto-incrementing ID, coordinates, and an optional number of nuts found in that spot. The coordinates takes advantage of the <code>GEOGRAPHY</code> type given to us by PostGIS. That type takes two parameters:</p>
<ol>
<li>A type modifier. Valid choices here are <code>POINT</code>, <code>LINESTRING</code>, <code>POLYGON</code>, <code>MULTIPOINT</code>, <code>MULTILINESTRING</code>, and <code>MULTIPOLYGON</code>. We’re representing points on a coordinate system, so we passed in <code>POINT</code>.</li>
<li>A unique Spatial Reference Identifier (SRID). 4326 is the SRID for the geographic latitude/longitude coordinate system of the Earth and is a default for many PostGIS types, so we’ll keep that as our SRID as well.</li>
</ol>
<p>Now that we have our geography column, we can start adding some locations that we’ve foraged in. I see you live in New York City’s Central Park. That wouldn’t be my first choice, but… Whatever makes you happy! Maybe you get a lot of free food from the tourists. Anyway, let’s add a few spots in the park that you’ve searched. Five should be enough:</p>
<pre><code class="language-sql">--- A little island in Turtle Pond
INSERT INTO foraging_spots (nuts, coordinates) VALUES (4, ST_GeographyFromText('POINT(-73.968504 40.779741)'));

--- Around the Obelisk
INSERT INTO foraging_spots (nuts, coordinates) VALUES (9, ST_GeographyFromText('POINT(-73.965393 40.779640)'));

--- One of the softball fields in the Great Lawn
INSERT INTO foraging_spots (nuts, coordinates) VALUES (17, ST_GeographyFromText('POINT(-73.966256 40.780602)'));

--- In the Shakespeare Garden
INSERT INTO foraging_spots (nuts, coordinates) VALUES (12, ST_GeographyFromText('POINT(-73.969861 40.779859)'));

--- In &quot;The Ramble&quot;, whatever that is!
INSERT INTO foraging_spots (nuts, coordinates) VALUES (7, ST_GeographyFromText('POINT(-73.969046 40.776146)'));
</code></pre>
<p>Okay! Now we have five foraging spots in our new table. This uses the <code>ST_GeographyFromText</code> function given to us by PostGIS. This takes a string representation of a <code>POINT</code> and uses it to construct a native Geography that can then be stored in our table. So this ends up taking the form of <code>ST_GeographyFromText('SRID=4326;POINT(longitude latitude)'))</code>. In our examples, we left out the SRID declaration because an SRID of 4326 is assumed; we don’t have to specify it.</p>
<p>Now, let’s say you go out on another run. You remember that you found a good number of nuts out in the Great Lawn, so you venture there again. It’s full of softball fields! You remember that you searched near one of them, but you don’t remember which. You want your database to be able to tell you if you start searching near the same field as last time. Your first thought is probably to put a uniqueness constraint<sup class="footnote-ref"><a href="#fn1" id="fnref1">1</a></sup> on your coordinates field:</p>
<pre><code class="language-sql">CREATE UNIQUE INDEX coordinates_gix ON foraging_spots (coordinates);
</code></pre>
<p>Now, let’s see what happens when we try to create a point with the same coordinates as last time:</p>
<pre><code class="language-sql">--- That same softball field in the Great Lawn
INSERT INTO foraging_spots (coordinates) VALUES (ST_GeographyFromText('POINT(-73.966256 40.780602)'));
--- ERROR:  duplicate key value violates unique constraint &quot;coordinates_gix&quot;
--- DETAIL:  Key (coordinates)=(0101000020E61000009A982EC4EA63444015E46723D77D52C0) already exists.
</code></pre>
<p>Our index works! … If you input the exact same coordinates. However, you are quite thorough in your searching. When you’ve added a point to your database, you know that you conducted a search that extends in a 50 meter radius around that point. Because geographic coordinates are so precise, it’s unlikely that you would try to search in the <em>exact</em> same spot. But if you try to start a new search even a single meter away from the old one…</p>
<pre><code class="language-sql">--- That same softball field in the Great Lawn, but shifted eeever so slightly...
INSERT INTO foraging_spots (coordinates) VALUES (ST_GeographyFromText('POINT(-73.966246 40.780612)'));
--- INSERT 0 1
</code></pre>
<p>Uh oh. That point is <em>very</em> close to our original point in the Great Lawn… Definitely within 50 meters of it. So how do we prevent ourselves from this sort of overlap? Unfortunately, PostGIS doesn’t seem to give us any sort of index or constraint that lets us do this easily… However, PostGIS does give us a nice function we can try to utilize: <code>ST_DWithin()</code>.</p>
<p><code>ST_DWithin()</code> takes two geographies and, in the case of our SRID, a number of meters. It will return <code>true</code> if the two geographies are within that number of meters of each other, and <code>false</code> if not. This isn’t usable from an index or uniqueness constraint, but we can use it in a different way. At a high level, we’re going to create a sort of check that raises an exception if a foraging spot is within 50 meters of an existing one. Then, we’re going to make PostgreSQL call that function every time a foraging spot has been added or modified.</p>
<pre><code class="language-sql">CREATE FUNCTION check_distance() RETURNS trigger AS $check_distance$
  BEGIN
    IF (SELECT 1 FROM foraging_spots WHERE ST_DWithin(NEW.coordinates, foraging_spots.coordinates, 50)) THEN
      RAISE EXCEPTION 'You already search this spot! Go somewhere else!';
    END IF;

    RETURN NEW;
  END;
$check_distance$ LANGUAGE plpgsql;

CREATE TRIGGER check_distance
  BEFORE INSERT OR UPDATE ON foraging_spots
  FOR EACH ROW
  EXECUTE PROCEDURE check_distance();
</code></pre>
<p>Okay, there’s a lot going on here! First, we’ve just created a function named <code>check_distance()</code> that will compare a new foraging spot with all existing foraging spots. The new foraging spot is helpfully assigned to <code>NEW</code> because we’ll be using the function as a trigger (more on that in a bit). We execute a query as a conditional check:</p>
<p><code>SELECT 1 FROM foraging_spots WHERE ST_DWithin(NEW.coordinates, foraging_spots.coordinates, 50)</code>. If the new foraging spot is found to be within 50 meters of any existing foraging spot, that statement returns <code>1</code>, passes the condition, and will raise an exception with a helpful message. If the two spots are sufficiently far apart, the condition fails and our function simply returns the new foraging spot.</p>
<p>Finally, we use the <code>CREATE TRIGGER</code> statement to tell PostgreSQL that we want this function called before any <code>INSERT</code> or <code>UPDATE</code> statement is called on the <code>foraging_spots</code> table. We want it to be called <code>FOR EACH ROW</code> as opposed to <code>FOR EACH STATEMENT</code>, since a single statement could insert multiple rows. For each row added in an <code>INSERT</code> statement or modified in an <code>UPDATE</code> statement, <code>check_distance()</code> will be called with <code>NEW</code> assigned with the new or updated values.</p>
<p>We should be good to go with this new distance check. Let’s remove that last foraging spot that we added in error, and try adding it again:</p>
<pre><code class="language-sql">--- Delete our bad foraging spot (it should just be the last one)
DELETE FROM foraging_spots WHERE id IN (SELECT id FROM foraging_spots ORDER BY id DESC LIMIT 1)
--- DELETE 1

--- Now, try to add it again
INSERT INTO foraging_spots (coordinates) VALUES (ST_GeographyFromText('POINT(-73.966246 40.780612)'));
--- ERROR:  You already search this spot! Go somewhere else!
</code></pre>
<p>Great! Now when you add a location that you’re going to start foraging in, your database will tell you if it’s within 50 meters of a spot you already searched. If it is, you should probably go somewhere else! Happy hunting!</p>
<section class="footnotes">
<ol>
<li id="fn1">
<p>It’s worth mentioning that, under normal circumstances, we would want to create a GiST (Generalized Search Tree) index. This is because PostGIS data types don’t play well with the typical B-Tree indices that PostgreSQL use by default. However, PostgreSQL doesn’t support unique GiST indices so we must stick with the default. <a href="#fnref1" class="footnote-backref">↩</a></p>
</li>
</ol>
</section>]]>
    </content>
    <published>2015-01-30T02:44:00Z</published>
    <updated>2022-12-29T22:31:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/distance-constraints-with-postgresql-and-postgis"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/330014320358327334</id>
    <title>Deploying Discourse with Capistrano</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p><em>Note: The Discourse team now provides Docker images that greatly simplify the process of deploying Discourse, so you should just use their <a href="https://github.com/discourse/discourse/blob/main/docs/INSTALL-cloud.md" target="_blank" rel="nofollow noopener noreferrer">official installation guide</a>. I’ll leave this post up for posterity but, unlike my other posts, I’m unlikely to update or maintain it.</em></p>
<!--more-->
<p>With the recent release of <a href="http://www.discourse.org/" target="_blank" rel="nofollow noopener noreferrer">Discourse</a>, an excellent piece of forum software by <a href="http://codinghorror.com/" target="_blank" rel="nofollow noopener noreferrer">Jeff Atwood</a>, <a href="http://eviltrout.com/" target="_blank" rel="nofollow noopener noreferrer">Robin Ward</a> and <a href="http://samsaffron.com/" target="_blank" rel="nofollow noopener noreferrer">Sam Saffron</a>, I was eager to get an installation up and running myself. With a simple layout it seemed like just what I needed as support and discussion forums for a side project. Discourse is still considered to be squarely in the beta phase, so I had a hard time finding any guides or tutorials that fit my needs for deploying Discourse to a VPS. I did find a <a href="https://github.com/baus/install-discourse" target="_blank" rel="nofollow noopener noreferrer">nice tutorial</a> by Christopher Baus, but I wanted to get a setup using Capistrano rather than <code>init.d</code> because that feels much more Railsy to me. Read on for how I got a Digital Ocean droplet running Discourse on thin with zero-downtime deployments.</p>
<h2>Set up your VPS</h2>
<p><em>If you’re a Chef wiz and already have some recipes to provision servers for running Rails apps, you can skip ahead to the section on setting up Discourse. Chef is still on my learning TODO list, so I provisioned a server from scratch.</em></p>
<p>I started with a <a href="http://www.digitalocean.com/" target="_blank" rel="nofollow noopener noreferrer">Digital Ocean</a> droplet running the latest release of Ubuntu 13.04, but presumably any VPS running a recent release (12 04 or 12.10) of Ubuntu should do. Make sure your server has at least 1GB of RAM; you’ll need it to compile all of Discourse’s assets. I chose the 1GB / 1CPU plan, and it’s served me well thus far.</p>
<h3>Secure your server</h3>
<p><a href="http://www.linode.com/" target="_blank" rel="nofollow noopener noreferrer">Linode</a> already has an <a href="http://library.linode.com/securing-your-server" target="_blank" rel="nofollow noopener noreferrer">excellent guide</a> on hiding your server from prying eyes, so I suggest following it to the T. I don’t want to reinvent the wheel here, so this is what I followed exactly. I set my username to <code>goodbrews</code> for deploying anything related to that site.</p>
<h3>Change your server’s hostname</h3>
<p>Depending on what characters you included in your droplet’s name, DigitalOcean may not set the hostname correctly. To change my server’s hostname, I did the following:</p>
<pre><code class="language-sh">echo forums.goodbre.ws | sudo tee /etc/hostname
</code></pre>
<p>Because I deployed my installation to <code>forums.goodbre.ws</code>, I changed the first line of my <code>/etc/hosts/</code> file to:</p>
<pre><code>127.0.0.1 localhost forums.goodbre.ws
</code></pre>
<h3>Create a swapfile</h3>
<p>If you’re running on a VPS service that automatically provisions a swap (such as Linode), you can skip this. On Digital Ocean, however, you’ll need to do this yourself. 1GB of RAM was enough for my first deployment but, when I tried to run a second, I wasn’t able to allocate enough memory to compile assets. Creating a swap, however, has fixed this issue for me. Let’s create a 512MB swapfile:</p>
<pre><code class="language-sh">sudo dd if=/dev/zero of=/swapfile bs=1024 count=512k
sudo mkswap /swapfile
sudo swapon /swapfile
</code></pre>
<p>Then, you’ll need to edit <code>/etc/fstab</code> and paste in the following line:</p>
<pre><code>/swapfile       none    swap    sw      0       0
</code></pre>
<p>Finally, prevent your swapfile from being readable:</p>
<pre><code class="language-sh">sudo chown root:root /swapfile
sudo chmod 0600 /swapfile
</code></pre>
<h3>Install Ruby 2.0.0-p195 and bundler</h3>
<p>SSH into your VPS as your deployment user and install the necessary dependencies to build Ruby on Ubuntu:</p>
<pre><code class="language-sh">sudo apt-get install build-essential openssl libreadline6 libreadline6-dev \
  curl git-core zlib1g zlib1g-dev libssl-dev libyaml-dev libsqlite3-dev \
  sqlite3 libxml2-dev libxslt-dev autoconf libc6-dev libgdbm-dev \
  ncurses-dev automake libtool bison subversion pkg-config libffi-dev
</code></pre>
<p>Now, I personally enjoy using rbenv and ruby-build, so this tutorial will follow installation prodecures around those tools. Feel free to use RVM, chruby, or some other Ruby version manager if you wish.</p>
<pre><code class="language-sh"># Install rbenv
git clone git://github.com/sstephenson/rbenv.git ~/.rbenv

# Setup rbenv initialization
echo 'export PATH=&quot;$HOME/.rbenv/bin:$PATH&quot;' &gt;&gt; ~/.bash_profile
echo 'eval &quot;$(rbenv init -)&quot;' &gt;&gt; ~/.bash_profile

# Restart your shell to use rbenv
exec $SHELL -l

# Install ruby-build
git clone https://github.com/sstephenson/ruby-build.git ~/.rbenv/plugins/ruby-build

# Install Ruby 2.0.0-p195 and set it as your user's default
rbenv install 2.0.0-p195; rbenv global 2.0.0-p195
</code></pre>
<p><em>NOTE: the installation guide for <code>rbenv</code> mentions using <code>~/.profile</code> instead of <code>~/.bash_profile</code>, but I found the latter to work for me. You can try either if you have issues.</em></p>
<p>Finally, you’ll need bundler.</p>
<pre><code class="language-sh">gem install bundler --no-rdoc --no-ri
</code></pre>
<h3>Install nginx</h3>
<p>I prefer to run my Rails apps under nginx. Plus, Discourse comes with a sample configuration file that requires little tweaking. No brainer!</p>
<pre><code class="language-sh">sudo apt-get install nginx
</code></pre>
<h3>Install PostgreSQL, Redis, and other required libraries</h3>
<p>Discourse requires <a href="http://www.postgresql.org" target="_blank" rel="nofollow noopener noreferrer">PostgreSQL</a> and <a href="http://redis.io/" target="_blank" rel="nofollow noopener noreferrer">Redis</a> to run, so let’s install those as well:</p>
<pre><code class="language-sh">sudo apt-get install postgresql-9.1 postgresql-contrib-9.1 redis-server \
  libxml2-dev libxslt-dev libpq-dev make g++
</code></pre>
<p>Note that postgresql-contrib-9.1 will install the hstore extension for PostgreSQL, which Discourse requires. You’ll need to create a postgres role and Discourse’s production database:</p>
<pre><code class="language-sh"># Create your postgres role (I named mine goodbrews, go figure)
sudo -u postgres createuser goodbrews -s -P

# Create discourse's production database (I named mine discourse_production)
createdb -U goodbrews discourse_production
</code></pre>
<p>And now your VPS should be set up and ready to run Discourse! That means it’s
time to…</p>
<h2>Set up Discourse</h2>
<p>Because you’ll need to configure it yourself, Discourse is a prime candidate for <a href="https://github.com/discourse/discourse/fork" target="_blank" rel="nofollow noopener noreferrer">forking</a>. If you don’t have Ruby installed on your own computer, yet, I recommend installing rbenv locally. The same steps above will work (or follow rbenv’s <a href="https://github.com/sstephenson/rbenv" target="_blank" rel="nofollow noopener noreferrer">installation guide</a>), so make sure you have Ruby 2.0.0-p195 installed.</p>
<p>Then, clone your fork locally (on your own computer).</p>
<pre><code class="language-sh">git clone git@github.com:$GITHUB_USERNAME/discourse.git
cd discourse
gem install bundler
bundle install
</code></pre>
<h3>Edit the Discourse configuration files</h3>
<p>Copy the sample configuration files…</p>
<pre><code class="language-sh">cp config/database.yml.sample config/database.yml
cp config/nginx.sample.conf config/nginx.conf
cp config/redis.yml.sample config/redis.yml
cp config/environments/production.sample.rb config/environments/production.rb
</code></pre>
<p>… and edit them with your own information. <code>database.yml</code> should, of course, be edited to use the postgres role that you configured earlier (I also changed the name of the production database to <code>discourse_production</code> as I mentioned earlier).</p>
<p><code>redis.yml</code> is unlikely to change if you’ve followed this guide and are running it on the same server.</p>
<p><code>production.rb</code> is where you’ll want to configure how your Discourse installation’s email gets sent out. Personally, I use <a href="http://mandrillapp.com/" target="_blank" rel="nofollow noopener noreferrer">Mandrill</a> and configured my app to hit their API.</p>
<p><code>nginx.conf</code> is trickier depending on your needs, however. For instance, I set my discourse installation up to use HTTPS and provisioned SSL certificates from <a href="https://startssl.com/" target="_blank" rel="nofollow noopener noreferrer">StartSSL</a>. For the purposes of this guide, we’ll stick with basic HTTP. I prefer deploying my applications into the home directory of my deployment user to ensure proper ownership, so my <code>nginx.conf</code> file ended up looking like:</p>
<pre><code class="language-nginx">upstream discourse {
  server unix:///home/goodbrews/discourse/shared/sockets/thin.0.sock max_fails=1 fail_timeout=15s;
  server unix:///home/goodbrews/discourse/shared/sockets/thin.1.sock max_fails=1 fail_timeout=15s;
  server unix:///home/goodbrews/discourse/shared/sockets/thin.2.sock max_fails=1 fail_timeout=15s;
  server unix:///home/goodbrews/discourse/shared/sockets/thin.3.sock max_fails=1 fail_timeout=15s;
}

server {
  listen 80;
  gzip on;
  gzip_min_length 1000;
  gzip_types application/json text/css application/x-javascript;

  server_name forums.goodbre.ws;

  sendfile on;

  keepalive_timeout 65;

  location / {
    root /home/goodbrews/discourse/current/public;

    location ~ ^/t\/[0-9]+\/[0-9]+\/avatar {
      expires 1d;
      add_header Cache-Control public;
      add_header ETag &quot;&quot;;
    }

    location ~ ^/assets/ {
      expires 1y;
      add_header Cache-Control public;
      add_header ETag &quot;&quot;;
      break;
    }

    proxy_set_header  X-Real-IP  $remote_addr;
    proxy_set_header  X-Forwarded-For $proxy_add_x_forwarded_for;
    proxy_set_header  X-Forwarded-Proto $scheme;
    proxy_set_header  Host $http_host;

    # If the file exists as a static file serve it directly without
    # running all the other rewite tests on it
    if (-f $request_filename) {
      break;
    }

    if (!-f $request_filename) {
      proxy_pass http://discourse;
      break;
    }
  }
}
</code></pre>
<p>The <code>max_fails=1 fail_timeout=15s</code> statements on the end of the upstream thin servers is what will help us achieve zero-downtime deployments. As we’re phasing out old thin processes with new ones, any failures on an old thin process will cause nginx to cease handing it requests. It will, instead, hand the request to an old thin process that’s either still running or one of the new thin processes that has already started.</p>
<h3>Configure thin</h3>
<p>You’ll also want a thin configuration file to make your life easier:</p>
<pre><code class="language-sh">thin config -C config/thin.yml --servers 4 -e production
</code></pre>
<p>This will generate a file, <code>config/thin.yml</code>, which you should then go edit. Make sure to set the <code>chdir</code>, <code>log</code>, <code>pid</code>, and <code>socket</code> keys to the correct paths. My thin configuration looks like this:</p>
<pre><code class="language-yaml">---
chdir: /home/goodbrews/discourse/current
environment: production
address: 0.0.0.0
port: 3000
timeout: 30
log: /home/goodbrews/discourse/shared/log/thin.log
pid: /home/goodbrews/discourse/shared/pids/thin.pid
socket: /home/goodbrews/discourse/shared/sockets/thin.sock
max_conns: 1024
max_persistent_conns: 100
require: []
wait: 30
servers: 4
daemonize: true
onebyone: true
</code></pre>
<p>Make sure you add the <code>onebyone: true</code> key. This is the secret sauce that will make your Thin servers restart… You guessed it… One by one. Zero downtime!</p>
<h3>Create a secret token</h3>
<p>Discourse doesn’t ship with a secret token set, since it’s not recommended to be version controlled. Therefore, you’ll need to generate one and set it yourself:</p>
<pre><code class="language-sh">bundle exec rake secret
</code></pre>
<p>Open up <code>config/initializers/secret_token.rb</code> and paste the generated secret
somewhere (probably line 10). You can delete the rest of the file.</p>
<h3>Set up Capistrano</h3>
<p>Add the following to Discourse’s <code>Gemfile</code>:</p>
<pre><code class="language-ruby">gem &quot;capistrano&quot;, require: nil
gem &quot;capistrano-rbenv&quot;, require: nil
</code></pre>
<p>Create a <code>Capfile</code> in Discourse’s root directory:</p>
<pre><code class="language-ruby">load &quot;deploy&quot; if respond_to?(:namespace)
load &quot;deploy/assets&quot;
Dir[&quot;vendor/plugins/*/recipes/*.rb&quot;].each { |plugin| load(plugin) }
load &quot;config/deploy&quot;
</code></pre>
<p>Create your deployment recipe file at <code>config/deploy.rb</code>. Here’s what mine looks like:</p>
<pre><code class="language-ruby"># Require the necessary Capistrano recipes
require &quot;capistrano-rbenv&quot;
require &quot;bundler/capistrano&quot;
require &quot;sidekiq/capistrano&quot;

# Repository settings, forked to an outside copy
set :repository, &quot;git@github.com:goodbrews/forums.git&quot;
set :deploy_via, :remote_cache
set :branch, fetch(:branch, &quot;master&quot;)
set :scm, :git
ssh_options[:forward_agent] = true

# General Settings
set :deploy_type, :deploy
default_run_options[:pty] = true

# Server Settings
set :user, &quot;goodbrews&quot;
set :use_sudo, false
set :rails_env, :production
set :rbenv_ruby_version, &quot;2.0.0-p195&quot;

role :app, &quot;forums.goodbre.ws&quot;, primary: true
role :db,  &quot;forums.goodbre.ws&quot;, primary: true
role :web, &quot;forums.goodbre.ws&quot;, primary: true

# Application Settings
set :application, &quot;discourse&quot;
set :deploy_to, &quot;/home/#{user}/#{application}&quot;

namespace :deploy do
  # Tasks to start, stop and restart thin. This takes Discourse's
  # recommendation of changing the RUBY_GC_MALLOC_LIMIT.
  desc &quot;Start thin servers&quot;
  task :start, :roles =&gt; :app, :except =&gt; { :no_release =&gt; true } do
    run &quot;cd #{current_path} &amp;&amp; RUBY_GC_MALLOC_LIMIT=90000000 bundle exec thin -C config/thin.yml start&quot;, :pty =&gt; false
  end

  desc &quot;Stop thin servers&quot;
  task :stop, :roles =&gt; :app, :except =&gt; { :no_release =&gt; true } do
    run &quot;cd #{current_path} &amp;&amp; bundle exec thin -C config/thin.yml stop&quot;
  end

  desc &quot;Restart thin servers&quot;
  task :restart, :roles =&gt; :app, :except =&gt; { :no_release =&gt; true } do
    run &quot;cd #{current_path} &amp;&amp; RUBY_GC_MALLOC_LIMIT=90000000 bundle exec thin -C config/thin.yml restart&quot;
  end

  # Sets up several shared directories for configuration and thin's sockets,
  # as well as uploading your sensitive configuration files to the serer.
  # The uploaded files are ones I've removed from version control since my
  # project is public. This task also symlinks the nginx configuration so, if
  # you change that, re-run this task.
  task :setup_config, roles: :app do
    run  &quot;mkdir -p #{shared_path}/config/initializers&quot;
    run  &quot;mkdir -p #{shared_path}/config/environments&quot;
    run  &quot;mkdir -p #{shared_path}/sockets&quot;
    put  File.read(&quot;config/database.yml&quot;), &quot;#{shared_path}/config/database.yml&quot;
    put  File.read(&quot;config/redis.yml&quot;), &quot;#{shared_path}/config/redis.yml&quot;
    put  File.read(&quot;config/environments/production.rb&quot;), &quot;#{shared_path}/config/environments/production.rb&quot;
    put  File.read(&quot;config/initializers/secret_token.rb&quot;), &quot;#{shared_path}/config/initializers/secret_token.rb&quot;
    sudo &quot;ln -nfs #{release_path}/config/nginx.conf /etc/nginx/sites-enabled/#{application}&quot;
    puts &quot;Now edit the config files in #{shared_path}.&quot;
  end

  # Symlinks all of your uploaded configuration files to where they should be.
  task :symlink_config, roles: :app do
    run  &quot;ln -nfs #{shared_path}/config/database.yml #{release_path}/config/database.yml&quot;
    run  &quot;ln -nfs #{shared_path}/config/newrelic.yml #{release_path}/config/newrelic.yml&quot;
    run  &quot;ln -nfs #{shared_path}/config/redis.yml #{release_path}/config/redis.yml&quot;
    run  &quot;ln -nfs #{shared_path}/config/environments/production.rb #{release_path}/config/environments/production.rb&quot;
    run  &quot;ln -nfs #{shared_path}/config/initializers/secret_token.rb #{release_path}/config/initializers/secret_token.rb&quot;
  end
end

after &quot;deploy:setup&quot;, &quot;deploy:setup_config&quot;
after &quot;deploy:finalize_update&quot;, &quot;deploy:symlink_config&quot;

namespace :db do
  desc &quot;Seed your database for the first time&quot;
  task :seed do
    run &quot;cd #{current_path} &amp;&amp; psql -d discourse_production &lt; pg_dumps/production-image.sql&quot;
  end
end

after  &quot;deploy:update_code&quot;, &quot;deploy:migrate&quot;
</code></pre>
<p>These should be all of the necessary Capistrano recipes you need to run your first deployment. Let’s do that now!</p>
<h2>Deploy Discourse</h2>
<p>Okay, you’ve got your server set up and Discourse configured… Let’s do this!</p>
<pre><code class="language-sh"># Set up the server's directory structure for Capistrano
cap deploy:setup

# Seed the database you created earlier
cap db:seed

# Do the first deploy! No server is running yet, so do a cold deployment
cap deploy:cold
</code></pre>
<h2>Configure Discourse</h2>
<p>Congratulations! You should now have Discourse running on your VPS. Now that your forums are up and running, you’ll want to follow the <a href="https://github.com/discourse/discourse/wiki/The-Discourse-Admin-Quick-Start-Guide" target="_blank" rel="nofollow noopener noreferrer">Quick Start Guide</a> that the Discourse team has written up.</p>
<h2>Keeping Discourse up-to-date</h2>
<p>You’ll, of course, want to keep your copy of Discourse up-to-date with the main repository. To do so, <code>cd</code> into your local copy and set up an upstream remote:</p>
<pre><code class="language-sh">git remote add upstream git@github.com:discourse/discourse.git
git fetch upstream
git merge upstream/master
git push origin master
cap deploy
</code></pre>
<p>And there you go. Whenever you want to update, just fetch upstream, merge it into master, push, and redeploy.</p>
<h3>Anything here wrong?</h3>
<p>Please let me know if the above steps didn’t work for you. It’s possible that I accidentally left a step out of my process or that the process has changed and I don’t know about it. Reply to <a href="http://meta.discourse.org/t/deploy-discourse-to-an-ubuntu-vps-using-capistrano/6353" target="_blank" rel="nofollow noopener noreferrer">my thread on meta.discourse.org</a> where I’ve posted this guide, or shoot me an email using the link at the top.</p>]]>
    </content>
    <published>2013-05-02T17:42:00Z</published>
    <updated>2022-12-29T21:08:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/deploying-discourse-with-capistrano"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/314446741631927333</id>
    <title>From 1.5 GB to 50 MB: Debugging Memory Usage in Redis</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>Back when I was still working on <a href="https://github.com/davidcelis/goodbre.ws" target="_blank" rel="nofollow noopener noreferrer">goodbre.ws</a> (well… rewriting, really), there was one big issue I was dealing with. Really big. Big enough to have taken down the entire site semi-permanently without me having access to more expensive servers. Long story short, my Redis database grew out of control and ballooned to 1.5 GB. The day before publising this for the first time, I reduced that memory usage to a cool 50 MB.</p>
<!--more-->
<p>In 2012, goodbre.ws was featured in <a href="https://www.huffpost.com/entry/goodbrews-beer-recommendations-exploration-website_n_1930567" target="_blank" rel="nofollow noopener noreferrer">The Huffington Post</a> and <a href="https://lifehacker.com/goodbrews-keeps-track-of-the-beer-you-like-suggests-br-5947790" target="_blank" rel="nofollow noopener noreferrer">Lifehacker</a>; with those features came a small horde of new users, and I quickly found myself with 7000 new accounts. This was quite a change from humble beginnings with only a couple hundred friends, classmates and colleagues. Unfortunately, with all of these new people came a few problems. First, my background jobs to refresh recommendations slowed waaay down. I eventually discovered an I/O bottleneck in the background worker that was hitting both PostgreSQL and Redis more than it reasonably should have been. However, as more and more people were getting their recommendations, I saw my server’s RAM usage get worse and worse. It wasn’t long before the amount of RAM that Redis was trying to use had exceeded the amount of RAM on my server (1 GB). I couldn’t reasonably afford larger servers, especially at this rate of growth, and I was forced to take goobre.ws down.</p>
<p>I started doing a lot of thinking about my Redis usage and what could possibly be causing it to use so much memory. The first thing I considered was the length of my keys. Typical redis keys in my instance looked something like <code>recommendable:users:1234:liked_beers</code>. Okay. Multiply that by five for each user (for dislikes, bookmarks, hidden beers, etc.) and there’s a lot of repetition in the key names. They’re also quite long. Maybe Redis was eating memory by storing tens of thousands of really long key names in RAM? I decided to try shortening them to a more compact format: <code>u:1234:lb</code> for example.</p>
<p>With lots of hope, I renamed my keys and restarted Redis. Hopes dashed: that reduced memory usage by a meager 0.01 GB. That’s 10 MB which, for RAM, may be worth exploring again in the future. However, it obviously wasn’t my main problem.</p>
<p>Being a fairly junior engineer at the time, optimization wasn’t a rabbit hole I’d had to go down many times. I was hardly an expert, and I let my own self-consiousness and self-doubt get in the way of doing real testing. I immediately jumped to conclusions that maybe Redis wasn’t the tool I should be using. Maybe I should revert to storing ratings in PostgreSQL and accept what would certainly be a large performance hit during recommendation generation (Redis was perfect for this in my case because I was using <a href="https://davidcel.is/articles/collaborative-filtering-with-likes-and-dislikes/">set math in a binary rating system</a>).</p>
<p>I toyed with the idea of finding some other data store. At the time, I couldn’t find a comparable key-value store that had the features I needed from Redis, namely both sets and sorted sets with the various operations I relied on for matching user similarities. The SET and ZSET data structures were just far too perfect for my usage. But what could I do? Redis obviously was becoming too expensive for me. I would have to find something else.</p>
<p>I thought about moving my ratings into a Neo4j graph database. It could make for an interesting way of generating recommendations, like a simple graph traversal out from a user to connected (similar) users to find beers that those users like frequently. That might even be faster, but I worried that the recommendations themselves wouldn’t be as good.</p>
<p>I also thought about moving the ratings back into PostgreSQL and initializing some sort of Ruby Set mapping when the Rails app booted up, but that would probably take just as much memory if not more. I’d just be moving RAM usage from Redis into Ruby.</p>
<p>Finally, the day before originally writing this post, I did what I should have done in the first place: I downloaded a <a href="https://github.com/sripathikrishnan/redis-rdb-tools" target="_blank" rel="nofollow noopener noreferrer">memory profiling tool built for Redis</a> that would give me key-by-key memory usage stats. What I discovered was surprising, only because it outlined a problem I remember thinking about so long ago that I thought I had already addressed it.</p>
<p>My issue was how much data I was retaining in the sorted sets (ZSETs) I was creating. Each user got two ZSETs. One was used to store user similarities, pairing other users’ IDs with a calculated similarity value as the rank. The other ZSET stored recommendations, pairing beer IDs with the probability of the user liking that beer. In each ZSET, I was keeping those values for every other user and for every other beer. Multiply that by what became a database of 7000 users and 60000 beers and, well, you can guess what happened. Let’s just say that a lot of these sets were over 1 MB each.</p>
<p>I thought I was already truncating the ZSETs filled with similarity values by using a k-Nearest-Neighbor setting that I had introduced to <a href="https://github.com/davidcelis/recommendable" target="_blank" rel="nofollow noopener noreferrer">Recommendable</a>. That setting uses some specified number of similar users when generating recommendations as opposed to every user. Enabling that setting reduced the size of each similarity set from around 7000 values to 200 (100 similar users and 100 dissimilar users).</p>
<p>Additionally, I implemented a setting to specify how many recommendations should be kept at any one time for each user. I only ever show 10 recommendations, so maintaining those probabilities for every single beer was ridiculous. I reduced that to 100 as well so people can immediately get more recommendations if they rate their current ones. After truncating all of the sets to their specified lengths, I watched in awe as the memory Redis had been consuming dropped from 1.5 GB to 50 MB.</p>
<p>If you yourself are a <a href="https://github.com/davidcelis/recommendable" target="_blank" rel="nofollow noopener noreferrer">Recommendable</a> user, definitely make use of the <code>nearest_neighbors</code>, <code>furthest_neighbors</code>, and <code>recommendations_to_store</code> settings!</p>]]>
    </content>
    <published>2013-03-20T18:42:00Z</published>
    <updated>2022-12-29T20:57:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/from-1-5-gb-to-50-mb-debugging-memory-usage-in-redis"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/243763743421367332</id>
    <title>Stop Validating Email Addresses with Regex</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>Just stop, y’all. It’s a waste of your time and your effort. Put down your Google search for an <a href="http://www.google.com/search?q=email+regex" target="_blank" rel="nofollow noopener noreferrer">email regular expression</a>, take a step back, and breathe.<!--more--> There’s a famous quote that goes:</p>
<blockquote>
<p>Some people, when confronted with a problem, think, “I know, I’ll use regular expressions.” Now they have two problems.</p>
<p>— <a href="http://regex.info/blog/2006-09-15/247" target="_blank" rel="nofollow noopener noreferrer">Jamie Zawinski</a></p>
</blockquote>
<p>Here’s a fairly common code sample from Rails Applications with some sort of authentication system:</p>
<pre><code class="language-ruby">class User &lt; ActiveRecord::Base
  # This regex is from https://github.com/plataformatec/devise, the most
  # popular Rails authentication library
  validates :email, format: {with: /\A[^@]+@([^@\.]+\.)+[^@\.]+\z/}
end
</code></pre>
<p>If you’re experienced at Regex, this seems simple. If (like me when I first saw this) you AREN’T experienced at Regex, it takes a while to parse. But believe me, it can get way worse…</p>
<pre><code class="language-ruby">class User &lt; ActiveRecord::Base
  validates :email, format: {with: /^(|(([A-Za-z0-9]+_+)|([A-Za-z0-9]+\-+)|([A-Za-z0-9]+\.+)|([A-Za-z0-9]+\++))*[A-Za-z0-9]+@((\w+\-+)|(\w+\.))*\w{1,63}\.[a-zA-Z]{2,6})$/i}
end
</code></pre>
<p>Or even worse still…</p>
<pre><code class="language-ruby">class User &lt; ActiveRecord::Base
  validates :email, presence: true, email: true
end

class EmailValidator &lt; ActiveModel::EachValidator
  EMAIL_ADDRESS_QTEXT           = Regexp.new &quot;[^\\x0d\\x22\\x5c\\x80-\\xff]&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_DTEXT           = Regexp.new &quot;[^\\x0d\\x5b-\\x5d\\x80-\\xff]&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_ATOM            = Regexp.new &quot;[^\\x00-\\x20\\x22\\x28\\x29\\x2c\\x2e\\x3a-\\x3c\\x3e\\x40\\x5b-\\x5d\\x7f-\\xff]+&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_QUOTED_PAIR     = Regexp.new &quot;\\x5c[\\x00-\\x7f]&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_DOMAIN_LITERAL  = Regexp.new &quot;\\x5b(?:#{EMAIL_ADDRESS_DTEXT}|#{EMAIL_ADDRESS_QUOTED_PAIR})*\\x5d&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_QUOTED_STRING   = Regexp.new &quot;\\x22(?:#{EMAIL_ADDRESS_QTEXT}|#{EMAIL_ADDRESS_QUOTED_PAIR})*\\x22&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_DOMAIN_REF      = EMAIL_ADDRESS_ATOM
  EMAIL_ADDRESS_SUB_DOMAIN      = &quot;(?:#{EMAIL_ADDRESS_DOMAIN_REF}|#{EMAIL_ADDRESS_DOMAIN_LITERAL})&quot;
  EMAIL_ADDRESS_WORD            = &quot;(?:#{EMAIL_ADDRESS_ATOM}|#{EMAIL_ADDRESS_QUOTED_STRING})&quot;
  EMAIL_ADDRESS_DOMAIN          = &quot;#{EMAIL_ADDRESS_SUB_DOMAIN}(?:\\x2e#{EMAIL_ADDRESS_SUB_DOMAIN})*&quot;
  EMAIL_ADDRESS_LOCAL_PART      = &quot;#{EMAIL_ADDRESS_WORD}(?:\\x2e#{EMAIL_ADDRESS_WORD})*&quot;
  EMAIL_ADDRESS_SPEC            = &quot;#{EMAIL_ADDRESS_LOCAL_PART}\\x40#{EMAIL_ADDRESS_DOMAIN}&quot;
  EMAIL_ADDRESS_PATTERN         = Regexp.new &quot;#{EMAIL_ADDRESS_SPEC}&quot;, nil, &quot;n&quot;
  EMAIL_ADDRESS_EXACT_PATTERN   = Regexp.new &quot;\\A#{EMAIL_ADDRESS_SPEC}\\z&quot;, nil, &quot;n&quot;

  def validate_each(record, attribute, value)
    unless value =~ EMAIL_ADDRESS_EXACT_PATTERN
      record.errors.add(attribute, options.fetch(:message, &quot;is invalid&quot;))
    end
  end
end
</code></pre>
<p>Yeesh. Is something that complex really necessary? If you actually check the Google query I linked above, people have been writing (or trying to write) <a href="http://tools.ietf.org/html/rfc2822" target="_blank" rel="nofollow noopener noreferrer">RFC-compliant</a> regular expressions to parse email addresses for years. They can get ridiculously convoluted as in the case above and, according to the specification, are often too strict anyway.</p>
<p>Sections <a href="http://tools.ietf.org/html/rfc2822#section-3.2.4" target="_blank" rel="nofollow noopener noreferrer">3.2.4</a> and <a href="http://tools.ietf.org/html/rfc2822#section-3.4.1" target="_blank" rel="nofollow noopener noreferrer">3.4.1</a> of the RFC go into the requirements on how an email address needs to be formatted and, well, there’s not much you can’t do in your email address when quotes or backslashes are involved. The local string (the part of the email address that comes before the @) can contain any of these characters: <code>! $ &amp; * - = ^ ` | ~ # % ' + / ? _ { }</code></p>
<p>But guess what? You can use pretty much any character you want if you escape it by surrounding it in quotes. For example, <code>&quot;Look at all these spaces!&quot;@example.com</code> is a valid email address. Nice. Technically, even the <code>@</code> and domain aren’t strictly required. If omitted, the email address is assumed to be local and mail could still be deliverable to the same machine using that address.</p>
<p>For those reasons, I just run any email address against a much more lenient regular expression:</p>
<pre><code class="language-ruby">class User &lt; ActiveRecord::Base
  validates :email, format: {with: /@/}
end
</code></pre>
<p>Simple, right? While email addresses don’t technically have to have an <code>@</code> symbol, almost every application will require it; you’ll be sending email remotely, after all. This is often the most I do and, when paired with a confirmation field for the email address on your registration form, can alleviate most problems with user error. But what if I told you there were a way to determine whether or not an email is valid without resorting to regular expressions at all? It’s surprisingly easy, and you’re probably already doing it anyway.</p>
<h2>Just send them an email already</h2>
<p>No, I’m not joking. Just send your users an email. The activation email is a practice that’s been in use for years, but it’s often paired with complex validations that the email is formatted correctly. If you’re going to send an activation email to users, why bother using a gigantic regular expression?</p>
<p>Think about it this way: I register for your website under the email address <code>qwiufaisjdbvaadsjghb@gmail.com</code>. C’mon. That’s probably going to bounce, courtesy of the preeminent mail daemon, but the formatting is fine; it’s a valid email address. To fix this problem, you can implement an activation system where, after registering, I am sent an email with a link I must click. This is to verify that I actually own that email address before my account is activated. At this point, why keep parsing email addresses for their format? The result of sending an email to a badly formatted email address would be the same: it’ll bounce. If your user enters a bad email address, they won’t get the activation email and they’ll try to register again if they really care about using your site. Of course, having too many of your emails bounce can lead to your domain being marked as spam, so this approach should be combined with rate limiting to avoid being blacklisted.</p>
<p>So eschew your fancy regular expressions already. If you really want to check email addresses right on the signup page, you can include a confirmation field so they have to type it twice. Better yet, use a library like <a href="https://github.com/mailcheck/mailcheck" target="_blank" rel="nofollow noopener noreferrer">mailcheck</a> to perform a bit of client-side validation that catches some potential common errors. It comes down to is this: if your user enters a bad email address, you shouldn’t make it more of a problem for yourself than you have to. A complex regex validation on the email address doesn’t introduce an additional solution, it introduces an additional problem. If you really, really want to make sure people are typing in an email address, just use the <code>/@/</code> regular expression and call it done. If that makes you nervous, then check for the dot too: <code>/.+@.+\..+/i</code>. But then send an email! Anything more is overkill.</p>]]>
    </content>
    <published>2012-09-06T17:33:00Z</published>
    <updated>2022-12-29T20:02:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/stop-validating-email-addresses-with-regex"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/230344571089847331</id>
    <title>Internationalization and the Rails Inflector</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>Here’s a peek at an enhancement for <code>ActiveSupport::Inflector</code> that was introduced in Rails 4.0 and which I’m proud to have <a href="https://github.com/rails/rails/commit/7db0b073fec6bc3e6f213b58c76e7f43fcc2ab97" target="_blank" rel="nofollow noopener noreferrer">contributed</a>. The Inflector is the part of Rails responsible for a good amount of the cool stuff you can do with Strings: pluralization, singularization, titleization, humanization, tableization… The list goes on. Rails uses these methods extensively to map between, say, Model names, Controller names, database table names, and more. Let’s dive into the new stuff!</p>
<!--more-->
<p>Previously, the Inflector could handle only one set of rules at a time. Rails provides a lengthy list of singularization and pluralization rules for English, but what if somebody wants to specify how certain words in a different language should be pluralized? In Rails 3, they had two options: they could define the words as irregularities in <code>config/initializers/inflections.rb</code> or add them one by one into locale files. The former is bad because it involves mixing two languages into one set of rules, and the latter can lead to large, cluttered locale files when internationalizing a website.</p>
<p>As of Rails 4, however, I’m happy to offer a better solution for Rails developers in the process of internationalization: a multilingual Inflector. It can manage a complete set of inflection rules for each locale! Rails will still only provide a list of inflections in English, but it’s now much easier to specify your own sets of rules for additional locales, or even to have them provided to you via a gem. You can, as previously, specify these rules in <code>config/initializers/inflections.rb</code>:</p>
<pre><code class="language-ruby"># You can pass any locale code as a parameter. This defaults to `:en`,
# but here we’re adding rules for Spanish (`:es`)!
ActiveSupport::Inflector.inflections(:es) do |inflect|
  inflect.plural(/$/, 's')
  inflect.plural(/([^aeéiou])$/i, '\1es')
  inflect.plural(/([aeiou]s)$/i, '\1')
  inflect.plural(/z$/i, 'ces')
  inflect.plural(/á([sn])$/i, 'a\1es')
  inflect.plural(/é([sn])$/i, 'e\1es')
  inflect.plural(/í([sn])$/i, 'i\1es')
  inflect.plural(/ó([sn])$/i, 'o\1es')
  inflect.plural(/ú([sn])$/i, 'u\1es')

  inflect.singular(/s$/, '')
  inflect.singular(/es$/, '')

  inflect.irregular('el', 'los')
end
</code></pre>
<p>After specifying our ruleset for Spanish, we can see it in action by passing the same locale code to methods defined by the inflector:</p>
<pre><code class="language-ruby">'avión'.pluralize      # =&gt; 'avións'
'avión'.pluralize(:es) # =&gt; 'aviones'
'luz'.pluralize        # =&gt; 'luzs'
'luz'.pluralize(:es)   # =&gt; 'luces'
</code></pre>
<p>Note that the Inflector will still default to English and use <code>:en</code> as the locale unless specified, despite the existence of the <code>I18n.default_locale</code> configuration. This is to avoid breaking applications that have been wired internally to use English pluralization rules for mapping between Model, Controller, and table names.</p>
<h3>Where can I get rules for other locales?</h3>
<p>I’m happy to offer a solution for this too! I’ve created a gem called <a href="https://github.com/davidcelis/inflections" target="_blank" rel="nofollow noopener noreferrer">Inflections</a> that I’m hoping can serve as a central repository for inflection rules of all locales. Unfortunately, I’m not a master of linguistics. Aside from English, I am only comfortable with Spanish, and so that’s the only additional set of inflection rules I myself have provided (aside from a shorter list of English inflections that I believe to be stripped to the essentials). If you or anybody you know are a Rails developer and fluent in another language, please consider <a href="https://github.com/davidcelis/inflections/fork_select" target="_blank" rel="nofollow noopener noreferrer">forking Inflections</a>, making a list of inflections (<code>lib/inflections/&lt;locale&gt;.rb</code>) with tests (<code>test/&lt;locale&gt;_test.rb</code>) and <a href="https://github.com/davidcelis/inflections/pull/new/master" target="_blank" rel="nofollow noopener noreferrer">opening a pull request</a>.</p>]]>
    </content>
    <published>2012-07-31T16:50:00Z</published>
    <updated>2022-12-29T19:39:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/internationalization-and-the-rails-inflector"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/225734192133047330</id>
    <title>The State of Rails Inflections</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>Ah, the Rails Inflector; one way or another, we all know and love it. This little part of ActiveSupport has a lot of responsibility in our Rails applications, after all! It’s used to determine table names, class names, our resourceful routes, foreign keys… It’s a small part of ActiveSupport, but it has a <em>huge</em> footprint. Outside of Rails’ internal use of the Inflector, it also provides a lot of useful mechanisms for string manipulation to Rails developers. But how does the Inflector actually handle things like singularization and pluralization? English isn’t a regular language, after all!</p>
<!--more-->
<p>There are a lot of grammatical rules to consider when converting between the singular or plural form of various words, so what does the Inflector consider? There must be some magic involved, right? According to the documentation, Rails defines these inflections directly in ActiveSupport… <code>lib/active_support/inflections.rb</code> to be exact. Let’s take a looksy, shall we?</p>
<pre><code class="language-ruby">#--
# Defines the standard inflection rules. These are the starting point for
# new projects and are not considered complete. The current set of inflection
# rules is frozen. This means, we do not change them to become more complete.
# This is a safety measure to keep existing applications from breaking.
#++
#--
# Defines the standard inflection rules. These are the starting point for
# new projects and are not considered complete. The current set of inflection
# rules is frozen. This means, we do not change them to become more complete.
# This is a safety measure to keep existing applications from breaking.
#++
module ActiveSupport
  Inflector.inflections(:en) do |inflect|
    inflect.plural(/$/, &quot;s&quot;)
    inflect.plural(/s$/i, &quot;s&quot;)
    inflect.plural(/^(ax|test)is$/i, '\1es')
    inflect.plural(/(octop|vir)us$/i, '\1i')
    inflect.plural(/(octop|vir)i$/i, '\1i')
    inflect.plural(/(alias|status)$/i, '\1es')
    inflect.plural(/(bu)s$/i, '\1ses')
    inflect.plural(/(buffal|tomat)o$/i, '\1oes')
    inflect.plural(/([ti])um$/i, '\1a')
    inflect.plural(/([ti])a$/i, '\1a')
    inflect.plural(/sis$/i, &quot;ses&quot;)
    inflect.plural(/(?:([^f])fe|([lr])f)$/i, '\1\2ves')
    inflect.plural(/(hive)$/i, '\1s')
    inflect.plural(/([^aeiouy]|qu)y$/i, '\1ies')
    inflect.plural(/(x|ch|ss|sh)$/i, '\1es')
    inflect.plural(/(matr|vert|ind)(?:ix|ex)$/i, '\1ices')
    inflect.plural(/^(m|l)ouse$/i, '\1ice')
    inflect.plural(/^(m|l)ice$/i, '\1ice')
    inflect.plural(/^(ox)$/i, '\1en')
    inflect.plural(/^(oxen)$/i, '\1')
    inflect.plural(/(quiz)$/i, '\1zes')

    inflect.singular(/s$/i, &quot;&quot;)
    inflect.singular(/(ss)$/i, '\1')
    inflect.singular(/(n)ews$/i, '\1ews')
    inflect.singular(/([ti])a$/i, '\1um')
    inflect.singular(/((a)naly|(b)a|(d)iagno|(p)arenthe|(p)rogno|(s)ynop|(t)he)(sis|ses)$/i, '\1sis')
    inflect.singular(/(^analy)(sis|ses)$/i, '\1sis')
    inflect.singular(/([^f])ves$/i, '\1fe')
    inflect.singular(/(hive)s$/i, '\1')
    inflect.singular(/(tive)s$/i, '\1')
    inflect.singular(/([lr])ves$/i, '\1f')
    inflect.singular(/([^aeiouy]|qu)ies$/i, '\1y')
    inflect.singular(/(s)eries$/i, '\1eries')
    inflect.singular(/(m)ovies$/i, '\1ovie')
    inflect.singular(/(x|ch|ss|sh)es$/i, '\1')
    inflect.singular(/^(m|l)ice$/i, '\1ouse')
    inflect.singular(/(bus)(es)?$/i, '\1')
    inflect.singular(/(o)es$/i, '\1')
    inflect.singular(/(shoe)s$/i, '\1')
    inflect.singular(/(cris|test)(is|es)$/i, '\1is')
    inflect.singular(/^(a)x[ie]s$/i, '\1xis')
    inflect.singular(/(octop|vir)(us|i)$/i, '\1us')
    inflect.singular(/(alias|status)(es)?$/i, '\1')
    inflect.singular(/^(ox)en/i, '\1')
    inflect.singular(/(vert|ind)ices$/i, '\1ex')
    inflect.singular(/(matr)ices$/i, '\1ix')
    inflect.singular(/(quiz)zes$/i, '\1')
    inflect.singular(/(database)s$/i, '\1')

    inflect.irregular(&quot;person&quot;, &quot;people&quot;)
    inflect.irregular(&quot;man&quot;, &quot;men&quot;)
    inflect.irregular(&quot;child&quot;, &quot;children&quot;)
    inflect.irregular(&quot;sex&quot;, &quot;sexes&quot;)
    inflect.irregular(&quot;move&quot;, &quot;moves&quot;)
    inflect.irregular(&quot;zombie&quot;, &quot;zombies&quot;)

    inflect.uncountable(%w(equipment information rice money species series fish sheep jeans police))
  end
end
</code></pre>
<p>As you can see from the comment, this file is mostly no longer touched; the last commit that <em>did</em> touch this file was in 2017, but Rails hasn’t accepted changes to the rules since possibly Rails 2 or even Rails 1. So, although this is a snapshot of Rails’ inflections in the most recent release (<code>7.0.4</code>), it’s unlikely to ever change. And, well… Yikes. This brings me to what I really want to discuss: the state of Rails’ Singularization and Pluralization rules. I think it’s a mess.</p>
<h2>Pluralization in English is not regular</h2>
<p>There are only a few basic rules in English for pluralization. Because we’re speaking in terms of text, I’ll try to keep these rules based on characters rather than sounds. However, one important rule does depend on “<a href="http://en.wikipedia.org/wiki/Sibilant" target="_blank" rel="nofollow noopener noreferrer">sibilant</a>” sounds, which are defined as a sound made by directing air through the sharp edge of your teeth and your tongue (i.e. “sh”, “ss”, “dge”, etc.). While prevalent, this can be difficult to detect in text and there are definitely edge cases.</p>
<h3>The rules</h3>
<ul>
<li>If the word ends with a “sibilant” sound, the plural form ends with “es” (dish → dishes or kiss → kisses) or “s” if the word already ends with an “e” (such as fridge → fridges or judge → judges)</li>
<li>Most words that end with an “o” preceded by a consonant pluralize as “oes” (potato → potatoes, avocado → avocadoes).</li>
<li>Most words that end with a “y” preceded by a consonant pluralize as “ies” (lady → ladies, berry → berries)</li>
</ul>
<p>Aside from these rules, however, all other regular plurals are achieved by adding an “s”.</p>
<h3>Some exceptions</h3>
<ul>
<li>Words of foreign origin are exempt from the “oes” rule (piano → pianos, zero → zeros, kimono → kimonos).</li>
<li>Proper nouns that end with a y are exempt from the “ies” rule (Germany → Germanys, Cody → Codys).</li>
</ul>
<p>These are just two sets of exceptions, however, and these are moreso rules that are exceptions to other rules. English pluralization is riddled with other exceptions that are inconsistent:</p>
<ul>
<li>Some words that end in an “f” have that “f” mutated to a “v” during pluralization (calf → calves, shelf → shelves, leaf → leaves) due in part to the evolution of old/middle English to standard English.</li>
<li>Some words with double “o”s replace those “o”s with “e”s (goose → geese, foot → feet).</li>
<li>Many words are both singular and plural (buffalo, money, sheep, series, fish, coffee) and are therefore uncountable.</li>
<li>Some words can even be pluralized <em>multiple ways</em> depending on context (indices/indexes, staffs/staves)! That’s a case that the Rails Inflector can never hope to get right.</li>
<li>And, of course, some words are just plain irregular (child → children, man → men, mouse → mice, datum → data, etc.).</li>
</ul>
<p>How can Rails hope to consider all of these exceptions when English is such an irregular and fluid language? How should Rails handle the edge cases and irregularities? The answer is simple: it shouldn’t.</p>
<h2>The inflector should be based on rules, not exceptions</h2>
<p>The current inflections that Rails defines are riddled with both rules and exceptions. The file has become such a mess, and so many people were submitting pull requests (<a href="https://github.com/rails/rails/pull/7086" target="_blank" rel="nofollow noopener noreferrer">#7086</a> <a href="https://github.com/rails/rails/pull/345" target="_blank" rel="nofollow noopener noreferrer">#345</a> <a href="https://github.com/rails/rails/pull/3930" target="_blank" rel="nofollow noopener noreferrer">#3930</a> <a href="https://github.com/rails/rails/pull/3910" target="_blank" rel="nofollow noopener noreferrer">#3910</a> <a href="https://github.com/rails/rails/pull/6820" target="_blank" rel="nofollow noopener noreferrer">#6820</a> <a href="https://github.com/rails/rails/pull/2457" target="_blank" rel="nofollow noopener noreferrer">#2457</a> and the list goes on and on and on…) to either fix inflections or add new ones, that inflections in Rails are now frozen. From the documentation for <code>ActiveSupport::Inflector</code>:</p>
<blockquote>
<p>The Rails core team has stated patches for the inflections library will not be accepted in order to avoid breaking legacy applications which may be relying on errant inflections. If you discover an incorrect inflection and require it for your application, you’ll need to correct it yourself.</p>
</blockquote>
<p>This makes sense, but I feel that this situation is unfortunate, and it seem like the Rails core team agrees. Don’t get me wrong, many of these pull requests <em>should</em> be closed. A common response to these patches is “Rails cannot possibly include all inflections by default,” and I completely agree. Rails has already found itself in a situation where it has defined way too many inflections that are exceptions or irregularities, such as ox → oxen, crisis → crises, and the aforementioned case of index → indices (even though this pluralization is purely contextual). <em>Many</em> of the “rules” defined in Rails’ inflections are really exceptions. Some of these exceptions are narrow and affect only one or two words. Some of the exceptions admittedly make sense, but should instead be defined as irregularities rather than singular/plural inflections. I’ll gloss over why some of the current inflections <em>don’t</em> make sense:</p>
<ul>
<li>axis → axes, testis → testes: these are special rules that should be defined as irregularities.</li>
<li>octopus → octopi, virus → viri: these special rules are actually disputed, as octopuses and viruses are more used and accepted.</li>
<li>octopi → octopi, viri → viri, oxen → oxen: these words are not singular, so pluralization should not even be attempted.</li>
<li>buffalo → buffaloes, tomato → tomatoes, hive → hives, alias → aliases, status → statuses: these all follow regular pluralization rules and shouldn’t have needed to be defined as special cases.</li>
<li>matrix → matrices, vertex → vertices, index → indices: indices and indexes are both accepted depending on the context, and it’s likely that neither matrices nor vertices are used enough in Rails applications to warrant a special rule.</li>
<li>quiz → quizzes: similar to the above; this is an irregularity.</li>
<li>mouse → mice, louse → lice: exceptional words that are unlikely see the light of day in most Rails applications.</li>
<li>news ⇄ news: this could have just been defined as an uncountable rather than a pluralization rule.</li>
<li>ox → oxen: A special rule that should be an irregularity, but is also unlikely to be used in most Rails applications</li>
</ul>
<p>As you can see, we have numerous inflections defined as special cases even though they follow a regular pluralization rule. Many shouldn’t be defined in the first place, even as irregularities instead of singularization/pluralization rules. Many involve words that most Rails applications will never need to use. The “Zombie” rule, for example, was only added because a website devoted to Rails tutorials, <a href="http://railsforzombies.org/" target="_blank" rel="nofollow noopener noreferrer">Rails for Zombies</a>, noticed that <a href="https://github.com/rails/rails/pull/2457" target="_blank" rel="nofollow noopener noreferrer">generators were singularizing “zombies” as “zomby”</a> due to “zombies” being irregular. Perhaps a better idea would have been to take that as an opportunity to provide a quick lesson on inflections to new Rails users. Oh, and don’t even get me started on the ridiculously added, archaic plural form of “cow”: <code>inflect.irregular('cow', 'kine')</code></p>
<p>Of course, there are some exceptions and irregularities that make sense to define. To name a few, “child” is a frequently-used term in programming and computer science, “person” is a somewhat frequently-used model name in Rails applications, and “half” or “life” are also used frequently in programming depending on the area. I won’t argue that they should be defined within the framework as irregularities.</p>
<h2>Why not <code>.unfreeze</code>, or at least fix, the default inflections?</h2>
<p>When I originally wrote this post, Rails 4.0 was on its way, which I thought would be a great opportunity to clean up the default inflections. Isn’t a major release the perfect time to eschew the worry of breaking existing applications for the betterment of the framework? Rails is a huge piece of software; any upgrades should be done with caution. In fact, most major (or even <em>minor</em>) version bumps of Rails now involve detailed instructions and tools to help users upgrade their applications. Do inflections really need to be backwards compatible as long as the CHANGELOG, documentation, and upgrade guides/tools make these changes clear? Upgrading Rails versions has required extreme caution for many years now.</p>
<h2>Until then, a better set of defaults:</h2>
<p>Until Rails core decides its time to clean up inflections, I’ve provided a more sane set of singularization and pluralization rules in the form of a gem:</p>
<p><a href="https://github.com/davidcelis/inflections" target="_blank" rel="nofollow noopener noreferrer">davidcelis/inflections</a></p>
<p>Here’s the difference:</p>
<ul>
<li>4 pluralization rules (down from 21)</li>
<li>5 singularization rules (down from 27)</li>
<li>3 irregularities (down from 7)</li>
<li>1 uncountable (down from 10)</li>
</ul>
<p>Ahhh… Much better.</p>]]>
    </content>
    <published>2012-07-18T23:30:00Z</published>
    <updated>2022-12-29T18:10:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/the-state-of-rails-inflections"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/166992125752247329</id>
    <title>Collaborative Filtering with Likes and Dislikes</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>We’ve talked about some of the pitfalls of the <a href="https://davidcel.is/articles/why-i-hate-five-star-ratings/">five-star rating system</a> and how a binary system based on likes and dislikes can be much better, but what does using this kind of rating system look like in practice? How can we take a user’s likes and dislikes and use them to generate helpful recommendations? The answer, as with the five-star system, is through collaborative filtering, but we can rely on a methodology better suited to a binary system!</p>
<!--more-->
<h2>Collaborative filtering</h2>
<p>If you aren’t familiar with collaborative filtering, its a technique used in recommendation engines to predict (filter) the interest of a user by collecting data about the interests of many users (collaborating). There are many different types of collaborative filtering, but for our system of likes and dislikes, we’ll be focusing on memory-based collaborative filtering, which uses ratings submitted by users to calculate the similarity between those users.</p>
<p>Even within just the scope of memory-based collaborative filtering, there are a number of algorithms or techniques to calculate similarity. A few of the more widely used algorithms or formulae include <a href="https://en.wikipedia.org/wiki/Euclidean_distance" target="_blank" rel="nofollow noopener noreferrer">Euclidean Distance</a>, <a href="https://en.wikipedia.org/wiki/Pearson_product-moment_correlation_coefficient" target="_blank" rel="nofollow noopener noreferrer">Pearson’s Correlation</a>, <a href="https://en.wikipedia.org/wiki/Cosine_similarity" target="_blank" rel="nofollow noopener noreferrer">Cosine-based vector similarity</a>, and the <a href="https://en.wikipedia.org/wiki/K-nearest_neighbor_algorithm" target="_blank" rel="nofollow noopener noreferrer">k-Nearest Neighbor algorithm</a>. These are all well documented and multiple example implementations are available should you wish to know more. They’re all great for the heavily-used five-star system and, while they’d work fine for likes and dislikes, we have an interesting alternative that I feel is better suited! So let’s talk about an algorithm I don’t see used often but which works great for our binary system. Let’s talk about the Jaccard similarity coefficient!</p>
<h2>Jean-Luc Jaccard?</h2>
<p>No, no, no; we’re talking about <em>Paul</em> Jaccard, a botanist that performed research near the turn of the 20th century. Jaccard’s research led him to develop the <em>coefficient de communauté</em>, or what is known in English as the <a href="https://en.wikipedia.org/wiki/Jaccard_index" target="_blank" rel="nofollow noopener noreferrer">Jaccard index</a> (or Jaccard similarity coefficient).<sup class="footnote-ref"><a href="#fn1" id="fnref1">1</a></sup> The Jaccard index is a simple calculation of similarity between sample sets. Where the aforementioned collaborative filtering algorithms can quickly become mathematically complex, the Jaccard index is rather simple! It can be described as the size of the intersection between two binary sample sets divided by the size of the union between the same sample sets. Whew! That description might be a little difficult to follow, so here’s how to represent it in math:</p>
<p>$$
J(u_1,u_2)=\frac{\left |u_1 \bigcap u_2\right |}{\left |u_1\bigcup u_2 \right |}
$$</p>
<p>This formula can rather intuitively be used with likes and dislikes! Let’s say we’re comparing two users: u<sub>1</sub> and u<sub>2</sub>. How does one intersect two users? How does one union them? Well, we don’t want to intersect or union the people themselves; this isn’t Mary Shelly’s <em>Frankenstein</em>! If we’re using the Jaccard index for collaborative filtering, we want both of these operations to deal with the users’ ratings. Let’s say that the intersection is the set of items which both users have rated. The union would then be the combined set of items that <em>either</em> u<sub>1</sub> <em>or</em> u<sub>2</sub> has rated. But how does this work with the actual ratings? Let’s modify the formula a bit to deal with the likes and dislikes themselves:</p>
<p>$$
J(u_1,u_2)=\frac{\left |L_{u1} \bigcap L_{u2}\right |+\left |D_{u1} \bigcap D_{u2}\right |}{\left |u_1\bigcup u_2 \right |}
$$</p>
<p>Now we’re getting somewhere! What we’ve got now is looking more collaborative and filtery for sure. We find the number of items that both u<sub>1</sub> and u<sub>2</sub> like, add it to the number of items that both u<sub>1</sub> and u<sub>2</sub> dislike, and then divide that by the total number of different items that u<sub>1</sub> and u<sub>2</sub> have rated. This is a great start, but we can go even further to match users up.</p>
<h2>Birds of a feather flock together, and opposites repel</h2>
<p>We may be defining our similarity as our agreed upon interests and disinterests, but what about our <em>disgreements</em>? If our shared likes and dislikes are important factors in calculating our similarity, we can use discrepancies in our ratings to incorporate <em>disimilarity</em> into our calculations. To show you what I mean, let’s tweak the formula a bit more, shall we?</p>
<p>$$
J(u_1,u_2)=\frac{\left |L_{u1} \bigcap L_{u2}\right |+\left |D_{u1} \bigcap D_{u2}\right |-\left |L_{u1} \bigcap D_{u2}\right |-\left |D_{u1} \bigcap L_{u2}\right |}{\left |u_1\bigcup u_2 \right |}
$$</p>
<p>Whew! This looks a lot more complex than the original formula, but we can walk through it together. Now, in addition to finding the agreements between u<sub>1</sub> and u<sub>2</sub>, we’re finding their disagreements! The agreements between u<sub>1</sub> and u<sub>2</sub> are the same as before. Their <em>disagreements</em> are conversely defined as the number of items that u<sub>1</sub> likes but u<sub>2</sub> dislikes and vice versa. All we do is subtract the number of disagreements from the number of agreements, and divide by the total number of items liked or disliked across the two users.</p>
<p>Previously, when we were calculating similarity only based on agreements, our coefficient would have been bounded between 0 and 1. However, now that we’re accounting for disagreements, our bounds have expanded to being between -1 and 1. You would have a -1.0 similarity value with your polar opposite (e.g. your evil twin that has rated the same items as you, but each one differently) and a 1.0 similarity value with a very fresh clone of yourself (you have both rated the same items in the same ways).</p>
<h2>Okay, read my mind!</h2>
<p>Now that we can reduce the tender, loving relationship between two people to a cold, indifferent number, let’s use that number to predict whether you’ll like or dislike something. Neat! Let’s say we want to predict how you’ll feel about <em>thing</em>. We get every user in our system that has rated <em>thing</em> and start calculating a kind of hive-mind sum. Don’t be afraid, though; this isn’t <em>really</em> a hive mind or intelligent AI! Anyway, if a user liked <em>thing</em>, we add your similarity value with them to the sum. If they disliked it, we subtract instead! The idea behind this is that if someone with tastes similar to yours likes <em>thing</em>, you’ll probably like it too. If they dislike <em>thing</em>, you’re less likely to enjoy it. Likewise, if a user who has a low or negative similarity coefficient with you has rated <em>thing</em>, you’re likely to rate it in the opposite way. Finallym we take this sum and divide it by the total number of people that have rated <em>thing</em>. Done! Like before, let’s let the math speak too:</p>
<p>$$
P(you, thing)=\frac{\sum_{i=1}^{n_L} J(you, u_i) - \sum_{i=1}^{n_D}J(you, u_i)}{n_L + n_D}
$$</p>
<p>In this equation: <em>thing</em> is the thing we want to know if <em>you</em> will like, <em>n<sub>L</sub></em> is the number of users that have liked <em>thing</em>, and <em>n<sub>D</sub></em> is the number of users that have disliked <em>thing</em>.</p>
<h2>Math is cool but how about some code?</h2>
<p>That’s fair. You’ve been very patient and I appreciate you reading all of that! Heck, even if you skipped all the way here, you’re here nonetheless. So here’s a little pseudo-implementation of Jaccardian collaborative filtering (in Ruby, of course)!</p>
<pre><code class="language-ruby">require &quot;set&quot;

class User
  # The collections of objects this user likes and dislikes. These are both
  # best represented using a Set.
  attr_reader :likes, :dislikes

  def similarity_with(user)
    # Set#&amp; is the set intersection operator.
    agreements = (self.likes &amp; user.likes).size
    agreements += (self.dislikes &amp; user.dislikes).size

    disagreements = (self.likes &amp; user.dislikes).size
    disagreements += (self.dislikes &amp; user.likes).size

    # Set#| is the set union operator
    all_items = (self.likes + self.dislikes) | (user.likes + user.dislikes)

    return (agreements - disagreements) / all_items.size.to_f
  end

  def prediction_for(item)
    sum = 0.0
    item.liked_by.each { |user| sum += self.similarity_with(user) }
    item.disliked_by.each { |user| sum -= self.similarity_with(user) }

    rated_by = (item.liked_by + item.disliked_by).size

    return sum / rated_by
  end
end
</code></pre>
<p>This is more or less the way I do things in <a href="https://github.com/davidcelis/recommendable" target="_blank" rel="nofollow noopener noreferrer">recommendable</a> and, while it was still online, <a href="https://github.com/davidcelis/goodbre.ws" target="_blank" rel="nofollow noopener noreferrer">goodbre.ws</a>. I did, however, tweak the algorithm in one major way. For example, in that last stage of calculating the similarity values, I actually divide by <code>self.likes.size + self.dislikes.size</code>. With this change, the similarity value becomes dependent on the number of items that <code>self</code> has rated, but not the number of items that <code>user</code> has rated. As such, this makes their similarity values not be reflective:</p>
<pre><code class="language-ruby">self.similarity_with(user) == user.similarity_with(self)
# =&gt; false unless self.ratings.size == user.ratings.size
</code></pre>
<p>My reasoning behind this is that newer users who have not had a chance to submit likes and dislikes for many objects shouldn’t be punished for simply being new; recommendations for new users can be pretty bad! Say I’ve submitted ratings for five items, you’ve submitted ratings for fifty, and four of these items are the same. If we share the same ratings for three of those items, I want our similarity coefficient to be high. I’m new here, and it’ll potentially help me get better recommendations faster. On the other hand, your fifty ratings means you’ve seen things. You don’t really need the same jump start that I do, so your similarity value with me can stand to be lower.</p>
<h2>The Conclusioning</h2>
<p>The Jaccard index can be a very intuitive way to compare people when your rating system is binary. The other algorithms I mentioned are pretty cool too, but likes, dislikes and set math were just made for each other. They’re like peanut butter and jelly. Bananas and Nutella™. Bored people and reality television. It’s a beautiful partnership that I hope can last forever.</p>
<script type="text/javascript" id="MathJax-script" async src="https://cdn.jsdelivr.net/npm/mathjax@3/es5/tex-mml-chtml.js">
</script>
<section class="footnotes">
<ol>
<li id="fn1">
<p>Although this statistic is named for Paul Jaccard, it was actually first developed by geologist <a href="https://en.wikipedia.org/wiki/Grove_Karl_Gilbert" target="_blank" rel="nofollow noopener noreferrer">Grove Karl Gilbert</a> in 1884; Jaccard independently developed and popularized the same statistic in 1912. It was then independently developed for a third time by T. Tanimoto in 1958, leading the statistic to be known occasionally as the Tanimoto index or Tanimoto coefficient. <a href="#fnref1" class="footnote-backref">↩</a></p>
</li>
</ol>
</section>]]>
    </content>
    <published>2012-02-07T21:10:00Z</published>
    <updated>2022-12-30T00:51:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/collaborative-filtering-with-likes-and-dislikes"/>
  </entry>
  <entry>
    <id>https://davidcel.is/posts/164797665899447328</id>
    <title>Why I Hate Five-Star Ratings</title>
    <content type="html" xml:lang="en">
      <![CDATA[<p>When I was originally developing <a href="https://github.com/davidcelis/goodbre.ws" target="_blank" rel="nofollow noopener noreferrer">goodbre.ws</a> (and, later, <a href="https://github.com/davidcelis/recommendable" target="_blank" rel="nofollow noopener noreferrer">recommendable</a>), the very first thing I had to decide was how users would rate items. Would I give them a standard five star system? Maybe something with more granularity, like allowing for half stars? Or perhaps the humble thumbs up or down? Truth be told, going with a binary thumbs up or down system of likes and dislikes was an easy choice. After all, I <em>hate</em> the five-star rating system.</p>
<!--more-->
<h2>The ★★★★★ scale</h2>
<p>At its core, any star rating scale is just a numeric rating scale with some number of options. The five-star scale is arguably the most classic of its kind (psychologically, five options “feels” like a nice, round number), so it’s not surprising that a lot of websites use it. Most big e-commerce sites like Amazon, eBay, or stores powered by Shopify allow you to rate products using five stars. IMDB uses a ten-star scale, which may as well be a 5-star scale that allows half stars. Untappd uses a five-star scale but with the granularity of quarter stars, which gives you 20 options. There are a lot of numeric systems, but they all have the same, core issues.</p>
<h3>Ambiguity and uncertainty of the scale</h3>
<p>One of my big gripes about numeric scales like the five-star scale is the ambiguity behind the ratings that you are allowed to give. What exactly distinguishes between three stars and four stars? What’s enough to push your rating up to that next star? What’s enough to pull it down? Because of a lack of clarity, star ratings can end up being highly subjective. It’s easy to end up with two people who give an item the same three-star rating when they actually feel differently about that item. Some websites attempt to handle this reasonably; back when Netflix still used a five-star scale, they presented some explanatory text for each star when hovering over while rating a movie:</p>
<ol>
<li>★☆☆☆☆ (Hated it)</li>
<li>★★☆☆☆ (Didn’t like it)</li>
<li>★★★☆☆ (Liked it)</li>
<li>★★★★☆ (Really liked it)</li>
<li>★★★★★ (Loved it)</li>
</ol>
<p>Eventually, Netflix stopped displaying this hover text, instead letting ambiguity creep back in. That being said, even the explanatory text itself can come off as subjective. What does it mean to “really” like a movie? Why are the intervals between the options unequal, with there being no “Really disliked it” option and with the typically neutral three-star text being very much <em>not</em> neutral? Explanatory text can help if done correctly, but it can also add to the subjectivity of submitted ratings.</p>
<h3>Unreliability of ratings</h3>
<p>Because a star rating scale iteslf is so ambiguous and uncertain, the ratings end up reflecting that ambiguity and uncertainty. Many users will not use this scale as intended even with intent given in the form of explanatory text. Some users <em>will</em> use the scale as intended, but that usage is always based on their subjective ability to understand the way the scale should be used.</p>
<p>Despite this, recommendation systems will accept these ratings as statistically accurate communications. Websites with huge samples of users and ratings are less likely to be negatively affected by the unreliable nature of these ratings; as sample sets grow, that unreliability can become normalized. Smaller websites and recommendation systems experiencing the <a href="https://en.wikipedia.org/wiki/Cold_start_(recommender_systems)" target="_blank" rel="nofollow noopener noreferrer">cold start</a>, however, will suffer due to the subjective nature of their small rating samples.</p>
<h3>Binary voting is already happening</h3>
<p>Despite being given a scale with five possible ratings, most people tend to vote in a binary fashion anyway. Back in 2009, YouTube <a href="http://youtube-global.blogspot.com/2009/09/five-stars-dominate-ratings.html" target="_blank" rel="nofollow noopener noreferrer">published some interesting data</a> concerning the ratings that videos had been receiving. As it turns out, a huge majority of videos would receive mostly five-star ratings. I think that YouTube’s takeaway from this data was spot on:</p>
<blockquote>
<p>Seems like when it comes to ratings it’s pretty much all or nothing. Great videos prompt action; anything less prompts indifference.</p>
</blockquote>
<p>The second highest rating was, of course, one star; this is a great example of binary voting in the works. A lot of people give mostly five-star ratings for things they like. If they don’t like a thing, they either give it one star or just bounce without rating the thing at all. I’ve also spoken to numerous friends and acquaintances who admit to giving almost exclusively four-star ratings to things they like, and three-star ratings to things that are “just ok”. In fact, this is a wide-spread phenomenon on Tabelog, a popular website used in Japan to rate restaurants. So many reviews on Tabelog stick to the middle of the road that, on average, <a href="http://tabelog.com/help/score/" target="_blank" rel="nofollow noopener noreferrer">most of its ratings are distributed between 3.1 and 3.5 stars</a>.</p>
<p>YouTube toyed with the idea of switching their rating system to a “favorites” system to “declare your love for a video”, but ultimately settled on likes and dislikes.</p>
<h2>The binary scale (and why it’s better)</h2>
<p>Binary rating scales have gained a lot of popularity. As mentioned earlier, YouTube has now operated on a thumbs up or down rating scale for a long time. Netflix also eventually dropped their star-rating system in favor of likes and dislikes. Reddit and other social networking sites like it use upvotes and downvotes. But what makes a binary system better than a five-star system?</p>
<h3>Less ambiguity</h3>
<p>The binary rating scale removes a large amount of ambiguity present in the star rating systems. Five or more subjective rating options are aggregated down into two options based on words that are easily understandable by native speakers of the language. It is much easier for a person to declare, “Hey, I like this thing” than it is to determine, “Well, I like this thing… But do I ‘three stars’ like it, or do I ‘four stars’ like it?”</p>
<h3>Less subjectivity</h3>
<p>A large amount of subjectivity is also removed. Ratings given that are based directly on feelings are much more likely to match than ratings given based on numbers. This can simplify a lot of situations in which two people may have similar feelings about something but rated it different:</p>
<ul>
<li>Me: “I liked this thing and rated it four stars.”<br/>
Friend: “I liked this thing and rated it five stars.”</li>
<li>Me: “I liked this thing and rated it three stars.”<br/>
Friend: “I didn’t like this thing, so I only gave it three stars.”</li>
</ul>
<p>Our feelings about something are clearly not conveyed well by more granular ratings, and they also don’t match. As I posited earlier, this will normalize as data sets grow, but this does not change that we have no way of knowing whether or not the underlying ratings are truly indicative of agreement. Given a binary scale, however, agreement is much more clear: “We both liked this thing” or “we both disliked this thing.”</p>
<h3>People are already doing this!</h3>
<p>Remember, even within a five-star system or other numeric systems, <em>people are pretty much already rating in this way</em>. Why fight it?</p>
<h3>No middle ground</h3>
<p>Of course, the likes and dislikes are not without their own flaws. Most notably, unless its explicitly added, there isn’t an obvious neutral ground in a binary rating system outside of abstenance. It’s often an all-or-nothing situation in which you either like something or you don’t. This may or may not be an issue for you as the implementor. Personally, when I’m ready to rate an item, I can almost always manage to categorize it into a like or dislike even if its very close. However, if I were to truly feel 100% neutral about something, I would likely ignore that thing and move on rather than rate it. If I have no feelings either way, why would I want it affecting my recommendations?</p>
<h2>tl;dr</h2>
<p>Embrace the binary rating system. It’s much less ambiguous and subjective than its stellar cousin, and it’s much easier for the user to deal with in general. Feelings themselves are more easily comparable than numbers indirectly based on feelings and can lead to more accurate recommendations.</p>]]>
    </content>
    <published>2012-02-01T19:50:00Z</published>
    <updated>2022-12-29T17:23:00Z</updated>
    <link rel="alternate" type="text/html" href="https://davidcel.is/articles/why-i-hate-five-star-ratings"/>
  </entry>
</feed>
