The Web of a lifetime
Our corner of the internet—the Fermi way
I was pleased to discover the next newsletter to which Manu will be dedicating some of his (abundant, as he admitted) spare time. While reading the announcement post, I felt the need to share a couple of thoughts. I couldn’t convince him that a trillion is not a large number at all—he told me he’s not sure integers are truly infinite because he’s never tried to count all of them. But beyond this somewhat witty exchange, the discussion became interesting when we tried to figure out what might be the actual number of unique web pages an average person visits in their lifetime. Even the very definition of “unique” isn’t exactly straightforward (but we’ll come back to that).
Manu cited a couple of numbers that sound reasonable: the entire WWW might have about 1.5 billion pages, but only about half a million (perhaps a little less) are active. This gives us the scale of the problem, but it’s not such useful information: we’re not interested in a particular property of the internet—that here we consider as a graph —but rather in the most common patterns for exploring it. Or, if you will, what’s the average size of the portion of this graph that an (average) person explores over their internet lifetime.
If you’re wondering how it’s possible to answer such a question, you’re on the right track to meet one of the most important physicists of all time: Enrico Fermi. Yes, him again. If it’s the first time for you, it’s gonna be exciting, I promise.
A quick online search confirms that there is no comprehensive data, nor data that covers a timespan of adequate length, that can help us answer that question with real precision. Sifting through the web a bit, I found a couple of studies that tried to extract people's browsing profiles from their browser histories, or that studied how unique the “fingerprints” we leave on the web while browsing are, and what are the implications. Another one attempts to quantitatively analyze how much of the web is actually explored—alas, by studying a tiny portion of it, but otherwise such an in-depth study would probably not have been possible.
After skimming through these sources and browsing around, I was convinced that the only realistic approach was precisely the one so strongly advocated by Fermi. That is, to arrive at a numerical answer by making a series of approximations—often rough, but sensible—starting from some information we are reasonably certain of. The goal is to arrive at an equally reasonable order of magnitude. It’s likely that it won’t be the right answer—we might be off by a factor of two, five, or if we’re unlucky, ten. But if our reasoning is sound, we won’t be too far off either.
If you’d like to follow my reasoning, here’s a friendly warning: there will be a little more math. Not too much, but even Fermi had to pull out some heavy artillery when the problem got rather complicated. And here we’re not trying to estimate the number of golf balls it would take to fill a school bus. Let’s set aside the balls and put together the first two pieces of our reasoning.
Exploring how?
We need to go back to the definition of unique web page. Not in the sense of “a single one”, but rather one “that doesn’t repeat,” meaning distinct. Because on one thing we can agree: any conscious human being doesn’t browse the web completely at random. We all end up in Wikipedia’s dense network of rabbit holes, but if we tried, we could retrace our steps. The average person doesn’t just bounce around the internet randomly, clicking on any link that happens to be under their mouse cursor or within tapping distance on their smartphone screen.
A second fact: the web is not a uniformly occupied space. There are extremely dense areas and areas that are almost empty. If we go back to the concept of a graph, some portions have a few elements with an astronomical number of connections to many other elements. Other sections have a few sparse elements that are almost completely isolated from everything else.
The third thing that doesn’t require much proof is that we humans are creatures of habit. We always go back to the same supermarket, the same barber, doctor, or dentist. We like to walk the same route, read the news from the same source (or from a selection we’ve curated over time), look for recipes from the same food blogger, or check the weather forecast from the service that best met our expectations during our last vacation. Sometimes there’s a way to objectively evaluate the quality of these choices; other times, it’s all completely subjective.
In the first link I mentioned, there’s a finding that (partially) connects all these observations. The study showed that users who had visited 50 or more distinct websites over a two-week period exhibited fairly recognizable web activity. Those users who visited 150 domains were practically unmistakable.
Discovering species
This problem seems quite peculiar, given that it concerns the internet—a quite peculiar human product, in some ways—but it really isn’t. It is very, very similar to another activity.
Let’s imagine we are a naturalist who has landed on a deserted island. We vaguely know its size, but we have no idea how many or what animal species live there. We set our minds to cataloging them all, one by one. At first, it will be hard to keep up with all the discoveries we make: we might record dozens of species a day. As time goes on, however, we will have explored an ever-larger portion of that animal kingdom. Some species might be inaccessible to us because we lack the right tools or are unable to reach certain areas of the island. After a long enough time, our knowledge will progress very slowly: it will never truly stop, but we wouldn’t be surprised if months went by without encountering any new species. All already seen, recorded, and neatly cataloged in our file cabinet (obviously an analog one, complete with elegant charcoal drawings).
The trend of the total number of new species discovered during our time on the island is called the species discovery curve (or collector’s curve). It is well-known in ecology, and is even connected to a law that correlates a word’s frequency of use with its popularity in a vocabulary (Zipf’s law).
Here’s the first important result of our reasoning: the discovery rate of new species is proportional to how many still remain unknown. This statement can take the form a mathematical equation, from which we can derive an equally important relationship. If we denote the total number of species known to us at time as
where and are two parameters that we can estimate from the data we have collected. The parameter has the dimensions of time and is particularly interesting: it is the time we need to discover about 63% of all existing species.
There is another form that this equation can take, which is much more useful for our problem on the web. We may not necessarily be able to estimate the total number of existing species, but we might be able to estimate the discovery rate, that is, how quickly changes over time—that is, its first derivative.
Discovering our Web
Let’s make the most obvious assumption: web exploration is practically the same process as that of our imaginary naturalist, with a few differences:
| Species discovery | Web discovery |
|---|---|
| The island | The web (or the “accessible web”) |
| Species living on the island | Unique websites |
| Personal “discoverable web” (the sites we could plausibly encounter) | |
| Time to saturate 63% of our discoverable web | |
| Cumulative unique websites visited by time |
We now need to make two important estimates: and . The data we have suggests that a moderately active web user visits between 50 and 150 (perhaps 200) distinct websites over the course of two weeks. We could assume the order of magnitude of as a good Fermi estimate. However, we know for a fact that not all of these domains will be new, never-before-seen ones. If we estimate that only 50% are, that means 50 domains every two weeks, i.e.,
Let’s therefore take 103 as our estimate for . This number might be an overestimate, but let’s remember its meaning: it’s the initial discovery rate. Just like the naturalist’s work, it’s reasonable for it to start very quickly and decrease over time. What will happen after 5 years? Here too, we have to make an estimate: it will certainly have decreased, but by how much? By a factor of ten? A hundred? This allows us to estimate the other parameter, .
| reduced by a factor | Estimate of |
|---|---|
| 10 | years |
| 100 | years |
Interestingly, in both cases, our estimate of —the largest portion of the web we can discover—remains on the order of 103.
There is an extra complication that we cannot completely ignore: this model assumes nearly unchanged behavior throughout one’s life. We know this is not the case for most of us. Our interests change, our tastes change, and therefore what we search for and what interests us most changes. Life stages bring significant behavioral changes, reflected in our interests or what we consider most important. A new job, a new hobby (or side-job), having children, or adopting a dog are all events that will significantly influence our behavior as web surfers. We could therefore assume that the discovery process is actually composed of two independent processes, each with its own set of parameters. One will be faster, the other slower. We would have to re-estimate and for the second process as well. If we assume that the slow component of discovery is 10 times slower than the first, we get a result that seems surprising: our is always on the order of 103. This suggests that, perhaps, at least the order of magnitude of this number is indeed reasonable. But is it really?
The Web is not an island
The crucial point is that the model applicable to our naturalist’s work is based on the hypothesis: discovery rate is proportional to the undiscovered pool, with uniform sampling. While this makes perfect sense in the naturalist’s case, it starts to fall apart when we consider some intrinsic aspects of the web:
- Sampling is not uniform. We follow links, search results, recommendations, all heavily biased toward content more or less popular by some degree.
- Most of our web activity isn’t sampling at all. It’s habitual return to the same well-known sites.
- The pool is not fixed. Websites appear and disappear, algorithms of various types and purposes reshape what’s findable or even reachable.
| Bias | Reality |
|---|---|
| Browsing = sampling | ~90% is habitual return to known sites |
| Uniform sampling from pool | Web is a graph where most areas are invisible |
| Fixed, static pool | New sites emerge; life transitions reopen discovery |
It’s not easy to estimate the quantitative contribution of these effects with precision. However, we can say one thing: they will all be percentage factors unlikely to lead to huge corrections of one or more orders of magnitude. They are factors that only add or subtract a contribution to the final result. Even in our extreme ignorance, our Fermi estimate of 103 remains solid. For a casual internet user, this number might be close to 1000, while for a researcher or a tech worker the number could reach 5000. We know for certain that it’s an upper-bounded number due to a saturation effect similar to that observed in species cataloging.
We know that the species accumulation model is wrong for web browsing, but it gave us a few insights on why. It’s a mathematical scaffold that forces us to name our assumptions: non-uniform sampling, fixed pool, and exploration-dominated discovery. Each assumption reveals a correction. The final estimate isn’t a triumphant derivation, but a negotiation between a tractable model and an intractable reality.
I’m convinced that Fermi would be proud of this exercise. The goal wasn’t precision, but building an estimate we could trust even while acknowledging everything we didn’t know.