<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Georg's Log - algorithm</title><link href="https://gms.tf/" rel="alternate"></link><link href="https://gms.tf/feeds/algorithm.atom.xml" rel="self"></link><id>https://gms.tf/</id><updated>2020-11-22T16:00:00+01:00</updated><entry><title>Perfect Hashing</title><link href="https://gms.tf/perfect-hashing.html" rel="alternate"></link><published>2020-11-22T16:00:00+01:00</published><updated>2020-11-22T16:00:00+01:00</updated><author><name>Georg Sauthoff</name></author><id>tag:gms.tf,2020-11-22:/perfect-hashing.html</id><summary type="html">&lt;p&gt;The beauty of &lt;a href="https://en.wikipedia.org/wiki/Perfect_hash_function"&gt;perfect hashing&lt;/a&gt; is that you never have to deal with
any collisions during item lookup.
I recently created &lt;a href="https://github.com/gsauthof/phashtable"&gt;libphashtable&lt;/a&gt;, a perfect hashing hash table library for
C/C++ which focuses on minimizing item lookup latency jitter.
This article presents benchmarking results that show how its lookup
times …&lt;/p&gt;</summary><content type="html">&lt;p&gt;The beauty of &lt;a href="https://en.wikipedia.org/wiki/Perfect_hash_function"&gt;perfect hashing&lt;/a&gt; is that you never have to deal with
any collisions during item lookup.
I recently created &lt;a href="https://github.com/gsauthof/phashtable"&gt;libphashtable&lt;/a&gt;, a perfect hashing hash table library for
C/C++ which focuses on minimizing item lookup latency jitter.
This article presents benchmarking results that show how its lookup
times compare to those of a traditional hash table.&lt;/p&gt;
&lt;h2 id="overview"&gt;Overview&lt;a class="headerlink" href="#overview" title="Permanent link"&gt;&amp;para;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;The &lt;a href="https://github.com/gsauthof/phashtable"&gt;libphashtable README&lt;/a&gt; contains some details on the libphashtable
design and on hashing background.
This diagram provides a short summary on how libphashtable works:&lt;/p&gt;
&lt;p&gt;&lt;img alt="libphashtable lookup scheme" src="https://gms.tf/image/libphash-lookup.svg"&gt;&lt;/p&gt;
&lt;h2 id="results"&gt;Results&lt;a class="headerlink" href="#results" title="Permanent link"&gt;&amp;para;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;&lt;img alt="Perfect Hashing Boxenplot" src="https://gms.tf/image/perfect-hashing-bplot.svg"&gt;&lt;/p&gt;
&lt;p&gt;This graph is a &lt;a href="https://seaborn.pydata.org/generated/seaborn.boxenplot.html"&gt;boxenplot&lt;/a&gt; (a &lt;a href="https://en.wikipedia.org/wiki/Box_plot"&gt;boxplot&lt;/a&gt; variant, a.k.a. letter-value
plot) that describes all measured item lookup access times.
It's a comparison of a standard hash-table, i.e. &lt;code&gt;std::unordered_map&lt;/code&gt; (umap), against libphashtable (ptable), using different item hash functions on a high-end CPU (Intel Xeon Gold 6246) vs. a low-end CPU (Intel Atom C3768).
Again, the libphashtable README's &lt;a href="https://github.com/gsauthof/phashtable#measurements"&gt;Measurements Section&lt;/a&gt; contains further
details on the benchmark setup.&lt;/p&gt;
&lt;p&gt;The median is marked by a dark grey line that is part (or on top)
of the biggest box. The biggest box contains 50 % of the values.
The next smaller boxes contain the next 25 %, 12.5 % etc. of the
measured values. Outliers are drawn in a diamond shape, where, of
course, multiple outliers may be drawn on top of each other.&lt;/p&gt;
&lt;p&gt;Thus, the lower the median the better, less boxes are better than
more, flat boxes are better than higher ones, less outliers
are better than more, etc.&lt;/p&gt;
&lt;p&gt;As expected, using a traditional hash table leads to much latency
jitter.
It's not just outliers, e.g. on Xeon 50 % of the lookups are
distributed over a 5 to 10 ns wide range or so.
While e.g. on the Atom CPU, 50 % of the lookups are
distributed over a 20 ns range or so.
The very simple and old SDBM item hash function over the board yields very good results.
Also, using a more expensive item hash function doesn't really
have a good trade off here, such as less collisions due to a better
distribution of its range, i.e. the boxes and outliers basically
are just shifted without being compressed.&lt;/p&gt;
&lt;p&gt;The graph shows that libphashtable indeed yields a 'perfect' item
lookup latency distribution.
That means the boxes are so flat that the median line covers
them all and there aren't any outliers, in most configurations.
In general, using the SDBM hash function as item hash function is
a safe choice, especially when targeting a low-end CPU.&lt;/p&gt;</content><category term="algorithm"></category><category term="algorithm"></category><category term="datastructure"></category><category term="C"></category><category term="C++"></category></entry></feed>