<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/feed.atom.xml" media="screen"?>
<feed xml:lang="en-US" xmlns="http://www.w3.org/2005/Atom">
  <id>tag:speakerdeck.com,2005:/tomnicholas</id>
  <link rel="alternate" type="text/html" href="https://speakerdeck.com"/>
  <link rel="self" type="application/atom+xml" href="https://speakerdeck.com/tomnicholas.atom"/>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1542991</id>
    <published>2026-05-18T10:26:53-04:00</published>
    <updated>2026-05-18T10:29:05-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/virtual-chunks-at-earthmover-from-theory-to-production"/>
    <title>Virtual Chunks at Earthmover: From Theory to Production</title>
    <content type="html">Talk given to the ESIP Cloud Computing Cluster</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/b49fd7f7f4334eda8c7c1e07a7fd7b51/preview_slide_0.jpg?39438741" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1526829</id>
    <published>2026-04-08T10:14:39-04:00</published>
    <updated>2026-04-08T10:16:56-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/shortcutting-cloud-migrations-with-virtualizarr-icechunk-and-earthmover"/>
    <title>Shortcutting Cloud Migrations with VirtualiZarr, Icechunk, and Earthmover</title>
    <content type="html">Talk given at AMS 2026 in Houston, Texas</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/8a53de8a7f8845d988dc1fc2634a7b8f/preview_slide_0.jpg?39020686" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1416776</id>
    <published>2025-07-30T10:08:10-04:00</published>
    <updated>2025-07-30T10:16:07-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/open-science-the-cloud-native-way"/>
    <title>Open science the cloud-native way</title>
    <content type="html">Talk given to a group of Oxford University Research Software Engineers.</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/1e88a02dd97049caaf190b7316135318/preview_slide_0.jpg?36101411" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1401795</id>
    <published>2025-07-11T10:22:17-04:00</published>
    <updated>2025-07-11T10:25:54-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/virtualizarr-plus-icechunk-talk-at-scipy-2025"/>
    <title>VirtualiZarr + Icechunk talk at SciPy 2025</title>
    <content type="html"></content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/c5bca85e920c411ba6098b2d3edec6aa/preview_slide_0.jpg?35838345" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1399396</id>
    <published>2025-07-09T16:59:54-04:00</published>
    <updated>2025-07-09T17:01:04-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/cubed-talk-at-scipy-2025"/>
    <title>Cubed talk at SciPy 2025</title>
    <content type="html">https://github.com/cubed-dev/cubed</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/ef54deab87694577855b38a61df10ede/preview_slide_0.jpg?35798778" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1365392</id>
    <published>2025-05-06T12:44:10-04:00</published>
    <updated>2025-05-06T12:48:15-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/virtualizarr-and-icechunk-build-a-cloud-optimized-datacube-in-3-lines"/>
    <title>VirtualiZarr &amp; Icechunk: Build a cloud-optimized datacube in 3 lines</title>
    <content type="html">Talk given at the Cloud-Native Geospatial Forum 2025</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/1b60a5c03c2245d4af237302ace1f1b2/preview_slide_0.jpg?34976797" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1324206</id>
    <published>2025-02-12T20:59:04-05:00</published>
    <updated>2025-02-12T21:03:30-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/frost-federated-registry-of-scientific-things-at-pangeo-showcase"/>
    <title>FROST: Federated Registry of Scientific Things (@ Pangeo Showcase)</title>
    <content type="html">Talk slides from Pangeo Showcase February 12th 2025
https://discourse.pangeo.io/t/pangeo-showcase-frost-federated-registry-of-scientific-things-feb-12-2025/4861

Blog post mentioned: https://hackmd.io/@TomNicholas/H1KzoYrPJe

Abstract:

Context: 
The easiest way to store and provide access to big scientific datasets is via ARCO data in S3-compatible cloud object storage. We now have scalable cloud-optimised formats that are version-controlled at rest in object storage (particularly Icechunk for arrays and Iceberg for tables). This is huge, as even dynamically-updated datasets can now be distributed via raw S3, with no other server needed. All the data providers who are paying attention are about to put their data in these formats, but then they will try to advertise the S3 URLs to the world via ad-hoc data catalogs.

Problem: Everyone’s catalogs are disconnected from everyone else’s.

This means:
- No cross-org discoverability (e.g. NASA catalog users won’t see NOAA datasets or vice versa).
- No cross-org tracking of updates (e.g. NOAA datasets derived directly from NASA datasets won’t automatically know if the NASA datasets have been updated upstream).
- Risk of “catalog wars” where platform services compete to make more and more comprehensive “meta-catalogs” which merely track (outdated) links to other orgs’ data.
- Risk that if one platform does win everyone might feel locked in to it via the social network effect.

Solution: Federated catalog protocol with cross-org publish-subcribe model.
- Cross-org discoverability enabled via displaying the contents of the dataset entries being broadcast,
- Cross-org tracking of updates to datasets enabled the same way,
- No need to compete to make a better catalog, as anyone can easily consume and display the entire global catalog, including updates,
- Federated trust model allows proliferation of high-quality centralized services, whilst also guarding against platform lock-in.

How do we build it?: 
Not sure exactly, but the problem is analogous to creating Federated alternatives to centralized social media (i.e. Bluesky/Mastodon vs Twitter). Perhaps we can piggyback off of Bluesky’s ATproto or Mastodon’s ActivityPub?</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/7567eb25ab71405bb6eba98af971d16b/preview_slide_0.jpg?33824610" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1282153</id>
    <published>2024-11-21T10:35:26-05:00</published>
    <updated>2024-11-21T10:38:12-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/virtualizarr-talk-at-met-office"/>
    <title>VirtualiZarr talk at MET Office</title>
    <content type="html">VirtualiZarr talk with section on Icechunk integration, given to the MET Office Architecture Guild.</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/4faa6339cc1e4f1eae40ae61c34b6cf5/preview_slide_0.jpg?32713287" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1183322</id>
    <published>2024-05-15T13:59:41-04:00</published>
    <updated>2024-05-15T14:00:50-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/virtualizarr-create-virtual-zarr-stores-using-xarray-syntax"/>
    <title>VirtualiZarr: Create virtual Zarr stores using xarray syntax</title>
    <content type="html">Long-form talk on the VirtualiZarr package as an alternative to kerchunk for creating virtual Zarr stores which point to archival data (e.g. many netCDF files).

Recording will be posted here (https://discourse.pangeo.io/t/pangeo-showcase-virtualizarr-create-virtual-zarr-stores-using-xarray-syntax/4127)

See the VirtualiZarr repository for more details (https://github.com/TomNicholas/VirtualiZarr)</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/08201d3ddd6c41f780b6295b3d45c6f3/preview_slide_0.jpg?30103983" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1118426</id>
    <published>2023-12-08T11:10:13-05:00</published>
    <updated>2023-12-08T11:11:11-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/whats-next-for-pangeo"/>
    <title>What's next for Pangeo?</title>
    <content type="html">Talk given at the Pangeo Showcase on 6th December 2023.

Intended to start a community discussion, which was then recorded here:

https://discourse.pangeo.io/t/pangeo-showcase-whats-next-for-pangeo/3870</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/958362d142aa4548ba1c52cd55799ae0/preview_slide_0.jpg?28120018" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1106805</id>
    <published>2023-11-15T13:32:07-05:00</published>
    <updated>2023-11-15T13:49:48-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/cubed-pangeo-showcase"/>
    <title>Cubed: Bounded-Memory Serverless Array Processing (Pangeo showcase)</title>
    <content type="html">Long-form talk on using the Cubed package as an alternative to dask.array for processing large datasets in Xarray.

Recording will be posted here (https://discourse.pangeo.io/t/pangeo-showcase-cubed-bounded-memory-serverless-array-processing-in-xarray/3836)

See this blog post for more details (https://xarray.dev/blog/cubed-xarray)</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/01259e756ec445e7922568b508bc65fa/preview_slide_0.jpg?27839252" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/1052064</id>
    <published>2023-07-17T13:32:03-04:00</published>
    <updated>2023-07-17T13:32:50-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/cubed-xarray-lightning-talk-at-scipy2023"/>
    <title>Cubed-xarray lightning talk at SciPy2023</title>
    <content type="html">Short 4-minute talk on using the Cubed package as an alternative to dask.array for processing large datasets in Xarray.

Given as a lightning talk at the SciPy Conference 2023 in Austin, TX.

See this blog post for more details (https://xarray.dev/blog/cubed-xarray)</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/de4f1bba22134944b98fa1dee9d31f61/preview_slide_0.jpg?26394487" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/975513</id>
    <published>2023-01-12T01:23:56-05:00</published>
    <updated>2023-01-12T01:24:45-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/pangeo-for-plasma"/>
    <title>Pangeo for Plasma</title>
    <content type="html">Presentation at the BOUT++ workshop 2023 at LLNL, on why the fusion plasma physics analysis community should learn lessons from the success of the Pangeo model in the geosciences community.</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/a190572ae44e4660bffca3f126fb5451/preview_slide_0.jpg?24031724" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/975204</id>
    <published>2023-01-11T13:23:08-05:00</published>
    <updated>2023-12-11T10:50:14-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/xgcm-staggered-grids-topologies-and-ufuncs-in-python"/>
    <title>xGCM: Staggered grids, topologies, and ufuncs in python</title>
    <content type="html">(This talk was given at the AMS conference in Denver 2023)

Staggered grids such as Arakawa grids are ubiquitous in climate models. Analysis and post-processing tools using finite-volume operators must respect staggering to get correct results. This presents a challenge for users of General Circulation Model (GCM) data, whose analysis routines must follow the idiosyncrasies of particular GCM grids.

The xGCM package [1] is designed to solve this problem, by extending xarray’s data model with information about the grid. It encodes the positions of variables along each axis of the grid, so that operations can respect differences in staggering between variables. It also encodes the topology of the grid, understanding how the spherical Earth is divided into different regions in various models.

XGCM has recently been upgraded by introducing the concept of “Grid Ufuncs”, which are analogous to numpy ufuncs but grid-aware. Users can define their own grid ufuncs, then apply them to their data. We hope that this extensible model can be built upon by scientists who work post-processing all types of climate models.

We present an overview of xGCM and its capabilities, before showing an example of using it on a multi-TB scale oceanographic dataset.

[1] https://github.com/xgcm/xgcm</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/f0a615a3bbd54ce99d27c3b4b4a33b65/preview_slide_0.jpg?24023219" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/974167</id>
    <published>2023-01-09T16:12:40-05:00</published>
    <updated>2023-01-09T16:13:10-05:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/xarray-datatree-hierarchical-data-structures-for-multi-model-science"/>
    <title>xarray-Datatree: Hierarchical Data Structures for Multi-Model Science</title>
    <content type="html">Talk from the AMS 2023 conference.

Abstract:

Real scientific workflows often require working with many heterogeneous but related datasets. Examples in geoscience include: (1) scenario simulations by many different climate models in the same intercomparison project, (2) simulation data at multiple resolutions from a convergence scan or sub-grid-scale study, and (3) observational + simulation data of the same region.

There is a need for a general high-level data structure which can organize such data in an accessible way, whilst still being flexible enough to adapt to the user’s mental model of their data. It should also be intuitive, so that simple operations such as calculating average climatologies are still simple to express. It should also serialize to a commonly-used data format, so as not to create backwards compatibility problems.

The new xarray-datatree [1] package solves these problems, by providing a tree-like hierarchical data structure that is general enough to be useful in a wide variety of cases. Datatree extends xarray - generalizing xarray.Dataset to build upon an interface that many geoscientists are already familiar with. Analysis operations can be mapped over a whole tree, allowing simple operations to be expressed intuitively, even over complex heterogeneous datasets.

Datatree is inspired by netCDF: Xarray’s highest-level object is currently an xarray.Dataset, which stores collections of arrays with a shared coordinate system and corresponds to a single group in a netCDF file. A DataTree object is instead a structured hierarchical collection of Datasets, and would map to multiple netCDF groups. Therefore serialization to and from netCDF files is possible with datatree, so backwards compatibility is maintained.

We will explain the model of datatree, its relation to netCDF &amp; Zarr, and how to use the data structure to simplify your own work. We will also give examples of using datatree with real geoscience datasets, such as CMIP6 model data. [2]

[1] https://github.com/xarray-contrib/datatree
[2] https://medium.com/pangeo/easy-ipcc-part-1-multi-model-datatree-469b87cf9114
</content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/771fad3599c54d23820cd485dd0bfb8d/preview_slide_0.jpg?24001222" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <entry>
    <id>tag:speakerdeck.com,2005:Talk/905477</id>
    <published>2022-08-08T09:41:40-04:00</published>
    <updated>2022-08-08T11:06:36-04:00</updated>
    <link rel="alternate" type="text/html" href="https://speakerdeck.com/tomnicholas/scipy-2022-can-we-analyse-the-largest-ocean-simulation-ever"/>
    <title>Scipy 2022: Can we analyse the largest ocean simulation ever?</title>
    <content type="html"></content>
<media:thumbnail url="https://files.speakerdeck.com/presentations/f664cc92782649ebb9a335507e95d10f/preview_slide_0.jpg?22300100" width='' height='' xmlns:media='http://search.yahoo.com/mrss/'></media:thumbnail>    <author>
      <name>Tom Nicholas (@tomnicholas)</name>
    </author>
  </entry>
  <title>Tom Nicholas (@tomnicholas) on Speaker Deck</title>
  <updated>2026-05-18T10:26:53-04:00</updated>
</feed>
