<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:cc="http://cyber.law.harvard.edu/rss/creativeCommonsRssModule.html">
    <channel>
        <title><![CDATA[Stories by Kovid Rathee on Medium]]></title>
        <description><![CDATA[Stories by Kovid Rathee on Medium]]></description>
        <link>https://medium.com/@kovidrathee?source=rss-13d513db037------2</link>
        <image>
            <url>https://cdn-images-1.medium.com/fit/c/150/150/0*_CwYR2OmNap47tQO.jpg</url>
            <title>Stories by Kovid Rathee on Medium</title>
            <link>https://medium.com/@kovidrathee?source=rss-13d513db037------2</link>
        </image>
        <generator>Medium</generator>
        <lastBuildDate>Wed, 29 Jul 2026 05:12:25 GMT</lastBuildDate>
        <atom:link href="https://medium.com/@kovidrathee/feed" rel="self" type="application/rss+xml"/>
        <webMaster><![CDATA[yourfriends@medium.com]]></webMaster>
        <atom:link href="http://medium.superfeedr.com" rel="hub"/>
        <item>
            <title><![CDATA[An Epistolary Bandish in Maru Bihag]]></title>
            <link>https://medium.com/khayaliya/an-epistolary-bandish-in-maru-bihag-d9a63b9e70fe?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/d9a63b9e70fe</guid>
            <category><![CDATA[indian-classical-music]]></category>
            <category><![CDATA[maru-bihag]]></category>
            <category><![CDATA[vasantrao-deshpande]]></category>
            <category><![CDATA[khayaliya]]></category>
            <category><![CDATA[khayal]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Tue, 17 Feb 2026 13:31:20 GMT</pubDate>
            <atom:updated>2026-02-17T13:31:20.187Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*QxVtVzR34NEJO2nx8FEcjQ.jpeg" /></figure><h4>KHAYAL</h4><h4>A few sentences on Vasantrao Deshpande’s masterful rendition of <em>Likh Bheji Main Patiyaan</em></h4><p>Remembering a piece I wrote for <a href="https://www.lensculture.com/veronique-lerebours">Véronique Lerebours</a>’ blog back in 2011, I now think about how little music I had heard, though at the time it felt like I was uncovering, discovering, and, at times, unearthing new music from the depths of the internet (I still do that, but YouTube has made it easier now). The piece was called <a href="https://harmonyom.blogspot.com/2011/06/romantic-raaga-maru-bihag.html"><em>The Romantic Raga</em></a>, spelled with two a’s, as in Raaga. You’ll excuse this because that was the phase when I was writing <em>cuz </em>for <em>because</em> without any shame. Anyway.</p><p>In the piece, I write about how a few bandishes or, in fact, sets of bandishes, had captivated me because of their lyrical beauty and the renditions’ musical technicalities. I briefly mention renditions by Ajoy Chakraborty, Iqbal Ahmed Khan, Rajan &amp; Sajan Mishra, Jasraj, and Roshan Ara Begum, but completely miss mentioning Prabha Atre and Vasantrao Deshpande, both of whom took Maru Bihag to the limits of beauty.</p><p>After many weeks, I got back to listening to Khayal today, and stumbled upon a recording of Vasantrao Deshpande’s on YouTube, where he is singing two bandishes, the first one is</p><p>उन ही से जाय कहूँ<br>मोरे मन की बिथा</p><p>दरस बिन कछु ना सुहावे<br>अँखियन नीकी प्रीत गुनवंता</p><p>And the second bandish is why I have named this piece as I have. It talks about the practice or feeling of writing letters to one’s beloved. Once upon a time, romance did take the form of letters. Pining took the form of poetic prose and, often, of poetry itself. Here, in this bandish, it’s not just poetry; it is also set to a set of notes that capture the feeling of this whole scene. It’s sort of like some languages, like German, have words that explain a whole human experience. This rendition of Maru Bihag captures the romance in a similar way.</p><p>लिख भेजी मैं पतियाँ<br>तुम्हरे कारन जुग सी बीतत मोरी रतियाँ</p><p>जा रे जा कगवा इतना मोरा संदेसवा लिए जा<br>जरत मन जोवत बाट तुम्हरी ये अँखियाँ</p><p>Who knows, maybe the heroine of this love story is singing the bandish immortalised by Prabha Atre.</p><p>जागूँ मैं सारी रैना बलमा<br>रसिया मन लागे ना मोरा</p><p>We are fortunate to have several recordings of Vasantrao Deshpande singing the same two bandishes at different phases of his career. No surprise that they all sound beautiful.</p><ul><li><a href="https://www.youtube.com/watch?v=vcHFAriTVcI">Recording №1</a> (~51 minutes, the one that prompted me to write this blog, c. not mentioned in the recording)</li><li><a href="https://www.youtube.com/watch?v=ii-YrqfvyZ4&amp;list=RDii-YrqfvyZ4&amp;start_radio=1">Recording №2</a> (~7 minutes, drut bandish only, c. 1960s)</li><li><a href="https://www.youtube.com/watch?v=-fcfsTnpKVU">Recording №3</a> (~46 minutes, from Diwakar Tole’s archives, c. it feels like slightly late in his career, but not mentioned in the recording)</li><li><a href="https://www.youtube.com/watch?v=2O2inNxmJwg">Recording №4</a> (~28 minutes, c. not mentioned in the recording)</li><li><a href="https://www.youtube.com/watch?v=qzFPzow9pm8">Recording №5</a> (~40 minutes, probably the most famous of them all because it also has video, c. <a href="https://www.discogs.com/release/28893298-Dr-Vasantrao-Deshpande-Vocal-Classical">released in 1982</a>)</li><li><a href="https://www.youtube.com/watch?v=uSPKUK9j3Uo">Recording №6</a> (~38 minutes, at a concert in Pune, c. 1974)</li></ul><p>Listen to some of these recordings, and if you’re as captivated as I was, I’m sure you’ll be yearning for more, so I’ll leave you with some more interesting recordings/playlists of Vasantrao ji</p><ul><li>Farhan Amin’s <a href="https://youtube.com/playlist?list=PLoSKBcYkTfcYjE2Kwxkxqe3R224_00RLZ&amp;si=RMeTYwgEQtn5VNBt">beautifully curated playlist of Vasantrao ji’s recordings</a></li><li><a href="https://www.youtube.com/watch?v=o1jlOrv80Wo">Sureshchandra Nadkarni interviewing Vasantrao Deshpande for Akashvani</a> — look for his explanation about Bilawal Thaat ki Gunakali</li><li><a href="https://youtube.com/playlist?list=PLaZvqMQr-2WPan4x0s0iZ16cO14L-PMVF&amp;si=H9P61y6EXBt1cZsT">Vasantrao ji’s recordings from Diwakar Tole’s personal archives</a></li><li><a href="https://www.youtube.com/watch?v=segJr-uouoE&amp;list=RDiKAU1O-Jtt4&amp;index=5">Two compositions in Paraj</a> and <a href="https://www.youtube.com/watch?v=iKAU1O-Jtt4&amp;list=RDiKAU1O-Jtt4&amp;start_radio=1">a dadra</a> with Zakir Hussain and Sultan Khan, c. 1983</li><li><a href="https://www.youtube.com/watch?v=AtZ0bD27g4U">Vasantrao ji singing something in Marathi in Bhairavi</a></li></ul><p>Happy listening!</p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=d9a63b9e70fe" width="1" height="1" alt=""><hr><p><a href="https://medium.com/khayaliya/an-epistolary-bandish-in-maru-bihag-d9a63b9e70fe">An Epistolary Bandish in Maru Bihag</a> was originally published in <a href="https://medium.com/khayaliya">Khayaliya</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Data Ingestion Tools for Snowflake]]></title>
            <link>https://medium.com/nexifi/data-ingestion-tools-for-snowflake-c59c73f8539f?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/c59c73f8539f</guid>
            <category><![CDATA[openflow]]></category>
            <category><![CDATA[snowflake-data-ingestion]]></category>
            <category><![CDATA[apache-nifi]]></category>
            <category><![CDATA[snowflake]]></category>
            <category><![CDATA[data-ingestion]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Wed, 22 Oct 2025 22:36:38 GMT</pubDate>
            <atom:updated>2025-10-22T22:36:38.210Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*0Dr1YrDM6cprj0k0" /><figcaption>Photo by <a href="https://unsplash.com/@theblowup?utm_source=medium&amp;utm_medium=referral">the blowup</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><h4>SNOWFLAKE</h4><h4>Our recommendations on when to use Openflow, Fivetran, dltHub, and others to ingest data into Snowflake</h4><h3>Background</h3><p>Getting data into Snowflake has gotten easier over time. There is more flexibility in terms of supporting a variety of data sources, different ingestion frequencies, and also cost and performance optimization options. In addition to offering several native data ingestion methods, Snowflake also partners with and integrates with various third-party data ingestion tools, most notably <a href="https://fivetran.com/docs/destinations/snowflake">Fivetran</a>. However, there are others available on the market.</p><p>The most recent addition to this ecosystem is <a href="https://www.snowflake.com/en/product/features/openflow/">Snowflake Openflow</a>, a managed version of an open-source project that automates data flow between systems. This project is called <a href="https://nifi.apache.org">Apache NiFi</a>, which can be run either in your own cloud account using the <a href="https://docs.snowflake.com/en/user-guide/data-integration/openflow/setup-openflow-byoc">BYOC (Bring Your Own Compute)</a> option or using Snowflake’s <a href="https://docs.snowflake.com/en/developer-guide/snowpark-container-services/overview">Snowpark Container Services (SPCS)</a>.</p><p>Then, there’s dlt (not to be confused with <a href="https://www.databricks.com/discover/pages/getting-started-with-delta-live-tables">Databricks’ Delta Live Tables</a>). dlt (from <a href="https://dlthub.com">dltHub</a>) is an open-source Python-first data ingestion tool that has many popular data source connectors. If you don’t find the connector you need, <a href="https://dlthub.com/docs/tutorial/load-data-from-an-api">building one using Python is quite easy</a>. Before we dive deeper into when to use which ingestion tool, here’s a quick overview of these options/categories:</p><ul><li><strong>Snowflake’s native options</strong>: <a href="https://docs.snowflake.com/en/user-guide/data-load-snowpipe-intro">Snowpipe</a>, <a href="https://docs.snowflake.com/en/user-guide/snowpipe-streaming/data-load-snowpipe-streaming-overview">Snowpipe Streaming</a>, <a href="https://other-docs.snowflake.com/en/connectors">Snowflake Connectors</a>, and now <a href="https://docs.snowflake.com/en/user-guide/data-integration/openflow/about">Snowflake Openflow</a>.</li><li><strong>Cloud-platform-aligned ingestion tools</strong>: <a href="https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-connect-kinesis-home.html">AWS Glue/Kinesis</a>, <a href="https://cloud.google.com/products/dataflow?hl=en">Google Cloud Dataflow</a>, and <a href="https://azure.microsoft.com/en-au/products/data-factory">Azure Data Factory</a>.</li><li><strong>Proprietary third-party ingestion tools</strong>: <a href="https://fivetran.com/docs/destinations/snowflake">Fivetran</a>, <a href="https://docs.airbyte.com/integrations/sources/snowflake">Airbyte</a> (managed), <a href="https://dlthub.com/docs/dlt-ecosystem/destinations/snowflake">dltHub</a>, and <a href="https://www.stitchdata.com/data-warehouses/snowflake/">Stitch</a>, among others.</li><li><strong>Open-source ingestion tools:</strong> <a href="https://dlthub.com">dlt (data load tool)</a>, Apache NiFi, Singer, etc.</li></ul><p>Let’s look at when to use which of these options!</p><h3>Snowflake’s Native Ingestion Tools</h3><p>There are four native options to ingest data in <a href="https://www.snowflake.com/en/">Snowflake</a>. We’ll look at all four. The first two assume that you already have infrastructure and tooling for landing/ingesting data into a cloud object store:</p><ul><li><a href="https://docs.snowflake.com/en/user-guide/data-load-snowpipe-intro"><strong>Snowpipe</strong></a> — when your data from various sources is already in cloud object storage, and you need the data from the cloud storage ingested into Snowflake in batches every few minutes or less frequently.</li><li><a href="https://docs.snowflake.com/en/user-guide/snowpipe-streaming/data-load-snowpipe-streaming-overview"><strong>Snowpipe Streaming</strong></a> — as the name suggests, this is the same as Snowpipe but for streaming data. Use this when you need to ingest data at a rate of every few seconds.</li></ul><p>The next two are ingestion options that help you get data directly from the data sources into Snowflake:</p><ul><li><a href="https://docs.snowflake.com/en/developer-guide/native-apps/connector-sdk/about-connector-sdk"><strong>Snowflake Native Application Connectors</strong></a><strong> — </strong>when you want to use pre-built connectors for services like <a href="https://docs.snowflake.com/en/connectors/google/gard/gard-connector-about">Google Analytics</a>, <a href="https://docs.snowflake.com/en/user-guide/data-integration/openflow/connectors/sharepoint/about">SharePoint</a>, MySQL, PostgreSQL, etc., use this option. Please note that only a limited number of connectors are available, and you cannot create your own.</li><li><a href="https://www.snowflake.com/en/product/features/openflow/"><strong>Snowflake Openflow</strong></a><strong> — </strong>when you want an open-source-based connector ecosystem and you need the option to bring your own cloud compute for ingestion while staying native to Snowflake, Openflow is a good option. This project is currently under <a href="https://docs.snowflake.com/en/user-guide/data-integration/openflow/version-history">active development, with updates arriving almost every week</a>.</li></ul><blockquote><strong><em>Bottomline</em></strong></blockquote><blockquote><em>Using Snowflake-native ingestion tool like Openflow makes sense you need an ingestion framework with a range of connectors and the possibility of creating new connectors while also giving you the freedom to run the ingestion on your own cloud compute.</em></blockquote><blockquote><em>The other ingestion options are good for specific use cases (ingesting from cloud storage or from a couple of services applications). Openflow is the only full-fledged direct data ingestion option that Snowflake has for a broad range of connectors.</em></blockquote><blockquote><em>Openflow is based on an open-source project called </em><a href="https://nifi.apache.org"><em>Apache NiFi</em></a><em>. One key thing to remember is that Apache NiFi can have a very steep learning curve as it can impact your team’s ability to use existing connectors, or build and maintain new connectors when required.</em></blockquote><h3>Cloud-platform-aligned ingestion tools</h3><p>If the ingestion tools aren’t Snowflake-native, they can be native to the cloud platform where Snowflake is running. The benefit of using a data ingestion service native to the cloud is that it can be easily integrated with other orchestration, observability, and testing services within that cloud. Here are the services you can use with the cloud platforms:</p><ul><li><a href="https://aws.amazon.com/glue/"><strong>AWS Glue</strong></a><strong> </strong>— has <a href="https://docs.aws.amazon.com/glue/latest/dg/connectors-chapter.html">30+ built-in connectors</a>, most of which are databases, and some of which are services like <a href="https://docs.aws.amazon.com/glue/latest/dg/connecting-to-mixpanel.html">Mixpanel</a> and <a href="https://docs.aws.amazon.com/glue/latest/dg/connecting-to-data-salesforce.html">Salesforce</a>.</li><li><a href="https://cloud.google.com/products/dataflow?hl=en"><strong>Google Cloud Dataflow</strong></a><strong> </strong>—<strong> </strong>is based on <a href="https://beam.apache.org">Apache Beam</a>, which has <a href="https://beam.apache.org/documentation/io/connectors/">50+ I/O connectors</a> for file systems, databases, message queues, data warehouses, etc.</li><li><a href="https://azure.microsoft.com/en-us/products/data-factory"><strong>Azure Data Factory</strong></a> — <a href="https://learn.microsoft.com/en-us/azure/data-factory/connector-overview">has over 100 connectors</a>, including those for data sources such as <a href="https://learn.microsoft.com/en-us/azure/data-factory/industry-sap-connectors">SAP</a>, <a href="https://learn.microsoft.com/en-us/azure/data-factory/connector-salesforce?tabs=data-factory">Salesforce</a>, <a href="https://learn.microsoft.com/en-us/azure/data-factory/connector-zendesk?tabs=data-factory">Zendesk</a>, and <a href="https://learn.microsoft.com/en-us/azure/data-factory/connector-dynamics-crm-office-365?tabs=data-factory">Dynamics 365</a>, among others.</li></ul><blockquote><strong><em>Bottomline</em></strong></blockquote><blockquote><em>Using these cloud-platform-aligned ingestion tools makes sense if the sources you have can be supported by these services. Most data sources that are not natively supported by these cloud-platform-aligned ingestion tools can be found in the cloud platform marketplaces.</em></blockquote><blockquote><em>Alternatively, you can build and maintain your own connectors, especially when dealing with very niche data sources for which you won’t find a well-maintained connector with any of the cloud ingestion tools. Having said that, it is important to keep in mind that building and maintaining connectors is a long-term investment.</em></blockquote><h3><strong>Proprietary third-party ingestion tools</strong></h3><p>Following the recent merger of dbt and Fivetran, <a href="https://www.linkedin.com/in/george-fraser-a0219230">George Fraser</a>, the CTO of Fivetran, discussed why it makes sense for dbt Core to be open-source, but not for Fivetran. He, in essence, said that nobody wants to own the painful process of building and maintaining connectors to services and applications that are ever-changing, which is why it is not a superb fit for open-source.</p><p>Fivetran has over 700 pre-built connectors, while Stitch has more than 300. These tools build and maintain connectors, so you don’t have to, providing a low-code experience for data ingestion. This comes in very handy when you don’t want to dedicate personnel solely to the task of ingestion and dealing with the day-to-day issues that come with it.</p><p>Assuming that one of the proprietary third-party tools has the connectors you want, you need to check for the following things:</p><ul><li>Does the tool provide an option to use your own compute for ingesting data, i.e., does it need to be temporarily or permanently stored in the tool’s cloud environment?</li><li>Cost models can be confusing, sometimes intentionally so. Therefore, you should gain a thorough understanding of what you’ll be billed based on past volumes.</li></ul><blockquote><strong><em>Bottomline</em></strong></blockquote><blockquote><em>There’s a clear benefit of using third-party proprietary tools, especially if you want to build and keep a lean data engineering team. With these tools solving the data ingestion, data analysts and engineers can focus on what matters the most — business use cases around reporting, analysis, etc.</em></blockquote><blockquote><em>While there are many proprietary data ingestion tools in the market, some even with full ETL capabilties, there are only a few that have a wide variety of connectors, so when considering a tool, make sure to look at the maturity of the connectors you need for your data ingestion workflows.</em></blockquote><h3><strong>Free and open-source ingestion tools</strong></h3><p>This is the option that gives you the most freedom in terms of controlling how and where your data is processed, as well as how much it costs to run. Having said that, despite my bias towards and liking for open-source ingestion tools, I recommend them to data teams with caution because the work required to develop using an open-source (or more likely, an open-core) project is often underestimated.</p><p>There are many open-source data ingestion tools, but three stand out for me:</p><ul><li><a href="https://dlthub.com">dlt (data load tool)</a> — is a dev tool for data ingestion; it is not the usual connector catalog in the market. It is known for its Python-first data ingestion philosophy and is a great option for data teams that are well-versed in Python and, in addition to using dlt’s official and community-maintained connectors, are happy to build their own, too. Doing that is way easier in dlt in comparison with some of the other options.</li><li>Apache NiFi — a mature and widely used <a href="https://nifi.apache.org/docs/nifi-docs/html/overview.html#nifi-architecture">JVM-based ingestion engine</a>, which Snowflake has recently adopted as the underlying engine for its new managed service called Openflow. NiFi does come with a steeper learning curve than the other options in this category, especially if you aren’t familiar with the JVM ecosystem.</li><li>Singer — is another very good option for a full-fledged ETL tool. <a href="https://news.ycombinator.com/item?id=13763681">It was open-sourced by Stitch eight years ago</a>. Singer was also the inspiration for the creation of <a href="https://github.com/meltano/meltano">Meltano</a>, which can be considered a more advanced open-source version of the same project.</li></ul><blockquote><strong><em>Bottomline</em></strong></blockquote><blockquote><em>If you want control and freedom, this is the best way to go but it comes with a cost. A better way to go about data ingestion is to use an open-source framework as the foundation but then use a managed service to run it for you. Openflow makes a good choice for managed open-source solution for Snowflake but is not fully Apache NiFi open-source needs a bit of getting used to.</em></blockquote><blockquote><em>If your team can write Python and you want your data team to use a developer-first ingestion framework, dlt might be a very good option. And it is fully open-source. Check out </em><a href="https://dlthub.com/blog/openflow"><em>this blog post</em></a><em> that addresses the dlt vs. Openflow question.</em></blockquote><blockquote><em>One thing to keep in mind is that it is more important to build your data infrastructure on open standards rather than open-source projects. This means that the choice of data ingestion tool is less important than the choice of file formats, table formats, control planes, among other things.</em></blockquote><blockquote><em>In other words, prefer open standards for your data infrastructure in general, and if you have the capacity and capability to run an open-source project for data ingestion, do that, otherwise use a managed version of an open-source project.</em></blockquote><h3>Closing Thoughts</h3><p>Traditionally, ingestion, transformation, modelling, and remodelling of data have be clubbed together under the ETL umbrella. That approach which is now shifting towards separating the concerns with tooling. An ingestion tool is meant to do only that, examples of which you read about in this article. A transformation tool like dbt only does transformation and doesn’t handle ingestion. Having these variations gives you more freedom to choose the kind of architecture you want to implement for your business.</p><p>For now, I’ll leave you with a few important questions you need to ask before deciding on a data ingestion tool for Snowflake:</p><ul><li>What cloud platform(s) do you run Snowflake on?</li><li>Aside from Snowflake, is all of your other data infrastructure on a particular cloud platform? In other words, do you utilize your cloud platform’s data services extensively?</li><li>Can a combination of Snowflake’s native options for data ingestion cover all the data sources you have?</li><li>What is the shape, size, and skillset of your data team? What are their priorities?</li><li>Is there a bias towards low-code/no-code tooling?</li><li>Does your team want more control and freedom in terms of choosing the compute and the process of ingestion into Snowflake?</li><li>What control do you have over the security and privacy of your data while it is being ingested?</li></ul><p>Answering these questions will translate into the decision points that I walked you through in the four sections that preceded this one. The situation with the data stack is that there are hundreds of options, which can be exhausting even to consider, let alone assess for use. That’s why narrowing down the good candidates is the first step. Examining the ingestion tools through these categories will hopefully help you do that more effectively.</p><p><em>Thanks to Paul Harmat, Kroum Klutchkov, and Adrian Brudaru for reading the draft of this blog post.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=c59c73f8539f" width="1" height="1" alt=""><hr><p><a href="https://medium.com/nexifi/data-ingestion-tools-for-snowflake-c59c73f8539f">Data Ingestion Tools for Snowflake</a> was originally published in <a href="https://medium.com/nexifi">Nexifi</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[An Ever New Jaijaiwanti]]></title>
            <description><![CDATA[<div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/khayaliya/an-ever-new-jaijaiwanti-3cda462d5f2c?source=rss-13d513db037------2"><img src="https://cdn-images-1.medium.com/max/2600/0*GHmFa0L4v8C_dJVx" width="5472"></a></p><p class="medium-feed-snippet">Notes from a baithak by Ishwar Ghorpade</p><p class="medium-feed-link"><a href="https://medium.com/khayaliya/an-ever-new-jaijaiwanti-3cda462d5f2c?source=rss-13d513db037------2">Continue reading on Khayaliya »</a></p></div>]]></description>
            <link>https://medium.com/khayaliya/an-ever-new-jaijaiwanti-3cda462d5f2c?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/3cda462d5f2c</guid>
            <category><![CDATA[indian-classical-music]]></category>
            <category><![CDATA[khayal]]></category>
            <category><![CDATA[jaijaiwanti]]></category>
            <category><![CDATA[raaga]]></category>
            <category><![CDATA[classical-music]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Mon, 08 Sep 2025 11:47:37 GMT</pubDate>
            <atom:updated>2025-09-08T12:30:05.959Z</atom:updated>
        </item>
        <item>
            <title><![CDATA[Ramashreya Jha’s Timeless Kirwani]]></title>
            <description><![CDATA[<div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/khayaliya/ramashreya-jhas-timeless-kirwani-42b5b106d9b7?source=rss-13d513db037------2"><img src="https://cdn-images-1.medium.com/max/2600/0*vt1xNSIUenBVhXKj" width="5225"></a></p><p class="medium-feed-snippet">Scattered thoughts on some renditions of compositions in Kirwani</p><p class="medium-feed-link"><a href="https://medium.com/khayaliya/ramashreya-jhas-timeless-kirwani-42b5b106d9b7?source=rss-13d513db037------2">Continue reading on Khayaliya »</a></p></div>]]></description>
            <link>https://medium.com/khayaliya/ramashreya-jhas-timeless-kirwani-42b5b106d9b7?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/42b5b106d9b7</guid>
            <category><![CDATA[ramashreya-jha]]></category>
            <category><![CDATA[khayal]]></category>
            <category><![CDATA[raag]]></category>
            <category><![CDATA[kirwani]]></category>
            <category><![CDATA[indian-classical-music]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Mon, 01 Sep 2025 12:11:41 GMT</pubDate>
            <atom:updated>2025-09-01T12:11:41.803Z</atom:updated>
        </item>
        <item>
            <title><![CDATA[Versent at DataEngBytes 2025]]></title>
            <link>https://medium.com/versent-tech-blog/versent-at-dataengbytes-2025-2fb63864f366?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/2fb63864f366</guid>
            <category><![CDATA[dataengbytes]]></category>
            <category><![CDATA[data-science]]></category>
            <category><![CDATA[data-engineering]]></category>
            <category><![CDATA[prompt-engineering]]></category>
            <category><![CDATA[tech-conference]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Wed, 30 Jul 2025 03:36:37 GMT</pubDate>
            <atom:updated>2025-07-30T03:36:37.959Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*yXWpfVvS9rDTf90I6ll3tQ.jpeg" /><figcaption>Jake Kerr from Versent during his presentation on PromptSQL — Image from the presentation, taken by Dima Marr</figcaption></figure><h4>DATA ENGINEERING</h4><h4>Two days in the life of Data and AI engineers</h4><p>The <a href="https://versent.com.au/data-insights/">Versent Data &amp; Insights</a> attended the main data engineering event of the year in Melbourne in late July. The conference caught my interest last year when they invited <a href="https://huyenchip.com/">Chip Huyen</a>, author of the book <a href="https://huyenchip.com/books/">AI Engineering</a>, for a talk. When I found out our team was going to the event, I checked the schedule of the two parallel tracks and immediately decided which talks were not to be missed and which ones I could probably skip.</p><p>During the two days, I ended up attending thirteen sessions and might have listened to a couple of others in bits and pieces. In my reading, Track 1 was mostly L100-type talks, although there were some exceptions. Track 2 was where L200 and, in some cases, L300 sessions were taking place. This blog post is a brief account of the two-day event, highlighting the aspects I enjoyed most. Here we go, session by session!</p><h3>Building Trino Data Pipelines with SQL or Python</h3><p>I started with <a href="https://www.linkedin.com/in/lestermartin/">Lester Martin</a>’s session, <a href="https://dataengbytes.com/sessions/S0018">Building Trino Data Pipelines with SQL or Python</a>, which was quite good. Lester took us through the history of Presto, how it compared to Spark, how it wasn’t a compute engine but a query engine, among other things, before he drew our attention to the complexities of a typical distributed ETL workflow, highlighting shuffles, sorts, partitions, stages, and whatnot.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*39dRJgFZGRPf0hsU1Whxhg.jpeg" /><figcaption>Lester Martin highlighting the concerns with distributed processing — Image from the presentation taken by the author.</figcaption></figure><p>Lester also highlighted how Trino’s fault-tolerant execution model, along with the support for both SQL and Python, more or less, covered the linguistic reach of the data engineering community. Then there are rebels who want to use Rust; it is <a href="https://thenewstack.io/rust-eats-pythons-javas-lunch-in-data-engineering/">becoming a thing anyway</a>, but I’ll leave that to another day. One of the things Lester said was really funny and really true, which was that even the data engineers who primarily use Python DataFrames end up using the .sql() syntax more often than not. SQL is such a simple and beautiful language. He said:</p><blockquote>We all like SQL or we hate SQL, but we all use SQL.</blockquote><p>Lester focused on writing SQL, which may not be necessary with Starburst’s (Trino’s managed data lakehouse solution) Python libraries: PyStarburst and Ibis. These libraries provide PySpark-like syntax, lazy evaluation, and the Python code you write is automatically converted into SQL, which means that most of the heavy lifting is then handled by Trino itself. It is important to note, though, that PyStarburst is only supported by Starburst Galaxy and Starburst Enterprise. I won’t go into much more detail about Ibis, but a good starting point if you’re interested in learning about Trino would be the following two links:</p><ul><li><a href="https://www.youtube.com/watch?v=ZwaVZplVmVA">An Overview of the Starburst Trino Query Optimizer</a></li><li><a href="https://www.youtube.com/watch?v=JNI5_aU2fn4">Ibis: One Library To Query Any Backend</a></li></ul><h3>Streaming Analytics with Apache Polaris</h3><p>The next session was led by Kamesh Sampath, a Lead Developer Advocate at Snowflake, who discussed using a fully open-source stack to create a streaming analytics dashboard—the stack he used:</p><ul><li><a href="https://kafka.apache.org/">Kafka</a> — for publishing real-time events on topics</li><li><a href="https://www.risingwave.com/">RisingWave</a> — the streaming data lakehouse platform</li><li><a href="https://polaris.apache.org/">Polaris</a> — the catalog for registering and discovering tables</li><li><a href="https://iceberg.apache.org/">Iceberg</a> — the table format used for organizing Parquet files</li><li><a href="https://streamlit.io/">Streamlit</a> — the analytics application engine</li></ul><p>The session focused on the idea of avoiding being locked into any particular technology and utilizing open standards and open-source technologies whenever possible. Polaris is the base for Snowflake’s Open Catalog and Snowflake Horizon Catalog. The open-source version of Polaris can be run on a small, simple server anywhere. It is language-agnostic and vendor-neutral, and focuses on metadata portability.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*K4GvVWIGt7jtBPsqSZBIGQ.jpeg" /><figcaption>A rough sketch of how the solution was wired up — Image from the presentation taken by the author.</figcaption></figure><p>It was a fun, fully functional and working demo, which, if you’re a data engineer, you probably know is a rarity. Find more about Kamesh’s work <a href="https://github.com/kameshsampath">here</a>.</p><h3>Simplifying Data Pipelines with Tableflow</h3><p>The next session was led by Confluent’s <a href="https://www.linkedin.com/in/prerna-tiwari/">Prerna Tiwari</a>, who took us through a Kafka-native, serverless, managed stream-to-table data pipeline service, emphasizing that managing streams is very difficult with traditional batch-based or micro-batch-based pipelines. Tableflow takes a <a href="https://www.ssp.sh/brain/declarative/">declarative approach to data pipelines</a>, which is where the data engineering world is going. Databricks also recently announced <a href="https://www.databricks.com/product/data-engineering/lakeflow-declarative-pipelines">Lakeflow Declarative Pipelines</a>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*uFv2eYyaWMNBQj9CdHl6RQ.jpeg" /><figcaption>“Declarative” everything is the in thing right now — Image from the presentation taken by the author.</figcaption></figure><p>Tableflow, in effect, bypasses all other data lakehouse ingestion patterns and feeds data directly into your data lakehouse, built on Iceberg or Delta Lake tables, essentially getting rid of the process of converting topics into schemas, which, as you would know, can be quite messy with evolving schemas. Tableflow was announced last year, and <a href="https://www.confluent.io/blog/latest-tableflow/">it is GA as of March 2025</a>. Tableflow also integrates with <a href="https://www.databricks.com/dataaisummit/session/streaming-meets-governance-building-ai-ready-tables-confluent-tableflow">Databricks</a>, <a href="https://quickstarts.snowflake.com/guide/snowflake-confluent-tableflow-iceberg/index.html?index=..%2F..index#0">Snowflake</a>, <a href="https://www.starburst.io/blog/tableflow-confluent-starburst/">Trino</a>, and other data platforms. The following image compares open table formats, stream-to-table data lakehouse solutions, general-purpose ETL solutions, and Tableflow on a range of features and capabilities.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*hNH2A6ikWXL9dBMpWd8hCQ.jpeg" /><figcaption>Prerna Tiwari presenting a comparison between various data processing and pipelining options and highlighting why Tableflow is better than the others — Image from the presentation taken by the author.</figcaption></figure><h3>On Creating a Data Catalog by Accident</h3><p>This was probably one of the best, certainly the funniest, and most relatable sessions of the two days. <a href="https://www.linkedin.com/in/maylynhu/?originalSubdomain=au">May-Lyn Hu</a>, a Data Lead at Pet Circle, presented a session with the following title: Automate Your Metadata, Deliver a Data Catalogue (by accident).</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*4eXrtcjea1R91DrLzZueeA.jpeg" /><figcaption>Image from <a href="https://www.linkedin.com/in/maylynhu/?originalSubdomain=au">May-Lyn Hu</a>’s on presentation on Pet Circle’s journey from not having a catalog to having a in-house self-developed catalog, by accient— Image from the presentation taken by the author.</figcaption></figure><p>May-Lyn perfectly captured the need for a data catalog by highlighting all the unnecessary conversations, bottlenecks, and frustrations people have gone through just to get an answer to what would seem to the business a fairly simple question!</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*wz53Yv8wocAP_Or_oqc3kA.jpeg" /><figcaption><a href="https://www.linkedin.com/in/maylynhu/?originalSubdomain=au">May-Lyn Hu</a>’s perfect depiction of the day-to-day interacts of the business with the data engineering and analytics teams and the frustrations that come with that — Image from the presentation taken by the author.</figcaption></figure><p>This led to the team asking the question — is this (cataloging) a solved problem? The answer was yes, but it was cost-prohibitive, or at least it was going to be significantly if not prohibitively expensive, as she cited a $1500 seat price. So, the real answer was to conceptually break down what a catalog is from a first principles approach. The team did that and determined that it was just a collection of metadata that needed to be maintained in a structured manner. The team chose YAML. This didn’t come without challenges, but the team successfully delivered a data engineer-facing catalog.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*6a3OdIcBOOG9wNZ3jjVq2w.jpeg" /><figcaption>Bernie is begging for your action: “Please automate the config” — Image from the presentation taken by the author.</figcaption></figure><p>What about the business? Wasn’t the actual problem that the business doesn’t know where data assets are, what they mean, and whether they can be trusted? Yes, it was. So, the team decided to build a frontend for the business, too. This frontend gave the business a way to search data assets, access their quality and trust scores, look up ownership and stewardship metadata, among other things.</p><p>May-lyn stressed that it wasn’t an easy journey to build this. It took about a year to build the backend and the frontend, which essentially means that this is not for everyone, a point I completely agree with. However, I also think that businesses with the engineering skillset and drive to pull this off should consider doing so.</p><h3>Problems That You Only Face with Scale</h3><p>The next session focused on a cost-optimization project at Block following their acquisition of Afterpay. It’s a story of how the data team reduced the data egress cost for their cross-region data processing set up by over half a million USD per year. The following is from the <a href="https://code.cash.app/project-teleport">Cash App blog</a>. It should give you an idea of the scale:</p><blockquote><em>Kafka pipelines managed by the Afterpay data team process over 9 TB of data daily and deliver data to critical business domains such as Risk Decisioning, Business Intelligence and Financial Reporting via ~200 datasets. In the legacy design, duplicate events from Kafka’s “at least once” delivery required downstream cleansing by Data Lake consumers and late-arriving records added substantial re-processing overheads.</em></blockquote><p>The talk brought back some memories of dealing with scale (although nowhere close to this), reading about the journey of other teams on blogs like <a href="https://highscalability.com/brief-history-of-scaling-uber/">High Scalability</a>. To know more about the challenge and solution at Afterpay, head over to the blog post for <a href="https://code.cash.app/project-teleport">Project Teleport: Cost-Effective and Scalable Kafka Data Processing at Block</a>.</p><h3>How to Build an Agent</h3><p><a href="https://www.linkedin.com/in/geoffreyhuntley/">Geoffrey Huntley</a>, who’s an AI engineer at <a href="https://sourcegraph.com/">Sourcegraph</a>, came in with a single goal — to convince everyone that they could build their code-editing agents and it wasn’t that difficult. Because I had built something of that sort, although without using <a href="https://modelcontextprotocol.io/docs/getting-started/intro">MCP</a>, I already knew that point he was trying to make. Nevertheless, it was good to see how simple an agent can be — essentially, a self-referencing agent completion request in a while loop.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*rJ5AS25ExRUS738rrjXHlg.jpeg" /><figcaption>Geoffrey Huntley’s presentation on “How to Build an Agent” — Image from the presentation taken by the author.</figcaption></figure><p>Sourcegraph has released its agentic coding tool called Amp, which is a Cursor and Windsurf alternative. There are some key distinctions between, say, Cursor and Amp:</p><ol><li>With Cursor, you can choose models; with Amp, you cannot. Amp believes it chooses the best models.</li><li>Amp doesn’t have any limit on token usage. Your token usage decides how much you pay.</li><li>Amp is also collaborative — it lets you share threads with other team members and collaborate with them.</li></ol><p>It’s an interesting approach to code editors. If you want to know more about Amp, I can point you to the following two resources:</p><ul><li><a href="https://ampcode.com/manual">Owner’s Manual — Amp by Sourcegraph</a></li><li><a href="https://ampcode.com/agents-for-the-agent">Agents for the Agent</a></li></ul><p>The entire slide deck that Geoffrey used during his presentation is also available on <a href="https://ghuntley.com/how-to-build-an-agent/">his blog</a>.</p><h3>PromptSQL — Using LLMs for Query Optimization</h3><p>And then, towards the end of Day 2, on centre stage, Versent’s own Jake Kerr shared his thoughts and experiences in using LLMs for query optimization. Jake broke down the problem of query optimization and told the audience how it’s not going to go away. Rather, he walked everyone through a step-by-step plan of optimizing queries using structured prompts created for investigating query optimization issues, highlighting those issues right within your SQL IDE, and then fixing those issues one by one.</p><p>The focus of this session was to highlight that raw prompting alone cannot get you very far with query optimization. But if you break down the problem into manageable chunks and use something like <a href="https://www.langchain.com/langgraph">LangGraph</a> to manage the state of the activities, you might be able to do that.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*yXWpfVvS9rDTf90I6ll3tQ.jpeg" /><figcaption>Jake Kerr from Versent during his presentation on PromptSQL — Image from the presentation, taken by Dima Marr.</figcaption></figure><p>Query optimization is just another code analysis problem that can and should be broken down into multiple steps, each of which can be assigned to several agents, each with its own separate tools. <a href="https://www.uber.com/en-AU/blog/query-gpt/">Uber’s QueryGPT</a> is a good example of this. It uses several types of agents: intent agent, table agent, column pruning agent, among other intermediate agents.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*W9v6M-Wq1rz7vxJd0Q3gjw.jpeg" /><figcaption>Slide highlighting the companies that have already started using LLM-based query optimization in production — Image from the presentation taken by the author.</figcaption></figure><p>The Database Group at the University of Pennsylvania also found that <a href="https://arxiv.org/html/2411.02862v1">LLMs are unreasonably effective with query optimization</a>, although this wasn’t necessarily an agent-based assessment. Find more about the LLMSteer project presentation <a href="https://neurips.cc/media/neurips-2024/Slides/103605.pdf">here</a> and the <a href="https://github.com/peter-ai/LLMSteer">associated code here on GitHub</a>.</p><p>Jake concluded his remarks by sharing some time-tested query optimization principles and rules of thumb, which I’ll repeat here for you: join reordering based on real cardinality, predicate pushdown, set-based replacements, redundancy elimination, and data type mismatches.</p><p>All in all, these were two days well spent with the team, meeting former colleagues, ISV partners, and dabbling in the interesting tools and technologies, and hearing stories of daily struggles, architectural best practices, and challenges of scale, among other things. I think it’s an event worth attending, regardless of your level in data engineering, ML, or AI.</p><h3>Epilogue</h3><p>Jake’s session marked the end of the two-day data engineering, ML, and AI fun getaway for the team. I stayed back for the after-party at Melbourne Park for a while. And then left for an even cooler Friday night event at the University of Melbourne, where Dr. Danielle Holmes, who is the Women in Physics Lecturer at the Australian Institute of Physics, was presenting the Marie Curie lecture as part of the July Lectures in Physics—the topic of her presentation, Quantum Century: Unlocking the Universe’s Secrets and Shaping our Future.</p><p>Dr. Holmes talked about 100 years of quantum theory and walked us through the history of how the theory came to be. She spoke about its practical importance in explaining phenomena like <a href="https://www.scientificamerican.com/article/how-migrating-birds-use-quantum-effects-to-navigate/">why birds don’t get lost in long migrations</a>, how stars shine, among other things. She then ventured into the future of quantum computing and how it will be critical to solving many of the world’s seemingly unsolvable problems, such as climate change, nuclear energy, molecular simulation, advanced cryptography, and machine learning. Those in Sydney can still catch the same lecture as part of the National Science Week. Details <a href="https://www.scienceweek.net.au/event/2025-marie-curie-lectures-quantum-century-unlocking-the-universes-secrets-and-shaping-our-future-with-dr-danielle-holmes/north-wollongong/">here</a>.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*rPoXMOmpgUTtn9gOl5wcFg.jpeg" /><figcaption>Dr. Danielle Holmes presenting the Marie Curie lecture for 2025 as part of the July Lectures in Physics at the University of Melbourne — Image from the presentation taken by the author.</figcaption></figure><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=2fb63864f366" width="1" height="1" alt=""><hr><p><a href="https://medium.com/versent-tech-blog/versent-at-dataengbytes-2025-2fb63864f366">Versent at DataEngBytes 2025</a> was originally published in <a href="https://medium.com/versent-tech-blog">Versent Tech Blog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Software Engineering of the Future]]></title>
            <link>https://medium.com/versent-tech-blog/software-engineering-of-the-future-d868b79517bb?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/d868b79517bb</guid>
            <category><![CDATA[generative-ai]]></category>
            <category><![CDATA[software-engineering]]></category>
            <category><![CDATA[future-technology]]></category>
            <category><![CDATA[future-of-work]]></category>
            <category><![CDATA[alan-turing]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Tue, 24 Sep 2024 23:12:13 GMT</pubDate>
            <atom:updated>2024-09-24T23:12:13.110Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*FuWH0ekvpTDvpmF1" /><figcaption>Photo by <a href="https://unsplash.com/@yogidan2012?utm_source=medium&amp;utm_medium=referral">Daniele Levis Pelusi</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><h4>GENERATIVE AI</h4><h4>Speculations about the future of my profession and why writing will be a key skill in it</h4><p>The <a href="https://openai.com/index/chatgpt/">initial public release of ChatGPT</a> is coming up on two years in November — yes, it’s only been two years, and the software world has transformed quite a lot since then. Very competitive LLMs (Large Language Models) have been <a href="https://github.com/Hannibal046/Awesome-LLM">open-sourced</a>, <a href="https://www.cursor.com/">generative AI code editors and companions</a> have propped up, and the most important thing, in my opinion, is that <a href="https://www.nature.com/articles/d41586-023-02361-7">the Turing test has been broken</a>.</p><p>With these improvements in generative AI, a software engineer&#39;s job has already changed quite a bit. While visiting <a href="https://stackoverflow.com/">StackOverflow</a> and other community forums used to be of real value a few years ago, you can now ask a well-trained coding assistant to help you with your query. Meanwhile, various well-funded startups are working on AIs that will find bugs in your codebase, search for a fix, and deploy the fix, too.</p><p>This means the change isn’t coming soon; it’s here. However, it will probably take years for most organizations to leverage generative AI for code in a real sense. There are a few reasons for that.</p><p>Despite generative AI&#39;s exponential improvements, organizations will need to navigate limitations around cost, infrastructure availability, code safety, governance, and sovereignty, among other things. These limitations will hamper organizations that are not engineering-driven. This is why, while generative AI should, by its design and philosophy, democratize coding, at least for the first few years, it might not have that effect. In fact, it might actually end up having a sort of Reverse Robinhood effect, where everyone in the software engineering world is passively working towards bettering the models for <a href="https://mitsloan.mit.edu/ideas-made-to-matter/how-generative-ai-can-boost-highly-skilled-workers-productivity">people and organizations that are already good at software engineering</a>.</p><p>There’s also a general hesitation to adopt these new generative AI-based tools as so much of the landscape is constantly changing. While the constant change <a href="https://stackoverflow.blog/2018/01/11/brutal-lifecycle-javascript-frameworks/">isn’t a complete shock to software engineers</a>, the rate of change with generative AI models hasn’t been seen before with any other software technology. The fact that organizations are <a href="https://www.forbes.com/sites/joemckendrick/2024/01/12/executives-cautiously-hit-the-brakes-on-artificial-intelligence/">slowing down spending on AI learning and development</a> doesn’t help either.</p><p>While there is a broad acceptance of AI&#39;s benefits, many software engineers revolt at the idea of AI replacing their jobs. This makes me think that we, as a society, have seen this movie before. During the Industrial Revolution, the workforce violently rallied against the automation of the mills for a prolonged period, eventually landing in jails. Hopefully, it won’t come to that. At the same time, it is important to raise social, legal, and ethical concerns around AI to promote a constructive dialogue that informs its governance framework.</p><p>Organizations will overcome these issues before the wider society does. While these things get resolved, organizations will see pockets of generative AI adoption, if not across the board. There are obvious things where it will have a tremendous boost in productivity — a tool for automation more than anything else.</p><p>Tools like Ansible did that with configuration management, Terraform with cloud and on-premises infrastructure management, and Kubernetes with deploying scalable services in a platform-agnostic world. These are all, at an elementary level, tools for automation. Generative AI is a new category of automation tools that will work with all existing software tools, including the ones mentioned above. It will help you write software configuration, infrastructure orchestration, data pipelines, and frontends, among other things.</p><p>Writing software has always been about <em>instructions</em> to the electrical magic underneath the silicon. These instructions were first written in FORTRAN, BASIC, and C++. More recently, they have been written in languages like Python, Go, and Rust. Writing code has been getting easier all these years with the simple goal of abstracting machine complexity. Well, now, with English (or quite likely your native language) as the ultimate layer of abstraction is going to be your first layer of instruction. Before generative AI, it never was. But it was definitely given serious thought over seven decades ago when Turing wrote his seminal paper — <a href="https://courses.cs.umbc.edu/471/papers/turing.pdf">Computing Machinery and Intelligence</a>.</p><h3>The New Programming Skill</h3><p>Now that it’s becoming increasingly clear that natural language will be the medium of instruction for writing software in the near future, the question is: What will be the key skill for a software engineer of the future?</p><p>My speculation is that though understanding the internals of hardware and all the abstractions that bring us to the natural language interface will remain important, the most important skill will be <em>writing</em>, that is, <em>writing in your natural language</em>.</p><p>To the reader who’s thinking — well, isn’t that obvious ?— I’d say it’s true that writing efficient, performant software, finding the right constructs, applying the right order of instructions, and choosing the correct level of abstraction make a good software engineer. In a similar vein, in the future, writing good natural language instructions (or prompts) and structuring them so that the underlying model responds the way you intend it to respond will make a good natural-language software engineer. Still, it seems unclear how much of the job will be just plain old writing. Probably a lot. But writing alone won’t be enough.</p><p>A software engineer&#39;s job will become an amalgamation of product and engineering, among other things. It will entail thinking about software, product, user experience, and more and putting all those ideas into words, the importance of which Paul Graham highlights in <a href="https://www.paulgraham.com/words.html">one of his essays</a>:</p><blockquote><em>Putting ideas into words doesn’t have to mean writing, of course. You can also do it the old way, by talking. But in my experience, writing is the stricter test. You have to commit to a single, optimal sequence of words. Less can go unsaid when you don’t have tone of voice to carry meaning.</em></blockquote><p>This applies to software, too—the search for a single, optimal sequence of words. If you think about it, it is similar to how software is done now, but not in natural language. A software engineer&#39;s job is to find the right constructs, the right order of execution, and the right balance between CPU, database, network, and cache, among other things.</p><p>Similarly, software engineering in the future will become the art of finding the right word or sequence of words to communicate an idea to the machine or an AI agent, the essence of the Flaubertian idea of <em>le mot juste</em>. Who would have thought that Gustave Flaubert would also contribute, in his own little way, to the field of software engineering? Probably Alan Turing!</p><p><em>Thanks to Soumya, Corin Lawson, Matt Dudley, Hannah Ryan, Emerald Leung, Nishant Virmani, and Jatin Malik for reviewing the draft of this post.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=d868b79517bb" width="1" height="1" alt=""><hr><p><a href="https://medium.com/versent-tech-blog/software-engineering-of-the-future-d868b79517bb">Software Engineering of the Future</a> was originally published in <a href="https://medium.com/versent-tech-blog">Versent Tech Blog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Geospatial Indexing and Partitioning in Grid Systems]]></title>
            <link>https://medium.com/versent-tech-blog/geospatial-indexing-and-partitioning-in-grid-systems-b7b9c310bfb0?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/b7b9c310bfb0</guid>
            <category><![CDATA[data-engineering]]></category>
            <category><![CDATA[geospatial]]></category>
            <category><![CDATA[geodesy]]></category>
            <category><![CDATA[geospatial-data]]></category>
            <category><![CDATA[indexing]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Fri, 30 Aug 2024 05:06:26 GMT</pubDate>
            <atom:updated>2024-08-30T05:06:26.563Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*Fuxz6Wg3aHVALkDG" /><figcaption>Photo by <a href="https://unsplash.com/@chuttersnap?utm_source=medium&amp;utm_medium=referral">CHUTTERSNAP</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><h4>DATA ENGINEERING</h4><h4>A brief guide to methods for creating a global grid system for geospatial indexing and partitioning</h4><p>All deterministic data processing is fundamentally based on searching and sorting. To process data efficiently, one needs to search and sort it by <a href="https://towardsdatascience.com/the-art-of-discarding-data-4948ae3b3d14">discarding all the data they don’t need</a> — that’s what indexes and partitions are for. The details of how indexes and partitions are implemented are usually abstracted away from end users, but different types of indexes and partitions work better with different underlying data types.</p><p>Relational databases and data warehouse platforms support indexing types like <a href="https://en.wikipedia.org/wiki/B-tree">B-tree</a>, <a href="https://en.wikipedia.org/wiki/B%2B_tree">B+ tree</a>, <a href="https://www.postgresql.org/docs/8.1/gist.html">GiST</a>, <a href="https://www.postgresql.org/docs/current/spgist.html">SP-GiST</a>, and <a href="https://hakibenita.com/postgresql-hash-index">Hash</a>. Some work better for point lookup queries on numbers and short text fields, while others work better for full-text searches, and so on. When it comes to geospatial data, one of the popular index types is the <a href="https://en.wikipedia.org/wiki/R-tree">R-tree</a>, about which the PostgreSQL documentation says the following:</p><blockquote><em>R-Trees break up data into rectangles, and sub-rectangles, and sub-sub rectangles, etc. It is a self-tuning index structure that automatically handles variable data density, differing amounts of object overlap, and object size.</em></blockquote><p>This has many variations in <a href="https://en.wikipedia.org/wiki/K-d_tree">kd-tree</a>, <a href="https://en.wikipedia.org/wiki/Quadtree">Quadtree</a>, <a href="https://en.wikipedia.org/wiki/Octree">Octree</a>, and other indexing methods that help you index geospatial data. These indexing methods are included in the database&#39;s low-level implementation, meaning that adding, removing, or changing the data can lead to expensive tree rebalancing and index rebuilding. This can both increase the computing cost and affect the performance of ongoing queries. Because of this limitation, various data processing systems have developed pluggable <a href="https://benfeifke.com/posts/geospatial-indexing-explained/">geospatial indexing and partitioning systems</a> that decouple indexing from core database operations. These are based on the idea that the surface of the Earth can be represented as a grid system. From Uber’s engineering blog:</p><blockquote><em>A global grid system usually requires at least two things: a map projection and a grid laid on top of the map. A map projection is needed for going from a three-dimensional location on Earth to a two dimensional point on a map. A grid is then overlaid on the map, forming a global grid system.</em></blockquote><p>In this article, I will outline some of these systems and their background, focusing on <a href="http://s2geometry.io/">Google’s S2</a>, <a href="https://h3geo.org/">Uber’s H3</a>, and <a href="https://learn.microsoft.com/en-us/bingmaps/articles/bing-maps-tile-system">Bing Maps’ Quadbin</a> indexing methods. But before we get started, let’s take a quick look at how indexing and partitioning work in geospatial data.</p><h3>Indexing vs. Partitioning in Grid Systems</h3><p>In the traditional database sense, indexing means creating a pointer table based on specific columns ordered a certain way. Partitioning, in the same context, is usually the physical separation of files under the same logical construct of a table.</p><p>While many types of indexing methodologies are available in the geospatial world, which we will discuss in a while, it is important to note that they are not <em>really </em>indexing techniques. Instead, they’re partitioning techniques.</p><p>Why do I say this? Indexes usually help you order data in a certain way to facilitate faster retrieval, while partitions help you segregate data for the same purpose. When it comes to geospatial data, what we’re doing with popular geospatial indexing techniques like S2, H3, and Quadbin is partitioning, not indexing.</p><p>That’s why Uber H3’s documentation says that H3 is a Hexagonal Hierarchical Spatial Index, which “enables users to partition the globe into hexagons for more accurate analysis.” Let’s understand some of the grid system partitioning methods now.</p><h3>Grid System Indexing and Partitioning Methods</h3><p>While many indexing (read partitioning) methods exist, four stand out because of their wide adoption and support in many data processing and visualization systems. Three of these four, S2, H3, and Quadbin, came from three large companies trying to solve maps for navigation, while Geohash was created by Canonical&#39;s CTO, Gustavo Niemeyer.</p><h4>Geohash</h4><p>It’s the oldest of the four methods we discussed. Geohash&#39;s author made it public in 2008. It is based on the equirectangular projection, which was invented thousands of years ago. It has its shortcomings, as do the others that we will discuss. After all, we are trying to map an irregularly shaped ellipsoid (our beautiful planet) to a perfect rectangle. What could go wrong? Indulge yourself in watching this short video that explains the trouble with projections.</p><iframe src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FkIID5FDi2JQ%3Ffeature%3Doembed&amp;display_name=YouTube&amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DkIID5FDi2JQ&amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FkIID5FDi2JQ%2Fhqdefault.jpg&amp;key=d04bfffea46d4aeda930ec88cc64b87c&amp;type=text%2Fhtml&amp;schema=youtube" width="854" height="480" frameborder="0" scrolling="no"><a href="https://medium.com/media/c24e88a33fffbe5d8fea4e7d7ac7f813/href">https://medium.com/media/c24e88a33fffbe5d8fea4e7d7ac7f813/href</a></iframe><p>Nevertheless, it is one of the most popular methods for implementing proximity searches. It is a system for encoding location in strings, not numbers while creating a hierarchical quadtree grid system. It is used for place identification, but a more popular, human-readable <em>sort of </em>hash would be what3words, a service that converts any lat long pair to a square of area 3m² identifiable with three words like ///candy.wire.dull. That’s our Versent Melbourne office, by the way.</p><p>Geohash can go to an arbitrary number of levels. When you zoom in, you break the squares into four, forcing them to be arranged into a tree that will later be used for search and traversal. This is where a Z-order curve comes in. It orders these smaller squares, automatically enabling you to perform a breadth-first search on your whole tree.</p><blockquote><em>Interestingly, Z-order curves are also used to reorganize data on disk to speed up queries. This feature is natively implemented in the Delta Lake file format. The </em><a href="https://docs.databricks.com/en/delta/data-skipping.html#delta-zorder"><em>Databricks documentation</em></a><em> says the following:</em></blockquote><blockquote><em>Z-ordering is a </em><a href="https://en.wikipedia.org/wiki/Z-order_curve"><em>technique</em></a><em> to colocate related information in the same set of files. This co-locality is automatically used by Delta Lake on Databricks data-skipping algorithms. This behavior dramatically reduces the amount of data that Delta Lake on Databricks needs to read.</em></blockquote><blockquote><em>The </em><a href="https://blog.cloudera.com/speeding-up-queries-with-z-order/"><em>Cloudera blog</em></a><em> also talks about using Z-order curves to speed up queries, and so do </em><a href="https://hudi.apache.org/blog/2021/12/29/hudi-zorder-and-hilbert-space-filling-curves/"><em>Apache Hudi</em></a><em>, </em><a href="https://aws.amazon.com/blogs/database/z-order-indexing-for-multifaceted-queries-in-amazon-dynamodb-part-1/"><em>AWS DynamoDB</em></a><em>, and </em><a href="https://aws.amazon.com/blogs/database/amazon-aurora-under-the-hood-indexing-geospatial-data-using-z-order-curves/"><em>Aurora MySQL</em></a>,<em> among others</em></blockquote><p>However, using Geohash (and Z-order curves to traverse) doesn’t come without challenges, mainly because of the choice of projection and the grid shape used. One major drawback is significant variations in squares at different latitudes. There are also limitations around inaccuracies in proximity searches closer to the boundary of the geohash square, i.e., points closer to the geohash boundary will end up having quite different hashes, hence unrecognizable.</p><h4>Google’s S2</h4><p>Geospatial processing can be compute-intensive. To mitigate this issue, Google’s open-source indexing library, S2, assumes Earth as a perfect mathematical sphere, making specific geospatial processing easier and less expensive.</p><p><a href="https://opensource.googleblog.com/2017/12/announcing-s2-library-geometry-on-sphere.html">Announcing the S2 Library: Geometry on the Sphere</a></p><p>S2 is the only one of the three methods we’re discussing that represents the data in a <a href="http://s2geometry.io/devguide/s2cell_hierarchy.html">hierarchy of quadrilaterals represented on a three-dimensional sphere</a>. Because of this, S2 shows less distortion towards the extremes of the northern and southern hemispheres.</p><p>S2 offers a wide range of shapes and sizes that you can use to partition your grid using spherical caps or discs, latitude-longitude rectangles, polylines, and polygons. This makes S2 a good choice for implementing proximity-based searches, simplifying geographies, and quick lookups.</p><p>Note that the percentage of area signifies the accuracy of a partitioning method it can successfully represent under all its partitions combined. S2 uses a function based on the continuous and space-filling Hilbert curve that visits every point in a unit square to break a sphere into smaller chunks; it’s how you traverse (and search) the grid on your map.</p><p><a href="https://www.pyblog.xyz/spatial-index-space-filling-curve">Spatial Index: Space-Filling Curves</a></p><p>This was worth mentioning because grid traversal, although with slightly different techniques, reappears when discussing H3 and Quadbin. For a deeper look, read the following:</p><ul><li><a href="https://cloud.google.com/bigquery/docs/grid-systems-spatial-analysis">Grid System Spatial Analysis in BigQuery</a></li><li><a href="https://blog.christianperone.com/2015/08/googles-s2-geometry-on-the-sphere-cells-and-hilbert-curve/">Google’s S2, geometry on the sphere, cells, and Hilbert curve</a></li></ul><p>Google’s S2 suffers, however, from the same drawbacks as Geohash. The grid cells aren’t equidistant from all their neighbours, which makes things difficult in many use cases. At least there’s a guarantee that a data point will always land in only one grid cell, unlike Uber’s H3, which can land in multiple cells, which we’ll come to a bit later. S2, in many ways, is better than Geohash but doesn’t have the portability, interoperability, and widespread support of Geohash, which is why it might be a hard sell.</p><h4>Quadbin</h4><p>Bing Maps launched in the same year as Google Maps, i.e., 2005, but took a different approach to mapping the globe when it used Quadtrees (conceptually quite similar to geohashes). Bing Maps’ Quadtree goes up to 23 levels. This implementation by Bing Maps is called Quadkey.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/851/1*DS9sYBhAXE4g6ZkeOnMVDg.png" /><figcaption>Quadkey indices and zoom levels — Image from <a href="https://learn.microsoft.com/en-us/azure/azure-maps/zoom-levels-and-tile-grid?tabs=csharp#quadkey-indices">Azure Maps documentation</a></figcaption></figure><p>Quadbin is a hierarchical index based on Quadkey. It stores grid cell values in a 64-bit number and goes a bit further concerning the level of detail to 27 levels, which can point to areas less than 1m². Bing Maps uses the Mercator projection, which comes with its challenges. While Bing Maps will be deprecated sometime in 2025, Azure Maps will continue to use the Spherical Mercator projection coordinate system.</p><h4>Uber’s H3</h4><p>While the rest of the systems we’ve discussed use rectangles and squares to create a mappable global grid, H3 uses hexagons, assuming the globe to be a regular icosahedron instead of a perfect sphere. This differs from Geohash (which uses an equirectangular projection), S2 (which uses a stereographic projection), and Quadbin (which uses a Mercator projection). The original blog post announcing the open-sourcing of H3 says the following about this choice:</p><blockquote><em>This projects from Earth as a sphere to an icosahedron, a twenty-sided platonic solid. An icosahedron-based map projection results in twenty separate two-dimensional planes rather than a single plane. The icosahedron can be unfolded in many ways, producing a two-dimensional map each time. H3, however, does not unfold the icosahedron to build its grid system, and instead lays its grid out on the icosahedron faces themselves, forming a geodesic discrete global grid system.</em></blockquote><p>I’ve already discussed H3 in another piece covering geospatial data ingestion, indexing, and analysis in Snowflake. I also discuss why hexagons have the edge when calculating grid distances, proximity searches, and navigation.</p><p><a href="https://medium.com/versent-tech-blog/getting-started-with-geospatial-data-in-snowflake-d4272e689200">Getting Started with Geospatial Data in Snowflake</a></p><p>Ben Feifke’s blog post about these systems has a section at the end that compares these three indexing/partitioning systems and discusses <a href="https://benfeifke.com/posts/geospatial-indexing-explained/#when-to-use-which-tool">when to use which one</a>.</p><h3>Support for Geohash, S2, H3, and Quadbin</h3><p>These indexing and partitioning systems are available in most databases and data warehouse systems. They can be used independently as they can be decoupled from the database&#39;s core indexing implementation. For example, Foursqaure’s layer library offers <a href="https://docs.foursquare.com/analytics-products/docs/layer-h3">H3</a>, <a href="https://docs.foursquare.com/analytics-products/docs/layer-s2">S2</a>, and <a href="https://docs.foursquare.com/analytics-products/docs/layer-hexbin">Hexbin</a> layers that can be used in analytical calculations or as overlays on top of existing geospatial visualizations.</p><p>The support for different grid systems, especially H3 and Geohash, is exceptionally promising across the data ecosystem. <a href="https://blog.rustprooflabs.com/2022/04/postgis-h3-intro">PostgreSQL</a>, <a href="https://www.youtube.com/watch?v=O184y7A-Hdg">Snowflake</a>, <a href="https://docs.databricks.com/en/sql/language-manual/sql-ref-h3-geospatial-functions.html">Databricks</a>, <a href="https://clickhouse.com/docs/en/sql-reference/functions/geo/h3">ClickHouse</a>, and even newer databases like <a href="https://tech.marksblogg.com/h3-duckdb-qgis.html">DuckDB</a> have extensions or built-in libraries. This means you can use the indexing and partitioning system that suits your use case.</p><h3>Conclusion</h3><p>My experience working with geospatial indexing and partitioning systems has been with PostgreSQL with PostGIS (the OG, I’d say), <a href="https://towardsdatascience.com/handling-geospatial-data-in-aws-a82ae364f80c">Redshift</a>, and MySQL, among others. Recently, though, I’ve been spending quite a bit of time exploring Snowflake for geospatial data, and it’s becoming more promising by the day. Check out the end of <a href="https://docs.snowflake.com/en/release-notes/2024/other/2024-06-28-geospatial-h3-functions-ga">June 2024 update</a> regarding new H3 functions to know more.</p><p>I hope this overview of geospatial indexing and partitioning systems has been useful to you. I’m working on my next post in this series exploring geospatial data, which will examine an actual use case using one of the database systems mentioned above.</p><p><em>Thanks to </em><a href="https://www.linkedin.com/in/kklutchkov/"><em>Kroum Kluctchkov</em></a><em>, </em><a href="https://www.linkedin.com/in/dima-marr-2ab26972/"><em>Dima Marr</em></a><em>, and </em><a href="https://www.linkedin.com/in/haymangahuja/"><em>Haymang Ahuja</em></a><em> for providing feedback on this blog post.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=b7b9c310bfb0" width="1" height="1" alt=""><hr><p><a href="https://medium.com/versent-tech-blog/geospatial-indexing-and-partitioning-in-grid-systems-b7b9c310bfb0">Geospatial Indexing and Partitioning in Grid Systems</a> was originally published in <a href="https://medium.com/versent-tech-blog">Versent Tech Blog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Nusrat’s Chain of Light]]></title>
            <description><![CDATA[<div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/khayaliya/nusrats-chain-of-light-b5a98378a3d3?source=rss-13d513db037------2"><img src="https://cdn-images-1.medium.com/max/2600/0*1GbWGAAPceFpldYn" width="4160"></a></p><p class="medium-feed-snippet">In the anticipation of Nusrat Fateh Ali Khan&#x2019;s lost album from 1990</p><p class="medium-feed-link"><a href="https://medium.com/khayaliya/nusrats-chain-of-light-b5a98378a3d3?source=rss-13d513db037------2">Continue reading on Khayaliya »</a></p></div>]]></description>
            <link>https://medium.com/khayaliya/nusrats-chain-of-light-b5a98378a3d3?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/b5a98378a3d3</guid>
            <category><![CDATA[qawwali]]></category>
            <category><![CDATA[khayal]]></category>
            <category><![CDATA[peter-gabriel]]></category>
            <category><![CDATA[indian-classical-music]]></category>
            <category><![CDATA[nusrat-fateh-ali-khan]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Fri, 23 Aug 2024 08:37:25 GMT</pubDate>
            <atom:updated>2024-08-24T07:53:00.547Z</atom:updated>
        </item>
        <item>
            <title><![CDATA[Getting Started with Geospatial Data in Snowflake]]></title>
            <link>https://medium.com/versent-tech-blog/getting-started-with-geospatial-data-in-snowflake-d4272e689200?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/d4272e689200</guid>
            <category><![CDATA[geospatial]]></category>
            <category><![CDATA[snowflake]]></category>
            <category><![CDATA[versent]]></category>
            <category><![CDATA[gis]]></category>
            <category><![CDATA[geospatial-data]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Wed, 07 Feb 2024 01:39:23 GMT</pubDate>
            <atom:updated>2024-02-07T01:39:23.781Z</atom:updated>
            <content:encoded><![CDATA[<figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/0*8zoUD83XRrzaB-eK" /><figcaption>Photo by <a href="https://unsplash.com/@martinirc?utm_source=medium&amp;utm_medium=referral">José Martín Ramírez Carrasco</a> on <a href="https://unsplash.com?utm_source=medium&amp;utm_medium=referral">Unsplash</a></figcaption></figure><h4>GEOSPATIAL</h4><h4>A brief overview of Snowflake’s native geospatial capabilities and how to extend them</h4><h3>Background</h3><p><em>Skip this section if you understand the fundamentals of geospatial data types and spatial reference systems.</em></p><p>Numeric, string, boolean, and other data types have reference systems. Numeric data types follow the decimal number system. String data types adhere to a character set (computationally represented by the hexadecimal system), while boolean data types follow the binary number system, and so on. You can efficiently perform basic manipulations and operations on these data types. You can also run an analysis on top of the data held in these data types. The underlying reference system powers all these operations performed using the data type.</p><p>Although geolocation data, at its most basic level, is just a pair of numbers — latitude and longitude — which can be stored in numeric data types, it is hard to perform real-world calculations on those numbers. The decimal number system alone isn’t enough for storing information about a place on Earth. A system analysing geospatial data must know that the two points lie on a <em>specific surface</em>, even to calculate something as simple as the distance between two points. That’s where a spatial reference system comes into the picture.</p><p>Now, most databases use two types of reference systems to map these points to a surface. The first assumes that the planet can be represented as a flat surface on a Cartesian or Euclidian plane, giving birth to the GEOMETRY data type. The second assumes that the planet can be modelled as a perfect sphere, making GEOGRAPHY data type the other option.</p><p>Using GEOGRAPHY data type makes more sense when calculating distances across the planet&#39;s surface. On the other hand, the GEOMETRY data type is more useful when visualising data on a two-dimensional map of a geographical area. The good thing is that, if your database allows, you can convert one data type to another using standard built-in functions. That’s it on the brief background. Let’s look at how Snowflake deals with geospatial data!</p><h3>Working with Geospatial Data in Snowflake</h3><p>In June 2021, Snowflake’s first geospatial data type — GEOGRAPHY— went GA. This May, they’ve also made the GEOMETRY datatype GA will now help geospatial data analysts and engineers have more flexibility in consuming the data through analysis and visualisation. To understand more about geometric representations in Snowflake, read <a href="https://www.snowflake.com/blog/blog-getting-started-geometry-data/">this post</a> on the Snowflake blog.</p><h4>Ingesting Geospatial Data into Snowflake</h4><p>Geospatial data can come in all shapes and sizes. The files you source into Snowflake may have a structure that isn’t very different from a CSV file, with the only difference being that the source files will have latitudes and longitudes. If the data is already processed and shaped up as having a geospatial representation, you can expect it to be in a well-known text format, where you can represent a pair of latitude and longitude as a POINT and a group of latitude and longitude pairs as a MULTILINE or a POLYGON.</p><p>There are also other standards and formats that Snowflake supports for ingestion, such as GeoJSON. You can look at the complete list <a href="https://docs.snowflake.com/en/sql-reference/data-types-geospatial#label-data-types-geospatial-io-formats">here</a>. There’s a difference between WKT/WKB and GeoJSON, which is worth repeating here from the Snowflake documentation —</p><blockquote><em>The WKT and WKB standards specify a format only; the semantics of WKT/WKB objects depend on the reference system — for example, a plane or a sphere.</em></blockquote><blockquote><em>The GeoJSON standard, on the other hand, specifies both a format and its semantics: GeoJSON points are explicitly WGS 84 coordinates, and GeoJSON line segments are supposed to be planar edges (straight lines).</em></blockquote><p>This goes back to the context I gave you about reference systems. Anyway, the long and short of it is that your geospatial data can be in various formats to be fit for ingestion. For successful ingestion, you must align or transform (using reference system transform functions like ST_TRANSFORM) your Snowflake table structure and data types according to the source data format and reference system.</p><h4>Geospatial Data Processing in Snowflake</h4><p>You may have several use cases for processing geospatial data, i.e., calculating proximity, adjacency, and overlap between points, lines, and polygons, among other things. Geospatial data aggregation is one of the most prominent use cases. Like how you calculate total sales, cumulative sales, etc., for an e-commerce business, you may need to aggregate your geospatial data from a point to a multipoint, a linestring, a polygon, and more, for a logistics business.</p><p>Here’s a simple example that comes from a free Snowflake marketplace <a href="https://app.snowflake.com/marketplace/listing/GZSTZBWGAEV/ahead-chicago-divvy-bike-station-status">data share</a> that has data about bike stations in Chicago:</p><pre>CREATE TABLE ANALYTICS.PUBLIC.STATION_INFO_FLATTEN_SOS AS<br><br>WITH bike_station_locations AS<br>(SELECT REPLACE(station_id,&#39;&quot;&#39;,&#39;&#39;) station_id,<br>        REPLACE(name,&#39;&quot;&#39;,&#39;&#39;) station_name,<br>        ST_MAKEPOINT(lon, lat) point<br>   FROM CHICAGO_DIVVY_BIKE_STATION_STATUS.PUBLIC.STATION_INFO_FLATTEN<br>  LIMIT 5)<br>  <br>SELECT ST_COLLECT(point) point<br>  FROM bike_station_locations;<br><br>SELECT * FROM ANALYTICS.PUBLIC.STATION_INFO_FLATTEN_SOS;</pre><p>The above piece of SQL reads the latitude and longitude values for the bike station location and converts them into a POINT data type. Once that’s done, it uses the ST_COLLECT aggregation function to convert several points to a MULTIPOINT, which sort of works like the LISTAGG function in Oracle or the GROUP_CONCAT function in MySQL. Here’s what the output will look like:</p><pre>{<br>  &quot;type&quot;: &quot;FeatureCollection&quot;,<br>  &quot;features&quot;: [<br>    {<br>      &quot;type&quot;: &quot;Feature&quot;,<br>      &quot;properties&quot;: {},<br>      &quot;geometry&quot;: {<br>        &quot;coordinates&quot;: [<br>          [<br>            -87.673688,<br>            41.871262<br>          ],<br>          [<br>            -87.6601406209,<br>            41.98974251144<br>          ],<br>          [<br>            -87.5970051479,<br>            41.81409271048<br>          ],<br>          [<br>            -87.57632374763489,<br>            41.78091096424803<br>          ],<br>          [<br>            -87.707322,<br>            41.921525<br>          ]<br>        ],<br>        &quot;type&quot;: &quot;MultiPoint&quot;<br>      }<br>    }<br>  ]<br>}</pre><p>Now, if you paste this into any GeoJSON visualiser like <a href="http://geojson.io">geojson.io</a>, you can see the points on a map, as shown in the image below:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*sd4Q7z3j6VsQx6m0w5TCtw.png" /><figcaption>Plotting five bike stations using Multipoint geospatial data type — Image by author</figcaption></figure><p>There’s a lot more to geospatial data processing. I’ll cover that in more detail in an upcoming post.</p><h4>Geospatial Search with Search Optimisation Service</h4><p>Most relational databases that support geospatial data also have constructs for geospatial indexing. Snowflake, on the other hand, doesn’t support indexes. Instead, you can use Snowflake’s Search Optimization Service (SOS) to speed up your geospatial search queries. Here’s how you enable search optimisation on a column:</p><pre>ALTER TABLE ANALYTICS.PUBLIC.STATION_INFO_FLATTEN_SOS <br>  ADD SEARCH OPTIMIZATION ON GEO(point);</pre><p>Like PostgreSQL, which supports search functionality using specific geospatially-indexed functions, Snowflake does that for <a href="https://docs.snowflake.com/en/user-guide/search-optimization-service#supported-predicates-with-geospatial-functions">some supported functions</a>. Currently, geospatial search optimisation only works for columns with GEOGRAPHY data type. GEOMETRY data type is not supported yet.</p><p>Snowflake’s Search Optimization Service helps speed up point look-up queries that answer questions like:</p><ul><li>Whether a geospatial object is fully contained within another geospatial object — a state within a country, a water body within a county, etc., using the ST_WITHIN geospatial function.</li><li>Whether two geospatial objects share any common portion — this turns out to be extremely useful and essential in mapping houses, localities, and communities using the ST_INTERSECTS geospatial function.</li></ul><p>This service works better if the data you seek is clustered together. If you don’t already have a geographical distribution column, you can use the ST_GEOHASH function to leverage the proximity grouping method that a geohash employs.</p><h4>Spatial Partitioning and Indexing with H3</h4><p>While the Search Optimisation Service enables point look-up queries, many more use cases require a space partitioning system as a discrete global grid. These use cases revolve mostly around location-based aggregation of moving vehicles, IoT devices, traffic lights, and other objects.</p><p>Although H3 (Hexagonal Hierarchical Geospatial Indexing System) is more popularly known as an indexing system, it aligns more with partitioning data, i.e., H3 partitions data like a range-based or a key-based partition scheme partitions your data in a relational database. The key difference here is that your partitions are hexagonal boundaries, which have benefits you can read about in the following blog — <a href="https://www.kontur.io/blog/why-we-use-h3/">Why We Use H3</a>.</p><blockquote><em>A quick note on Hexagons — and why they’re better to map to a sphere.</em></blockquote><blockquote><em>When you want to map the Earth, you want to devise the most efficient way of spatially partitioning it into blocks or sections that don’t leave any space uncovered. This can be achieved by creating a tessellation, a tiling, of polygons, which can either be triangles, squares, or hexagons. These are the only three shapes that can fill a plane or a spherical surface without leaving a gap. A hexagon is </em><a href="https://gis.stackexchange.com/questions/82362/what-are-the-benefits-of-hexagonal-sampling-polygons"><em>the most complex polygon that can fill a plane</em></a><em>, which makes it the most accurate when mapping to a spherical-ish surface because of the curvature.</em></blockquote><p>Read the following piece for more info on Hexagonal grids:</p><p><a href="https://strimas.com/post/hexagonal-grids/">Fishnets and Honeycomb: Square vs. Hexagonal Spatial Grids | Matt Strimas-Mackey</a></p><p>The value of an H3 index is subjective; it depends on two fixed values and one variable value, i.e., latitude, longitude, and the H3 index resolution. When you pass these three things to a function that calculates the H3 value, you uniquely identify a hexagon on the discrete global grid of hexagons.</p><blockquote>Using the values derived from the H3 spatial partitioning system, you can map grid hexagons ranging from approximately one m² to about 4.3 million km², represented with number of hexagons ranging from 569.7 trillion to 122, with resolutions ranging from 15 to 0.</blockquote><p>Here’s the grid at resolution 10 with the highlighted hexagon where our Melbourne office is located.</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*PHRWLdojpOxDfRkWNMl2MQ.png" /><figcaption>Hexagon with a resolution = 10, which maps to the Versent Melbourne office — Image by author.</figcaption></figure><p>Working with geospatial data, you’ll figure out that the visualisation aspect only comes at the end of the pipeline. Before that, it’s going to be, more or less, the same kind of data processing that gets done with any other form of data. Here’s a quick peek into how H3 values are created using latitude and longitude values at two different resolutions (for two different levels of aggregation):</p><pre>WITH bike_station_locations AS<br>(SELECT REPLACE(station_id,&#39;&quot;&#39;,&#39;&#39;) station_id,<br>        REPLACE(name,&#39;&quot;&#39;,&#39;&#39;) station_name,<br>        ST_MAKEPOINT(lon, lat) point,<br>        H3_POINT_TO_CELL(ST_MAKEPOINT(lon, lat),12) h3_index_12,<br>        H3_POINT_TO_CELL(ST_MAKEPOINT(lon, lat),10) h3_index_10<br>   FROM CHICAGO_DIVVY_BIKE_STATION_STATUS.PUBLIC.STATION_INFO_FLATTEN<br>  LIMIT 10)<br>  <br>SELECT * <br>  FROM bike_station_locations;</pre><p>Here’s what the output of this query would look like:</p><figure><img alt="" src="https://cdn-images-1.medium.com/max/1024/1*BzDzvt1FpxfPfTD3qLLOIg.png" /><figcaption>Output showing numeric H3 index values at different resolutions using the H3_POINT_TO_CELL function in Snowflake — Image by author.</figcaption></figure><p>The H3 spatial partitioning system not only helps address the problem of spatial data optimisation but also solves many practical problems, like identifying taxi supply and demand clusters, pinpointing accident-prone areas, and many more.</p><p>Snowflake’s geospatial capabilities are well-placed to solve any geospatial ingestion, processing, analysis, and visualisation problem that might come your way. As described in the article, Snowflake extensively supports geospatial data types, specialised indexing, partitioning systems, and more. While this was an introduction to the geospatial capabilities within Snowflake, in a couple of upcoming blogs, I’ll cover a step-by-step approach to ingesting, processing, optimising, converting, and visualising spatial data and how Snowflake handles the visualisation aspect of geospatial data with Streamlit. Meanwhile, you can <a href="https://blog.streamlit.io/how-to-analyze-geospatial-snowflake-data-in-streamlit/">check out this blog on Streamlit’s website</a>, which discusses exactly that.</p><h3>Further Reading</h3><ul><li><a href="https://carto.com/blog/from-postgresql-to-snowflake">Migrating Spatial Data from PostgreSQL to Snowflake</a></li><li><a href="https://docs.snowflake.com/en/sql-reference/functions-geospatial">Geospatial Functions in Snowflake</a></li><li><a href="https://h3geo.org/docs/core-library/restable/#cell-counts">Cell Counts and Resolutions in H3 Indexing</a></li><li><a href="https://blog.rustprooflabs.com/2022/04/postgis-h3-intro">H3 Indexes and PostGIS</a></li><li><a href="https://quickstarts.snowflake.com/guide/getting_started_with_geospatial_geography/index.html?index=..%2F..index#0">Snowflake Quickstart — Getting Started with Geospatial</a></li><li><a href="https://www.youtube.com/watch?v=yAYMyqZmZOk">The Defence Context, Geospatial Standards for Data Science — Dr Paul. Cripps</a></li><li><a href="https://slate.com/technology/2015/07/hexagons-are-the-most-scientifically-efficient-packing-shape-as-bee-honeycomb-proves.html">The Miraculous Space Efficiency of Honeycomb</a></li><li><a href="https://www.youtube.com/watch?v=thOifuHs6eY">Hexagons are the Bestagons</a>, and <a href="https://www.youtube.com/watch?v=4zWDLKWmBnE">Hexagons are NotSoGreatAgons</a></li></ul><h3>Credits</h3><p><em>Thanks to </em><a href="https://www.linkedin.com/in/matt-dudley-157688180/?originalSubdomain=au"><em>Matt Dudley</em></a><em>, </em><a href="https://www.linkedin.com/in/sagarsatishkulkarni/?originalSubdomain=au"><em>Sagar Kulkarni</em></a><em>, and </em><a href="https://www.linkedin.com/in/emeraldleung/"><em>Emerald Leung</em></a><em> for generously taking time out and providing feedback on this blog.</em></p><img src="https://medium.com/_/stat?event=post.clientViewed&referrerSource=full_rss&postId=d4272e689200" width="1" height="1" alt=""><hr><p><a href="https://medium.com/versent-tech-blog/getting-started-with-geospatial-data-in-snowflake-d4272e689200">Getting Started with Geospatial Data in Snowflake</a> was originally published in <a href="https://medium.com/versent-tech-blog">Versent Tech Blog</a> on Medium, where people are continuing the conversation by highlighting and responding to this story.</p>]]></content:encoded>
        </item>
        <item>
            <title><![CDATA[Meend Masterclass in Amsterdam]]></title>
            <description><![CDATA[<div class="medium-feed-item"><p class="medium-feed-image"><a href="https://medium.com/khayaliya/meend-masterclass-in-amsterdam-e02fa207cfb0?source=rss-13d513db037------2"><img src="https://cdn-images-1.medium.com/max/2600/1*nTi8rSrx0rsoM9AE2y0xQA.jpeg" width="6000"></a></p><p class="medium-feed-snippet">Sayeeduddin Dagar&#x2019;s beautiful musical interpretation of the flow of Ganga and Yamuna, 1999</p><p class="medium-feed-link"><a href="https://medium.com/khayaliya/meend-masterclass-in-amsterdam-e02fa207cfb0?source=rss-13d513db037------2">Continue reading on Khayaliya »</a></p></div>]]></description>
            <link>https://medium.com/khayaliya/meend-masterclass-in-amsterdam-e02fa207cfb0?source=rss-13d513db037------2</link>
            <guid isPermaLink="false">https://medium.com/p/e02fa207cfb0</guid>
            <category><![CDATA[khayal]]></category>
            <category><![CDATA[dhrupad]]></category>
            <category><![CDATA[indian-classical-music]]></category>
            <category><![CDATA[music]]></category>
            <category><![CDATA[classical-music]]></category>
            <dc:creator><![CDATA[Kovid Rathee]]></dc:creator>
            <pubDate>Sat, 29 Jul 2023 22:17:14 GMT</pubDate>
            <atom:updated>2023-07-29T22:56:18.777Z</atom:updated>
        </item>
    </channel>
</rss>