<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Cloud Blog</title><link>https://cloud.google.com/blog/</link><description>Cloud Blog</description><atom:link href="https://cloudblog.withgoogle.com/blog/rss/" rel="self"></atom:link><language>en</language><lastBuildDate>Wed, 07 Oct 2026 16:00:02 +0000</lastBuildDate><image><url>https://cloud.google.com/blog/static/blog/images/google.a51985becaa6.png</url><title>Cloud Blog</title><link>https://cloud.google.com/blog/</link></image><item><title>Introducing Google Cloud’s U4 compute: Enabling ultra-low latency trading</title><link>https://cloud.google.com/blog/topics/financial-services/ultra-low-latency-solution-with-u4-enables-high-velocity-trading/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In capital markets, operational success is built on predictable performance and low latency, where even tiny variations in packet timing alter queue priority and trade execution. At the same time, both exchange venues and their participants face mounting constraints in traditional on-premises data centers — from power and physical rack space limits to lengthy hardware procurement cycles. Increasingly, global capital markets seek the speed and determinism of physical co-location combined with the dynamic scalability and automation of the cloud.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At Google Cloud Next ‘26, we announced the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/networking/whats-new-in-cloud-networking-at-next26#:~:text=Ultra%20Low%20Latency%20Solution%20for%20financial%20exchanges"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Ultra Low Latency (ULL) Solution&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, now generally available, providing high-frequency trading workflows that run in the cloud with an ultra-low latency network and agility. The ULL Solution includes the new &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/participants/u4-machines"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;U4 machine family&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: sub;"&gt;,&lt;/span&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; also generally available today. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The ULL Solution is based on three infrastructure pillars:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Hardware-level networking:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Scaleable, hardware-based multicast data distribution for reliable market-data feeds&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Advanced networking observability:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Dynamic &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;network traffic capture with hardware-level timing accuracy, facilitating consistent&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;auditing, market replays, and absolute trade validation without impacting primary traffic performance&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Deterministic high-performance compute:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Processing of latency-critical execution tiers in a highly predictable, consistent amount of time, every single time&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Taken together, the ULL Solution’s compute, storage, networking, and observability provide financial exchanges and market participants with a number of technical capabilities: &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Bare metal performance:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Dedicated bare metal compute provides direct access to physical host resources, minimizing jitter and delivering predictable, low-latency execution for market-data feeds and order routing.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;High-performance storage options:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Local Titanium SSDs handle real-time transaction and tick logging on the host, complemented by scalable &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/disks/hyperdisks"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Hyperdisk&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for persistent market data archives and analytics.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Physical traffic isolation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The ULL trading fabric’s redundant A/B multicast market feeds are accessed through two independent and dedicated &lt;/span&gt;&lt;a href="https://cloud.google.com/titanium?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Titanium adapters&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, while an independent third Titanium adapter offloads telemetry, management, and provides access to Google Cloud services.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Hardware-accelerated multicast distribution: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The ULL network architecture supports hardware-level multicast feed ingestion. This allows participants to stream high-throughput market data directly to low-latency trading applications, bypassing traditional hypervisor-level virtual switches.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerated packet processing:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Support for &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/participants/use-onload"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OpenOnload&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/participants/use-dpdk"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;DPDK&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; enables Linux user-space networking to deliver predictable unicast and multicast packet handling while minimizing application code changes.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Precision timing and UTC synchronization: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Integration with Google Cloud’s&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;a href="https://cloud.google.com/blog/products/networking/understanding-the-firefly-clock-synchronization-protocol"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Firefly&lt;/span&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;clock synchronization system allows the solution to consistently&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;achieve&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/networking/understanding-the-firefly-clock-synchronization-protocol?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;sub-10 nanosecond network-level timestamping&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;and&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;better synchronization to UTC than the sub-100 microsecond regulatory requirement for financial exchanges.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Built-in telemetry&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;24/7 low-latency packet capture and seamless out-of-band packet brokering for regulatory compliance and real-time analytics helps ensure deep visibility without impacting primary traffic performance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;24-7 market ready&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The solution is &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;designed for continuous, round-the-clock trading readiness by isolating production workloads in a dedicated primary zone for live trading, while routine cloud maintenance and qualification testing occur in a secondary zone for updates and testing.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Compute in the ULL Solution is delivered by the new &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/participants/u4-machines"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;U4 machine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; family, which brings predictable performance and ultra-low latency compute in three specialized machine series: dual-socket bare metal instances with three physical NICs — &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;U4P for exchange operators and U4C for market participants&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; — alongside U4S high-performance VMs for operators, participants, and service providers.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Machine Series&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/cpu-platforms"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;CPU Platform&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;CPU Cores&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Memory&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;NICs&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Storage Options&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;U4P&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;(Bare Metal)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Intel 5th Gen Xeon (Emerald Rapids)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;120 Physical Cores&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;512 GB or 768 GB&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;3 Physical (1 Standard + 2 ULL)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Local Titanium SSD (12 TiB), Hyperdisk&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;U4C&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;(Bare Metal)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Intel 5th Gen Xeon (Emerald Rapids)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;120 Physical Cores&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;512 GB or&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;768 GB&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;3 Physical (1 Standard + 2 ULL)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Local Titanium SSD (12 TiB), Hyperdisk&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;U4S&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;(VM)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Intel 6th Gen Xeon (Granite Rapids)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;2 to 288 vCPUs&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Up to 2,232 GB&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Multi-vNIC (Up to 200 Gbps)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Local Titanium SSD (18 TiB), Hyperdisk&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We developed the U4 machine family to enable the world’s most technically demanding markets to run within Google Cloud and benefit from cloud services and scale. An example of this is our ongoing collaboration with CME Group, through which we are migrating listed derivatives markets to Google Cloud. Here, the U4C and U4P bare metal instances deliver direct physical co-location latency parity, providing predictable low-latency clock precision, and native hardware-multicast feed ingestion. This architecture demonstrates that core exchange systems and trading strategies can run in the cloud with the speed, consistency, and control that financial markets require.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/cme_group_g3mWTg1.max-1000x1000.jpg"
        
          alt="cme group"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="y5p56"&gt;&lt;i&gt;"Migrating the world’s leading derivatives marketplace to the cloud requires uncompromising performance, latency, and reliability. Google Cloud’s Ultra Low Latency Solution and the U4 instance families represent a breakthrough for financial market infrastructure. By combining deterministic, ultra-low latency networking and dedicated bare metal compute with the benefits of the cloud, market participants can execute and ingest market data and execute with the speed of traditional co-location and the agility of the cloud." -&lt;/i&gt; Pearce Peck-Walden, Managing Director Markets Engineering, CME Group&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Powered by Google Cloud-native infrastructure&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The ULL Solution is available in select &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/configuration-overview#locations"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;private Google Cloud regions&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, where trading teams can leverage core cloud infrastructure benefits such as rapid automated resource provisioning, on-demand capacity scaling, and dynamic fleet management. Trading teams can execute ultra-low latency trades on a dedicated, isolated ULL network, and offload data streams over an independent network interface to Google Cloud services such as BigQuery and Gemini Enterprise. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For exchange participants managing tick-to-trade workflows or processing real-time market data, microsecond variations are critical. As bare metal instances, U4P and U4C machine series don’t add any virtualization overhead for latency-sensitive execution tiers. Leveraging Google Cloud's Titanium offload system architecture, these bare metal instances provide direct access to the host resources. Compute Engine instances enable users to provision dedicated hardware through standard Google Cloud APIs, orchestrate deployments with infrastructure as code, apply Cloud Next Generation Firewall (NGFW) policies, and collect rich telemetry.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The U4S VM series complements these bare metal offerings for applications such as pre-trade risk validation and real-time analytics. Deployed alongside U4P and U4C instances, these VMs shorten transit hops across the trading architecture while offering elastic scaling from 2 to 288 vCPUs and networking up to 100 Gbps.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/prime_trading.max-1000x1000.jpg"
        
          alt="prime trading"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="y5p56"&gt;&lt;i&gt;“Our testing of the U4 instances validated that Google Cloud delivers the low latency and determinism required for our demanding exchange trading workloads. Having both bare metal and VM options allows us to evaluate the value of that performance relative to its cost, helping us align the right infrastructure to each of our trading strategies.”&lt;/i&gt; - Corbin Kidd, EVP/CTO, Prime Trading&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/28stone.max-1000x1000.jpg"
        
          alt="28stone"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="y5p56"&gt;&lt;i&gt;"In our tick-to-trade performance testing, the U4C bare metal instance demonstrated exceptional speed, showing that cloud compute and ultra-low latency networking can comfortably satisfy microsecond trading demands. But, what sets the U4 series apart isn't just raw processing speed, it’s structural determinism. When we stress-tested Google Cloud's U4 bare metal family, the performance curve remained virtually flat, which confirms that these cloud compute layouts can now offer the strict determinism institutional market participants require."&lt;/i&gt; - Christopher Wilson, Global Head of Experience Modernization, 28Stone&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Looking to the future, this low-latency architecture serves as the foundation for algorithmic trading, allowing firms to feed real-time markets directly into AI-driven trading models. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Storage and placement for trading workloads &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The U4 family supports Google Cloud storage options and placement policies, helping trading teams record trades at speed and minimize network latency between machines:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Local Titanium SSD for fast logging: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;For real-time trade logging and data caching without slowing down execution, U4P and U4C provide up to 12 TiB of direct-attached NVMe SSDs, and U4S VMs scale up to 18 TiB.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Hyperdisk for data archives: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;For historical market data, backtesting models, and regulatory archives, you can attach Hyperdisk Balanced and Hyperdisk Extreme volumes across all U4 series, scaling up to 32 volumes and 512 TiB per instance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Placement policies:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; To reduce physical network hops, compact placement policies keep related machines close together in the data center. Alternatively, spread placement policies place machines on separate racks to improve resiliency.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started with the ULL Solution&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With the ULL Solution, financial exchange operators, exchange participants, and trading service providers can execute trades with the speed and reliability they need today while positioning their businesses to leverage continuously evolving AI and infrastructure capabilities for the trading of tomorrow. The &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/participants/release-notes?hl=en"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ULL Solution&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is available in select Google Cloud regions. To learn more, check out the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/ull-solution/participants/configuration-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;public documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and reach out to our team to &lt;/span&gt;&lt;a href="mailto:exchanges@google.com"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;request access&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Wed, 07 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/financial-services/ultra-low-latency-solution-with-u4-enables-high-velocity-trading/</guid><category>Compute</category><category>Networking</category><category>Infrastructure</category><category>Storage &amp; Data Transfer</category><category>Infrastructure Modernization</category><category>Financial Services</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing Google Cloud’s U4 compute: Enabling ultra-low latency trading</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/financial-services/ultra-low-latency-solution-with-u4-enables-high-velocity-trading/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Disha Chopra</name><title>Product Manager, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Yarden Halperin</name><title>Product Manager, Google Cloud</title><department></department><company></company></author></item><item><title>Announcing MCP Toolbox Java SDK v1.0: Agentic data access for the enterprise</title><link>https://cloud.google.com/blog/topics/developers-practitioners/announcing-mcp-toolbox-java-sdk-v10-agentic-data-access-for-the-enterprise/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p style="text-align: center; font-size: 0.8rem;"&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Co-authors:&lt;/span&gt;&lt;/em&gt;&lt;br/&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Stenal Jolly&lt;/strong&gt;, Strategic Cloud Engineer, Google&lt;/span&gt;&lt;/em&gt;&lt;br/&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Anubhav Dhawan&lt;/strong&gt;, Software Engineer, Google&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Following the &lt;/span&gt;&lt;a href="https://medium.com/google-cloud/mcp-toolbox-v1-0-the-open-source-framework-for-secure-agentic-data-access-3c2199546ba8" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;landmark announcement&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; of MCP Toolbox v1.0, we're thrilled to announce that the &lt;a href="https://github.com/googleapis/mcp-toolbox-sdk-java" rel="noopener" target="_blank"&gt;MCP Toolbox Java SDK&lt;/a&gt; has officially reached version 1.0.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This release brings first-class, type-safe agent orchestration to one of the world's most widely adopted enterprise ecosystems. Java's mature architecture is purpose-built for rigorous demands, providing the high concurrency, strict transactional integrity, and robust state management required to safely scale mission-critical AI agents in production.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this post, we'll tell you about what's new in Java SDK v1.0, show you a real-world example, and help you get started with your own implementation.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;MCP: The universal interface&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, developers face a compounding integration bottleneck: if you have &lt;em&gt;N&lt;/em&gt; different AI models and &lt;em&gt;M&lt;/em&gt; enterprise data sources, you must build, secure, and maintain &lt;em&gt;N × M&lt;/em&gt; bespoke, custom connections. This lack of a unified integration layer forces engineering teams to rely on a fragmented web of ad-hoc pipelines. As a result, scaling an agentic architecture quickly becomes unsustainable, exposing sensitive enterprise databases to severe security vulnerabilities, fragmented access controls, and massive maintenance overhead.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Eliminating the fragmented web of custom integrations is the core problem solved by the Model Context Protocol (MCP). Acting as a universal interface—the "USB Type-C" for AI orchestration—MCP decouples models from data sources. Instead of writing custom or managed API integration code for every new model or database, developers write to a single, standardized protocol. This approach allows any MCP-compliant agent to securely and immediately interact with any MCP-enabled system. The MCP connection lets developers connect agents to real-world systems without building bespoke integrations for every new model.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;What's new in Java SDK v1.0: Built for production workloads&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When we announced the &lt;/span&gt;&lt;a href="https://medium.com/google-cloud/announcing-the-mcp-toolbox-java-sdk-2ed3171bbaa0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;public Beta&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for the MCP Toolbox Java SDK, our goal was to bring first-class, type-safe agent orchestration to enterprise Java environments. Since then, we've collaborated with developers and open-source contributors to harden our APIs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The v1.0 release marks a stable, backwards-compatible foundation suitable for enterprise workloads. Here's what's new and hardened since our v0.2 release:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Transport layer abstraction &amp;amp; &lt;/strong&gt;&lt;code&gt;&lt;strong style="vertical-align: baseline;"&gt;HttpMcpTransport&lt;/strong&gt;&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: We introduced a clean transport layer abstraction alongside &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;HttpMcpTransport&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. This feature decouples the core protocol logic from underlying HTTP clients, making it easy to swap network implementations or customize connection pooling.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Decoupled client authentication&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: To simplify enterprise security compliance, client authentication is now decoupled using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;CredentialsProvider&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;AuthMethods&lt;/code&gt; classes&lt;span style="vertical-align: baseline;"&gt;. Credentials are resolved asynchronously on every request, so teams can refresh tokens dynamically or plug in their own token source (Google OIDC via ADC ships in the box, anything else is a one-method interface).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Default parameter support&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Native support for default values in tool parameters, reducing prompt payload sizes and enhancing agent reliability.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pruning bound parameters&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Sensitive parameters that are bound server-side (like &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tenant_id&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) are now automatically stripped from exposed tool definitions so the LLM can't manipulate them.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Version selection &amp;amp; session tracking&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Standardized MCP version selection and robust session tracking ensure consistent protocol negotiation and conversation-state lifecycles.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;HTTP credential exposure warnings&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Added built-in detection that warns you at runtime when credentials are about to travel over a plaintext HTTP connection.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Generic client headers map&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Easily attach custom corporate proxy headers, transaction tracing IDs, or correlation metadata to all outgoing requests.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Get started with the Java SDK v1.0&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We designed the MCP Toolbox Java SDK to be frictionless for enterprise teams. Just add the following dependency to your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;pom.xml&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;&amp;lt;dependency&amp;gt;\r\n    &amp;lt;groupId&amp;gt;com.google.cloud.mcp&amp;lt;/groupId&amp;gt;\r\n    &amp;lt;artifactId&amp;gt;mcp-toolbox-sdk-java&amp;lt;/artifactId&amp;gt;\r\n    &amp;lt;version&amp;gt;1.0.0&amp;lt;/version&amp;gt;\r\n&amp;lt;/dependency&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd40d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Real-world example: The autonomous transit concierge&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To demonstrate the power of the Java SDK combined with AlloyDB, let's look at an enterprise use case.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Meet&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; Cymbal Transit&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, a fictitious intercity bus network. Customers don't want to click through nested dropdown menus to plan a trip. They want to ask:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"I need to get from New York to Boston tomorrow morning. Can I bring my Golden Retriever? If so, book me the fastest trip."&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To answer this question, an AI agent must cross-reference unstructured data (pet policies) with structured data (schedules and seat availability) and execute a transaction (booking)—all while maintaining the context of the conversation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The foundation: AlloyDB schema with native embeddings&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We used AlloyDB for this implementation because it handles relational data and high-dimensional vectors natively. Set up your database tables with these statements:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Enable necessary extensions for semantic search and embeddings\r\nCREATE EXTENSION IF NOT EXISTS vector;\r\nCREATE EXTENSION IF NOT EXISTS google_ml_integration;\r\n\r\n-- Table 1: Transit Policies (Unstructured Data for RAG)\r\nCREATE TABLE transit_policies (\r\n    policy_id SERIAL PRIMARY KEY,\r\n    category VARCHAR(50),\r\n    policy_text TEXT,\r\n    policy_embedding vector(768)\r\n);\r\n\r\n-- Table 2: Intercity Bus Schedules (Structured Data)\r\nCREATE TABLE bus_schedules (\r\n    trip_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),\r\n    origin_city VARCHAR(100),\r\n    destination_city VARCHAR(100),\r\n    departure_time TIMESTAMP,\r\n    arrival_time TIMESTAMP,\r\n    available_seats INT DEFAULT 50,\r\n    ticket_price DECIMAL(6,2)\r\n);\r\n\r\n-- Table 3: Booking Ledger (Transactional Action Data)\r\nCREATE TABLE bookings (\r\n    booking_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),\r\n    trip_id UUID REFERENCES bus_schedules(trip_id),\r\n    passenger_id VARCHAR(100),\r\n    status VARCHAR(20) DEFAULT &amp;#x27;CONFIRMED&amp;#x27;,\r\n    booking_time TIMESTAMP DEFAULT CURRENT_TIMESTAMP\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-sql&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd73d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Mapping intents to SQL: The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tools.yaml&lt;/code&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The MCP Toolbox lets you define custom tools securely. Rather than granting the LLM direct database access, a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tools.yaml&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; configuration maps natural language intents directly to parameterized, safe queries:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;kind: source\r\nname: alloydb\r\ntype: alloydb-postgres\r\nproject: my-project\r\nregion: us-central1\r\ncluster: my-cluster\r\ninstance: my-instance\r\ndatabase: postgres\r\n---\r\nkind: tool\r\nname: query-schedules\r\ntype: postgres-sql\r\nsource: alloydb\r\ndescription: Find available bus schedules between cities.\r\nparameters:\r\n  - name: origin\r\n    type: string\r\n    description: The departure city name.\r\n  - name: destination\r\n    type: string\r\n    description: The arrival city name.\r\n  - name: limit\r\n    type: integer\r\n    description: Maximum number of schedules to return.\r\n    default: 5\r\nstatement: |\r\n  SELECT CAST(trip_id AS TEXT) AS trip_id, departure_time, ticket_price\r\n  FROM bus_schedules\r\n  WHERE lower(origin_city) = lower($1) AND lower(destination_city) = lower($2)\r\n  ORDER BY departure_time ASC\r\n  LIMIT $3&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd77d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For the complete YAML file, see the &lt;/span&gt;&lt;a href="https://github.com/googleapis/mcp-toolbox-sdk-java/blob/main/demo-applications/cymbal-transit/tools.yaml" rel="noopener" target="_blank"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;tools.yaml&lt;/code&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; file in the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;mcp-toolbox-sdk-java&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; repository.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Stateful agent architecture in Spring Boot&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The hardest part of building conversational AI in enterprise applications is managing state: when a user asks, "What times are available?" and follows up with, "Book the 8 AM one," the agent must remember prior context across turns.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Using the Java MCP Toolbox SDK with Spring Boot and LangChain4j, we can cleanly maintain conversational memory in the HTTP Session and we can cleanly separate the agent into two declarative components:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A declarative agent interface that manages the prompt, tools, and conversational memory via an HTTP session.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A tool execution service that routes agent requests directly to the MCP Toolbox server.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;interface TransitAgent {\r\n    @SystemMessage({\r\n        &amp;quot;You are the Cymbal Transit Concierge.&amp;quot;,\r\n        &amp;quot;Use the \&amp;#x27;querySchedules\&amp;#x27; tool for finding schedules.&amp;quot;,\r\n        &amp;quot;Use \&amp;#x27;bookTicket\&amp;#x27; to execute transactions.&amp;quot;,\r\n        &amp;quot;Use \&amp;#x27;searchPolicies\&amp;#x27; to look up luggage and pet rules.&amp;quot;\r\n    })\r\n    String chat(@MemoryId String sessionId, @UserMessage String userMessage);\r\n}\r\n\r\n@Service\r\nclass TransitAgentTools {\r\n    // These methods automatically invoke our MCP Toolbox server!\r\n    @Tool(&amp;quot;Query specific schedules between an origin and destination city.&amp;quot;)\r\n    public String querySchedules(String origin, String destination) { ... }\r\n\r\n    @Tool(&amp;quot;Book a ticket for a passenger.&amp;quot;)\r\n    public String bookTicket(String tripId, String passengerName) { ... }\r\n\r\n    @Tool(&amp;quot;Search transit policies for luggage and pet rules.&amp;quot;)\r\n    public String searchPolicies(String query) { ... }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd5d10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Notice how the &lt;code&gt;@MemoryId&lt;/code&gt; annotation abstracts session tracking: Spring Boot automatically correlates conversational context to the user's HTTP session. Meanwhile, LangChain4j and the MCP Toolbox handle schema translation and tool routing behind the scenes—no handwritten if/else intent parsing required.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By pairing the MCP Toolbox Java SDK with LangChain4j, we achieve clean separation of concerns and effortless state management:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero-boilerplate session management&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;@MemoryId String sessionId&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; parameter binds conversation history directly to the user's HTTP session.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Declarative agent contract&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;TransitAgent&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; interface defines the model's persona and system instructions without complex prompt templating.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Type-safe tool execution&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;TransitAgentTools&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Spring service wraps remote MCP database tools as native Java methods.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This architecture ensures your agent remains modular: you can refine prompt guidance in the interface, manage user sessions automatically, and execute secure database queries through MCP Toolbox without tight coupling.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting the dots: Listing, invoking, and executing tools in Java v1.0&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Now let's look under the hood of the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;TransitAgentTools&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; interface. Inside those LangChain4j &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;@Tool&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; methods on our Spring &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;@Service&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, the MCP Toolbox Java SDK handles the heavy lifting—bridging your Java service methods to the MCP tools defined in the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tools.yaml&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; file. In just a few lines of type-safe code, we can initialize our client using the new v1.0 decoupled authentication and headers abstractions:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;// 1. Initialize the Client with Decoupled Auth and Custom Headers (v1.0)\r\nString serviceUrl = &amp;quot;https://toolbox-my-project-uc.a.run.app/mcp&amp;quot;;\r\nMcpToolboxClient mcpClient = McpToolboxClient.builder()\r\n    .baseUrl(serviceUrl)\r\n    .credentialsProvider(new GoogleCredentialsProvider(serviceUrl)) // Decoupled OIDC credentials\r\n    .headers(Map.of( // Generic client headers\r\n        &amp;quot;X-Correlation-ID&amp;quot;, &amp;quot;enterprise-session-abc123&amp;quot;,\r\n        &amp;quot;X-Client-Platform&amp;quot;, &amp;quot;Spring-Boot&amp;quot;\r\n    ))\r\n    .build();\r\n\r\n// 2. Listing Discoverable Tools\r\nmcpClient.listTools().thenAccept(tools -&amp;gt; {\r\n    System.out.println(&amp;quot;Successfully discovered &amp;quot; + tools.size() + &amp;quot; tools.&amp;quot;);\r\n});\r\n\r\n// 3. Invoking a Tool (Read-Only Data with Default Parameter Support)\r\n// &amp;quot;limit&amp;quot; is omitted: the SDK fills it from the default in the tool definition\r\nString schedules = mcpClient.loadTool(&amp;quot;query-schedules&amp;quot;)\r\n    .thenCompose(tool -&amp;gt; tool.execute(Map.of(\r\n        &amp;quot;origin&amp;quot;, &amp;quot;New York&amp;quot;,\r\n        &amp;quot;destination&amp;quot;, &amp;quot;Boston&amp;quot;)))\r\n    .join().text();\r\n\r\n// 4. Executing a Transactional Tool (Using Bound Parameters)\r\nAuthTokenGetter toolAuthGetter = () -&amp;gt; CompletableFuture.completedFuture(myIdToken);\r\n\r\nString bookingConfirmation = mcpClient.loadTool(&amp;quot;book-ticket&amp;quot;, Map.of(&amp;quot;google_auth&amp;quot;, toolAuthGetter))\r\n    // Bind the authenticated user context securely. bindParam returns a new immutable\r\n    // Tool, and the bound parameter is pruned from the definition exposed to the LLM!\r\n    .thenCompose(tool -&amp;gt; tool.bindParam(&amp;quot;passenger_name&amp;quot;, &amp;quot;Jane Doe&amp;quot;)\r\n        // Execute the mutable transaction\r\n        .execute(Map.of(&amp;quot;trip_id&amp;quot;, &amp;quot;123e4567-e89b-12d3-a456-426614174000&amp;quot;)))\r\n    .join().text();&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd52d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Secure by default: authentication and deployment&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Moving an AI agent to production requires rock-solid credential handling and an infrastructure that scales with demand. Let's look at how you can enforce credential safety across environments and deploy independently on Cloud Run.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Application Default Credentials (ADC) &amp;amp; safety&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By using the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GoogleCredentialsProvider&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; service&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, your Java app inherits its secure identity from its execution environment (whether local or in Google Cloud) through Application Default Credentials (ADC)—no hard-coded keys, with OIDC tokens minted and cached per audience under the hood. Furthermore, v1.0 offers HTTP Credential Exposure Warnings that automatically detect when credentials are about to travel over a plaintext HTTP connection and emit a runtime warning telling you to switch to HTTPS.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Deploying the fleet to Cloud Run&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because MCP Toolbox and the Spring Boot Agent are fully decoupled, they scale independently on Google Cloud Run to meet high concurrency and stateful conversation requirements.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To set up and configure Toolbox on Cloud Run, download the open-source MCP Toolbox for Databases and then follow the &lt;/span&gt;&lt;a href="https://mcp-toolbox.dev/documentation/deploy-to/cloud-run" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;deployment guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Get started today&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With MCP Toolbox Java SDK v1.0, enterprise Java teams can wire Spring Boot and LangChain4j agents to the Toolbox server, and through it, to AlloyDB and every other supported data source. When you use the toolbox, arguments are validated against the tool definition before they leave the JVM, authentication is decoupled, and custom headers are attached to every outgoing request. The following &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;implementation steps will help you get started&lt;/strong&gt;.&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Step 1: Add the Dependency&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To start building with the SDK, add the following dependency to your Maven project's &lt;code&gt;pom.xml&lt;/code&gt; file:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;&amp;lt;dependency&amp;gt;\r\n        &amp;lt;groupId&amp;gt;com.google.cloud.mcp&amp;lt;/groupId&amp;gt;\r\n        &amp;lt;artifactId&amp;gt;mcp-toolbox-sdk-java&amp;lt;/artifactId&amp;gt;\r\n        &amp;lt;version&amp;gt;1.0.0&amp;lt;/version&amp;gt;\r\n &amp;lt;!-- {x-version-update:mcp-toolbox-sdk-java:current} --&amp;gt;\r\n    &amp;lt;/dependency&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd6c10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Step 2: Explore Resources &amp;amp; Demos&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GitHub Repository&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;a href="https://github.com/googleapis/mcp-toolbox-sdk-java" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Java SDK for interacting with the MCP Toolbox for Databases&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Official Documentation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;a href="https://mcp-toolbox.dev/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP Toolbox for Databases&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/ul&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Demo Application&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: To experience the Java SDK V1.0 for MCP Toolbox latest, try the sample application &lt;a href="https://github.com/googleapis/mcp-toolbox-sdk-java/tree/main/demo-applications/cymbal-transit" rel="noopener" target="_blank"&gt;Cymbal transit project&lt;/a&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;Gradle implementation&lt;/h4&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;If your team uses Gradle instead of Maven, remember they will need to translate this dependency:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;implementation &amp;#x27;com.google.cloud.mcp:mcp-toolbox-sdk-java:1.0.0&amp;#x27;&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cbd4050&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Automatic version tracking&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;I&lt;span style="vertical-align: baseline;"&gt;f you copy this setup into automated internal repositories, k&lt;/span&gt;eep the XML comment &lt;code&gt;&amp;lt;!-- {x-version-update...} --&amp;gt;&lt;/code&gt; intact. It's required by the release manager's deployment scripts to automatically bump versions.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Now that you can integrate your modern agentic tools and servers to your enterprise Java applications with a stale MCP Toolbox Java SDK, get started today!&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 06 Oct 2026 19:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/announcing-mcp-toolbox-java-sdk-v10-agentic-data-access-for-the-enterprise/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/announcing-mcp-toolbox-java-sdk-v10-agentic-.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Announcing MCP Toolbox Java SDK v1.0: Agentic data access for the enterprise</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/announcing-mcp-toolbox-java-sdk-v10-agentic-.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/announcing-mcp-toolbox-java-sdk-v10-agentic-data-access-for-the-enterprise/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Abirami Sukumaran</name><title>Staff Developer Advocate, Google</title><department></department><company></company></author></item><item><title>Managed Apache Iceberg at scale: How Spanner powers Lakehouse runtime catalog</title><link>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-powered-by-spanner/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As customers modernize to lakehouse architectures, they are standardizing on open formats such as Apache Iceberg to create a shared data estate across compatible engines. This enables you to build AI-native lakehouses that turn data to semantic knowledge, enable proactive action, and operate at an agentic scale.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;One of the key components of a lakehouse is the catalog, and in the Apache Iceberg environment, that usually means the Iceberg REST Catalog. An Apache Iceberg catalog is responsible for maintaining table pointers, handling atomic commits, and serving as the single source of truth for table locations. But as large enterprise organizations modernize to lakehouses, they have begun to realize that they need a highly scalable and available managed catalog as a part of their lakehouse architecture. This becomes even more important as querying scales with agents. To support a large amount of repeated small queries from agents, you will need to build on top of a managed catalog that provides atomicity, consistency, availability and concurrency at massive scale. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this blog, we explore the challenges a managed catalog faces in modern cloud environments at agent scale, and show you how Google Cloud’s serverless &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/biglake-metastore-now-supports-iceberg-rest-catalog?e=a"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse runtime catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; can help address them. Powered by &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google Cloud’s always-on database with virtually unlimited scale, and built to meet the open Apache Iceberg REST catalog specification, the Lakehouse runtime catalog is the highly scalable and available foundation you need for the agentic era.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Challenges a managed catalog needs to solve in the Lakehouse&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When speaking with data engineers and infrastructure leads running production analytics at scale in a Lakehouse, the following core pain points consistently emerge and need to be solved by a managed catalog: &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Atomic commits and concurrency control:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Iceberg guarantees ACID transactions via optimistic concurrency control (OCC). A managed catalog must implement a bulletproof atomic compare-and-swap (CAS) operation to swap the current metadata pointer.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;High availability and operational maintenance:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Because queries fail immediately if the catalog is down, a managed catalog becomes a critical Tier-1 service.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Scaling the database backing the catalog:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Catalog architects typically face a difficult trade-off when choosing a backing database for table metadata and state. Traditional scale-up relational databases provide SQL and ACID transactions, but hit vertical CPU, memory, storage and connection limits under heavy concurrent read/write loads unless manually sharded which incurs a huge operational overhead; while scale-out database systems are either eventually consistent, hard to manage, not enterprise-ready, or all of the above. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Table maintenance coordination:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A catalog alone does not optimize data; you must build and operate ancillary pipelines for compaction, snapshot expiration, manifest rewriting, and orphan file cleanup.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Governance and security:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A catalog acts as the security gatekeeper. The catalog must implement and maintain:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Authentication protocols (e.g., OAuth2 token exchange, IAM federation)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Access control down to namespace and table levels&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Vended storage credentials (e.g., generating short-lived tokens so query engines don't need broad, direct storage credentials)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Lakehouse runtime catalog&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To solve these challenges, we built the Lakehouse runtime catalog (GA) with support for Iceberg Rest Catalog. The Lakehouse runtime catalog is a fully serverless, highly available, and unified metadata registry designed from the ground up to support modern open table formats like&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Apache Iceberg. By natively implementing the Apache Iceberg REST Catalog specification&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;,&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; the Lakehouse runtime catalog decouples metadata discovery from compute engines, helping ensure multiple Iceberg-compatible engines can access a shared data estate and enabling you to take your workloads to production sooner. We’ve helped many customers streamline the migration of their managed catalogs. For example, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/etsy-lakehouse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Etsy&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; migrated its catalog to Lakehouse runtime catalog, joining data in place to accelerate pipeline queries by 60%.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_bmVFk8w.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This approach offers a number of architectural benefits:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Open APIs&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Support for Iceberg Rest Catalog enables different teams to use their preferred analytics tools on a single, unified dataset.&lt;/span&gt;&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-engine interoperability&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Once registered, tables are immediately discoverable and queryable across Google Cloud Managed Service for Apache Spark, BigQuery, and open-source engines via standard REST interfaces. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/the-future-of-data-lakehouse-for-the-agentic-era?e=0"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Read/write interoperability&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; for Iceberg tables: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Leverage Iceberg-compatible engines such as BigQuery, Managed Spark to write to Iceberg tables registered in the Lakehouse runtime catalog. Customers can also use Managed Spark to write to Iceberg tables in external catalogs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Fully managed Iceberg storage with enterprise-grade features: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Use Google's &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/unveiling-new-bigquery-capabilities-for-the-agentic-era?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;differentiated infrastructure&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to run analytics with performance on Iceberg tables. This gives you the benefits of open-source flexibility plus performance, scale, governance, and multimodal processing. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero data copy&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Table definitions point directly to your existing data in the underlying object store. You do not move, rewrite, or duplicate your underlying data.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Bi-directional catalog federation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;across clouds:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Access data from &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-cross-cloud-connection-databricks"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Databricks Unity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-cross-cloud-connection-snowflake"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Snowflake Horizon&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-cross-cloud-connection-aws-glue"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AWS Glue&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with support for vended credentials and OIDC token exchange. This lets you bring Google AI &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/introducing-the-borderless-lakehouse?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;directly to your AWS and Azure data&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Secure access using credential vending&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The catalog supports multiple authorization mechanisms, letting you choose between &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/credential-vending"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;credential vending&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and end-user credentials. This means that you can access tables with modern mechanisms such as credential vending without needing direct access to the files in the underlying object store (Cloud Storage, AWS S3, Azure Blob Storage). &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;AI-powered context and governance: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The Lakehouse runtime catalog integrates directly with &lt;/span&gt;&lt;a href="https://cloud.google.com/products/knowledge-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/products/iam?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud IAM&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, allowing you to define trusted context for your agents, and apply table-level security consistently across all compute engines.&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Get out-of-the-box search, lineage, and insights for Iceberg tables in the catalog. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Atomic commits and concurrency control, high availability and scalability: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Backed by Google’s planet-scale infrastructure and Spanner, you get the high availability, concurrency and scale you need for your metadata to scale with your data. Support for Cloud Storage dual-region and multi-region buckets enables failover use cases. It also provides reduced TCO due to serverless and no-ops environments, and scalability for any workload size.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;What powers the Lakehouse runtime catalog?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Lakehouse runtime catalog is a highly available, concurrent and scalable catalog with strong consistency guarantees because it is built on top of &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Unlike traditional scale-up relational databases that hit vertical single-node ceilings, Spanner combines full relational SQL semantics and multi-table ACID transactions with the horizontal scale-out elasticity for both reads and writes of a NoSQL system. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Spanner makes the Lakehouse runtime catalog highly available through Spanner’s &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/instance-configurations#regional-configurations"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;regional configurations&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with up to 99.99% availability. Spanner also delivers out-of-the-box scalability for  Lakehouse runtime catalog: As a horizontally scalable database, Spanner does not require manual sharding and scales compute and storage independently and transparently. Spanner dynamically monitors data volume and query load, splitting and redistributing data ranges across nodes. Compute nodes scale dynamically based on CPU utilization and storage thresholds. Spanner automatically detects split-level overload and moves heavy splits away from overloaded nodes. Spanner also lets Lakehouse runtime catalog users eliminate the &lt;/span&gt;&lt;a href="https://dl.acm.org/doi/abs/10.1145/2491245?__cf_chl_tk=oi5rkZKtX5q2CR93K9mrIHsAp_oqRxwaGqqpqq99v9Y-1790923307-1.0.1.1-pNUQvRr4OBosWiOcfxGuTWTD8I3JW_8V0i5xLuCsErU" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;traditional trade-offs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; between relational consistency and distributed scalability, delivering Lakehouse runtime catalog’s industry-leading &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/true-time-external-consistency"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;consistency guarantees&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Lakehouse transactions are serializable — the order of transactions within the database is the same as the order in which clients observe the transactions to have been committed. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;This foundation allows Lakehouse users to operate at agent-scale. &lt;/span&gt; &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then, to further power agentic use cases, the Lakehouse runtime catalog integrates directly with Knowledge Catalog to easily discover lakehouse Iceberg tables and provide trusted context to agents. Knowledge Catalog leverages an efficient combination of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/full-text-search"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;full-text search&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and native&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/vector-search-overview"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;vector search&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provided by Spanner; this approach enables better recall for search retrieval, pairing lexical keyword searches with semantic embeddings in a single query. Because both index types are built on the identical base dataset, they update with strict, transactional ACID consistency alongside base table DML operations. This removes operational overhead such as managing sync pipelines, and external-vector and full-text search systems. In short, the Lakehouse runtime catalog provides faster time-to-market for your agentic use cases.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Combine analytical and operational workloads&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With Google Cloud’s borderless Lakehouse based on Apache Iceberg, you can combine your analytical data with your operational workloads. Use cases span combining data assets from your Lakehouse with OLTP data (from Spanner) for analytics, to low-latency serving applications where your Lakehouse assets are accessible in an operational database such as Spanner, to conversational analytics in first-party and third-party agents. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Below, in an example, you can see the Lakehouse runtime catalog with Spanner in action. Here, we combine analytical data for taxi trips in Manhattan (backed by Apache Iceberg) with operational data for taxi zones in Spanner to find the most congested traffic routes. The example also shows that you can also use a Conversational Analytics Agent to access the same data and get second-order insights.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_TvlKrOP.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="cjj42"&gt;Find out most congested traffic routes in Manhattan&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Modernize to the borderless Lakehouse&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Modernizing to Google Cloud’s&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/solutions/data-lakehouse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;minimizes data silos across analytics engines and agents, unifies multi-engine governance, provides trusted context to your agents and slashes operational TCO. Leveraging the Spanner-based Lakehouse &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-lakehouse-iceberg-rest-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;runtime catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; helps prepare your modern cloud environments to operate at agent-scale. To learn more and get started with a free trial, visit the &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/products/lakehouse?e=a"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; web&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;page and learn more about Spanner &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 06 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-powered-by-spanner/</guid><category>Databases</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Managed Apache Iceberg at scale: How Spanner powers Lakehouse runtime catalog</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-powered-by-spanner/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vinod Ramachandran</name><title>Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Mina Mikhail</name><title>Software Engineering Manager</title><department></department><company></company></author></item><item><title>Networking for AI inference model serving - GKE only and for all other backends</title><link>https://cloud.google.com/blog/topics/developers-practitioners/networking-for-ai-inference-model-serving-gke-only-and-for-all-other-backends/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprises and individual developers frequently run multiple AI inference models. The right architecture can simplify how the models are called while also providing centralized governance. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;In this post, we'll look at two reference architectures focused on networking AI inference model serving: one for &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/kubernetes-engine-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine (GKE)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and one all other backend types. First, we'll explore the commonalities between the reference architectures that you'll see later. Then we'll explore unique components of the architecture for GKE backends and finally, we'll go over the elements of the architecture for all backend types.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;The entry point&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can expose your model deployment behind a stable, secure, and reliable entry point that acts as the front end for inference calls. This entry point also acts as a control zone where policy, security, and logic can be enforced. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Both the Cloud Load balancer and the Inference Gateway provide entry point capability. These types of endpoints can terminate secure connections with TLS, integrate with API management components, extend functionally with service extensions, and capitalize on capabilities of Model Armor for added security.  &lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Common services in the designs &lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Both reference designs use these services:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/vpc/docs/private-service-connect#endpoints"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Private Service Connect inference endpoint&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Anchors the entry point inside your consumer Virtual Private Cloud (VPC) network. Traffic hits a private internal &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/vpc/docs/subnets#valid-ranges"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;IP address&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, keeping inference calls in your private network.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/apigee/docs/api-platform/get-started/what-apigee"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Apigee API Management&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; (Optional)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Integrates via an Apigee Extension Processor callout to handle client identity verification, rate limits, and quota enforcement before requests ever reach compute resources.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/model-armor/overview"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Model Armor&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Serves as an inline AI safety checkpoint, screening prompts and output completions against prompt injection and sensitive data leakage.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Design pattern serving on GKE only&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This section focuses on a GKE-only backend design.&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;To understand the full end-to-end concept, please read the entire architecture document &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/networking-for-ai-inference-gke"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Networking for AI inference model serving on GKE&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;The design pattern is based on this diagram:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/networking-ai-inference-model-serving-gke-.max-1000x1000.png"
        
          alt="networking-ai-inference-model-serving-gke-architecture"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition to the common services identified in the previous section, the design for GKE uses these components:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/about-gke-inference-gateway"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Inference Gateway&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Deployed as an internal Application Load Balancer (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gke-l7-rilb&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). It acts as a specialized ingress engine that parses incoming request payloads, evaluates HTTPRoute rules, and steers queries to appropriate model-serving targets.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Inference pools&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: A logical group containing replicas of the same model. When the Gateway receives a prompt, it evaluates HTTPRoute rules to select the appropriate inference pool based on the model identifier. Pools have an initial size and can be configured to autoscale dynamically.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Model replica sets&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Individual model replicas (inference server instances) deployed across single-node or multi-node GPU or TPU node pools. A replica set represents a uniform group of these model replicas.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The traffic flow GKE example&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A client application that uses this GKE-based architecture to call a backend model would go through a flow like this:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Ingress&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: A client application in the consumer VPC issues an OpenAI-compatible API call to the local Private Service Connect endpoint, routing directly to the GKE Inference gateway using a regional internal Application Load Balancer.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Payload inspection&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The Gateway reads the target model parameter specified in the request body and adds it to the HTTP headers.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Control plane validation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: If Apigee is used, it checks client credentials and quotas. Model Armor screens the prompt for policy violations or data leakage.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Backend selection&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The Gateway evaluates HTTPRoute mappings to identify the target pool, matches shared prefix cache context, and routes to the lowest-load GPU or TPU replica based on real-time Prometheus data.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Egress&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The replica runs the inference workload. Output tokens pass through Model Armor for final response verification before streaming back over the private connection.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Design pattern serving on all backends&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This section focuses on multiple backend types which can be used for inference, and it provides an overview of the architecture in the following diagram.&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;To understand the full end-to-end concept, please read the entire architecture document &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/networking-for-ai-inference"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Networking for AI inference model serving on all backends&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/networking-ai-inference-model-serving-all-.max-1000x1000.png"
        
          alt="networking-ai-inference-model-serving-all-backends"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For architectures spanning mixed environments such as GKE, Cloud Run, Agent Platform, on-premises data centers, or external clouds, these additional components are used:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/load-balancing/docs/l7-internal#load-balancer-mode"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Regional internal Application Load Balancer&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Serves as the central Layer 7 routing proxy that manages routing logic, SSL termination, and Service Extensions callouts.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://github.com/llm-d/llm-d-inference-payload-processor" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Inference Payload Processor&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; (Service Extensions)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: This is similar to body-based routing as used in the GKE Inference Gateway, but to enable the functionality on an Application Load Balancer a service extension is needed.&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;A lightweight Cloud Run callout inspects the JSON body of incoming OpenAI API requests, extracts the target model identifier, and writes an X-Gateway-Model-Name header to drive URL map routing.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/load-balancing/docs/negs"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Network Endpoint Group (NEG)&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Delivers flexible routing to heterogeneous backends based on the injected model header.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;All-backends traffic flow example&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A client application that uses this architecture to call a backend model would go through a flow like this:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Private ingress&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The client application targets the Private Service Connect endpoint over private IP address space. The regional internal Application Load Balancer receives the request.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Model name extraction&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The load balancer sends the payload to the Cloud Run body-based router callout, which inspects the JSON payload and injects the X-Gateway-Model-Name header.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Policy and safety enforcement&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The request passes to Apigee for identity and quota validation, then to Model Armor to scrub sensitive data and block malicious prompts.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;NEG routing&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The load balancer URL map inspects the model header and forwards the request to the matching backend NEG (Agent Platform, GKE, Cloud Run, Hybrid, or Internet).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Private delivery&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The target backend executes the model prompt, Model Armor screens the completion, and the result returns privately along the ingress path.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/networking-ai-inference-model-serving-flow.max-1000x1000.png"
        
          alt="networking-ai-inference-model-serving-flow"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;What's next&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Take a deeper dive into building AI workloads on Google Cloud:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Document set: &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/agentic-ai-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agentic AI architecture guides&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Architecture Center: &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/multi-agent-private-networking-patterns"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Multi-agent private networking patterns in Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Document: &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/machine-learning/general/netsec-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform networking access overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Want to ask a question, find out more, or share a thought? Please connect with me on &lt;/span&gt;&lt;a href="https://www.linkedin.com/in/ammett/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Linkedin&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 06 Oct 2026 12:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/networking-for-ai-inference-model-serving-gke-only-and-for-all-other-backends/</guid><category>Networking</category><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/networking-ai-inference-model-serving-hero.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Networking for AI inference model serving - GKE only and for all other backends</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/networking-ai-inference-model-serving-hero.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/networking-for-ai-inference-model-serving-gke-only-and-for-all-other-backends/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Ammett Williams</name><title>Developer Relations Engineer</title><department></department><company></company></author></item><item><title>Where mission meets moonshot: Join us at the Google Public Sector Summit 2026</title><link>https://cloud.google.com/blog/topics/public-sector/where-mission-meets-moonshot-join-us-at-the-google-public-sector-summit-2026/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="2qeha"&gt;Across the public sector, the question has shifted from “what’s possible?” to “how can we drive mission impact?” We’ll be answering this question at the &lt;a href="https://events.govexec.com/google-public-sector-summit/" target="_blank"&gt;Google Public Sector Summit&lt;/a&gt; on October 20th at the Ronald Reagan Building in Washington, D.C.&lt;/p&gt;&lt;p data-block-key="3bjuv"&gt;This year’s summit brings together public sector leaders, creators, and builders to architect what’s next and bridge the gap between ambitious innovation and secure, scaled execution. Together, we’ll explore how public sector organizations are harnessing Google’s secure, integrated AI stack—encompassing everything from planet-scale infrastructure to leading models and agentic platforms—to turn ambitious moonshot aspirations into real-world outcomes that drive mission impact.&lt;/p&gt;&lt;p data-block-key="2536b"&gt;The Google Public Sector Summit is a celebration of what’s possible when human ambition and powerful technology come together in support of mission. Keep reading for a preview of the event, and how to make the most of your onsite experience.&lt;/p&gt;&lt;h3 data-block-key="14s2m"&gt;&lt;b&gt;Opening Keynote: Where Mission Meets Moonshot&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="spra"&gt;Get inspired during our opening keynote which will feature &lt;b&gt;Karen Dahut&lt;/b&gt;, CEO of Google Public Sector,&lt;b&gt; Thomas Kurian&lt;/b&gt;, CEO of Google Cloud, and &lt;b&gt;Chris Hein&lt;/b&gt;, Field CTO of Google Public Sector. They will share real-world customer examples and live product demonstrations that showcase Google’s technology in action across the public sector. Attendees will also get the opportunity to hear from public sector leaders who are collaborating with Google and applying technology to transform missions and power moonshots.&lt;/p&gt;&lt;h3 data-block-key="7r63f"&gt;&lt;b&gt;General Session: Cloud Talks&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="3a53u"&gt;Join &lt;b&gt;Elizabeth Moon&lt;/b&gt;, Managing Director of Customer Engineering for Google Public Sector, as she hosts a series of talks that dive deeper into the critical pillars of public sector transformation including infrastructure, AI, and security. Hear from &lt;b&gt;Rich Sanzi,&lt;/b&gt; VP of Engineering for Google Cloud, &lt;b&gt;Michael Gerstenhaber&lt;/b&gt;, VP of Product Management for Gemini, and &lt;b&gt;Sandra Joyce&lt;/b&gt;, VP of Google Threat Intelligence, as they share strategies for establishing an infrastructure foundation for the agentic era, building safely with Google’s platforms, and implementing proactive defense at planet scale.&lt;/p&gt;&lt;h3 data-block-key="4mnuc"&gt;&lt;b&gt;Breakouts: Panel Discussions and Birds of a Feather&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="28co6"&gt;Hear from Google Public Sector executives, industry partners, and leaders across defense, federal civilian, higher education, and state and local government through a series of panels and birds of a feather discussions.&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="fl18p"&gt;&lt;b&gt;From fragmented to fluid: Connecting data to build the agentic future.&lt;/b&gt; Moderated by &lt;b&gt;Tony Orlando,&lt;/b&gt; Managing Director of Partner and Specialty Sales, this session will explore how agencies can break down legacy data silos and build a unified, AI-ready data foundation. Speakers include &lt;b&gt;Andrew Mapes&lt;/b&gt;, CDAO Principal Deputy Chief Digital and AI Officer, &lt;b&gt;Dr. Neil Jacobs&lt;/b&gt;, NOAA Administrator and Undersecretary of Commerce for Oceans and Atmosphere, and &lt;b&gt;Scott Alfieri&lt;/b&gt;, Accenture Global Lead for the Google Business Group. An additional speaker, &lt;b&gt;Adarryl Roberts&lt;/b&gt;, DLA CIO, has been invited.&lt;/li&gt;&lt;li data-block-key="e9l72"&gt;&lt;b&gt;Automating defense: Securing the agentic era against complex threats.&lt;/b&gt; Moderated by &lt;b&gt;Ron Bushar&lt;/b&gt;, Managing Director and Chief Security Officer, panelists will discuss how they’re applying autonomous AI and threat intelligence to protect dynamic environments at machine speed. Speakers include &lt;b&gt;Justin Fanelli&lt;/b&gt;, U.S. Navy CTO, &lt;b&gt;Jamie Wolff&lt;/b&gt;, DOE-NNSA CIO, Vice Admiral (Ret.) &lt;b&gt;TJ White&lt;/b&gt;, Texas Cyber Command Chief, and &lt;b&gt;Thomas Browning&lt;/b&gt;, Rhombus Power Senior Vice President of Strategic Capabilities.&lt;/li&gt;&lt;li data-block-key="7qu1g"&gt;&lt;b&gt;The agentic era: Advancing missions and transforming the way we work.&lt;/b&gt; Moderated by &lt;b&gt;Elizabeth Moon&lt;/b&gt;, this session will highlight how secure, cost-efficient agentic capabilities can scale impact and advance missions. Featured speakers include &lt;b&gt;Barry J. Schindler&lt;/b&gt;, USPTO Acting Commissioner for Patents, &lt;b&gt;Miro Humer&lt;/b&gt;, Case Western Reserve University CIO, and &lt;b&gt;Todd Johnston&lt;/b&gt;, Deloitte Managing Director, AI &amp;amp; Engineering and Defense, Security &amp;amp; Justice AI/Analytics Leader.&lt;/li&gt;&lt;li data-block-key="6o8e3"&gt;&lt;b&gt;Interactive Birds of a Feather discussions.&lt;/b&gt; Connect directly with a small group of peers and Google experts to address critical mission blockers, share best practices, and accelerate your agency's journey beyond the pilot stage. Join discussions across a range of topics including tokenomics, migration, and compliance.&lt;/li&gt;&lt;/ul&gt;&lt;h3 data-block-key="dtduu"&gt;&lt;b&gt;Demos, Lightning Talks and Labs&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="4edo0"&gt;Experience Google’s innovations through hands-on labs, Lightning Talks, and demos taking place in the Mission Exchange on the showfloor. Topics will include agentic threat defense, multimodal AI citizen services, and legacy database modernization with AI.&lt;/p&gt;&lt;p data-block-key="d0cqb"&gt;Attendees can also join two expert-led labs:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="fd9v5"&gt;&lt;b&gt;Getting Started with Gemini:&lt;/b&gt; This introductory lab will guide participants through the fundamentals of Gemini and Gemini Notebook to summarize documents, search data and streamline daily tasks. No technical background required.&lt;/li&gt;&lt;li data-block-key="1crt4"&gt;&lt;b&gt;Build Gemini Agents: From Prompt to Prototype:&lt;/b&gt; This lab will help you bring your real-world agency use cases to life by building an AI agent with Gemini, with support from Google Cloud specialists.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="aa7fn"&gt;Throughout the day, we will host Lightning Talks, which are dynamic 20-minute presentations featuring live demos and interactive Q&amp;amp;A that take place on the showfloor. These talks feature partner and customer speakers and showcase the "art of the practical" through live, real-world solutions.&lt;/p&gt;&lt;h3 data-block-key="4svn3"&gt;&lt;b&gt;Closing Session: What’s Your Moonshot?&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="b144u"&gt;Join us for an inspirational closing session that I’ll moderate with a special guest: NASA Administrator &lt;b&gt;Jared Isaacman&lt;/b&gt;. Together, we’ll define what a moonshot means for the public sector in the agentic era and share strategies for fostering a culture of courageous problem-solving and applying technology to supercharge human ambition and ingenuity.&lt;/p&gt;&lt;h3 data-block-key="3ns9g"&gt;&lt;b&gt;Register for the Google Public Sector Summit&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="abal6"&gt;Beyond the personalized programming and practical hands-on experiences, attendees will leave the event with a customized action plan for driving impact in their organization. &lt;a href="https://events.govexec.com/google-public-sector-summit/register/" target="_blank"&gt;Register today&lt;/a&gt; for the Google Public Sector Summit to gain the strategies, skills, and inspiration to turn your moonshot aspirations into a reality. Join us as we build what’s next for the public sector, together.&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 06 Oct 2026 07:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/public-sector/where-mission-meets-moonshot-join-us-at-the-google-public-sector-summit-2026/</guid><category>Public Sector</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/GPS_Summit2025_BlogPreview_5.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Where mission meets moonshot: Join us at the Google Public Sector Summit 2026</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/GPS_Summit2025_BlogPreview_5.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/public-sector/where-mission-meets-moonshot-join-us-at-the-google-public-sector-summit-2026/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Katharyn White</name><title>Director of Marketing, Public Sector</title><department></department><company>Google Cloud</company></author></item><item><title>AlloyDB: A unified database engine for hybrid search</title><link>https://cloud.google.com/blog/products/databases/simplify-ai-search-with-alloydb-hybrid-search-and-rrf/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For modern AI and RAG applications, achieving high search relevance requires that you combine at least two techniques: vector search for semantic context, and full-text search (FTS) for keyword precision. While AlloyDB for PostgreSQL supports both of these capabilities, managing them has traditionally required a more hands-on operational approach to ensure peak performance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The challenge is not executing the searches, but the subsequent fusion of result sets. Merging results from the vector query (distance scores) and the FTS query (relevance scores) requires complex SQL queries or custom code in the application layer. This often means maintaining a separate system for fusion, score normalization, and re-ranking.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This article details how AlloyDB AI's hybrid search eliminates this complexity. We will explore how recent updates allow you to:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Simplify hybrid search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Consolidate complex SQL queries, or multi-step application workflows into a single, high-performance SQL function powered by Reciprocal Rank Fusion (&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search#rank-fusion"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;RRF&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Optimize FTS performance:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Use the new RUM extension to achieve low-latency relevance ranking and efficient phrase matching by storing word positions directly in the index.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Leverage industry-standard ranking:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Utilize the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/native-bm25-search-in-alloydb-and-cloud-sql"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;new natively supported BM25 index&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for superior keyword-based scoring directly out of the box.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Expand search versatility:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Execute queries against specialized external clusters, including Elasticsearch, OpenSearch, and Solr, using the new external search Foreign Data Wrapper (FDW) without leaving the AlloyDB environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The problem of the multi-step workflow&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Before AlloyDB AI's native solution, achieving robust hybrid search was very demanding, especially for developers trying to keep this logic within the database using standard SQL. This approach required multi-step orchestration that was not only difficult to manage and maintain for two sources, but became virtually impossible to scale as additional sources were added. These steps included:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Executing vector search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Run a query using a vector column or vector index (for AlloyDB this can be a ScaNN index) to find top k results, generating vector scores.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Executing the FTS query:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Run FTS query e.g. by using a generalized inverted index, or GIN, to find top k results.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Normalizing scores (the brittle step):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Write complex SQL logic or custom application code to map both result sets onto a common scale. This logic is prone to breaking when data distributions change.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Performing cross-service joins and re-rankings:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Use application memory to perform a complex &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;FULL OUTER JOIN&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; on document IDs, apply a weighted summation of the normalized scores, and finally sort the combined results in cases where the FTS search was done in a separate system than the one used for vector search.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This decentralized approach led to fragile score logic, increased latency, high operational load, and a dependency on application expertise for maintaining search quality.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--medium
      
      
        h-c-grid__col
        
        h-c-grid__col--4 h-c-grid__col--offset-4
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_XqX6b2x.max-1000x1000.png"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Simplifying hybrid search architectures with AlloyDB&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The key to simplifying this is to adopt RRF, which, because it is inherently rank-based, cleverly bypasses the brittle step of score normalization entirely. RRF consists of two parts:&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;1. The single source of truth&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of a multi-step fetch and join process, you simply call one SQL function, providing your search components as a declarative JSON array:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;SQL&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;SELECT id, score \r\nFROM ai.hybrid_search(\r\n    search_inputs =&amp;gt; ARRAY[\r\n        -- Vector Component (Semantic) with dynamic embedding generation from natural language\r\n        $json${\r\n            &amp;quot;data_type&amp;quot;: &amp;quot;vector&amp;quot;,\r\n            &amp;quot;limit&amp;quot;: 10,\r\n            &amp;quot;table_name&amp;quot;: &amp;quot;documents&amp;quot;,\r\n            &amp;quot;key_column&amp;quot;: &amp;quot;doc_id&amp;quot;,\r\n            &amp;quot;vec_column&amp;quot;: &amp;quot;embedding&amp;quot;,\r\n            &amp;quot;distance_operator&amp;quot;: &amp;quot;&amp;lt;=&amp;gt;&amp;quot;,\r\n            &amp;quot;query_vector&amp;quot;: &amp;quot;ai.embedding(\&amp;#x27;text-embedding-005\&amp;#x27;, \&amp;#x27;alloydb search\&amp;#x27;)::vector&amp;quot;\r\n        }$json$::jsonb,\r\n        \r\n        -- Text Component (Keyword)\r\n        $json${\r\n            &amp;quot;data_type&amp;quot;: &amp;quot;text&amp;quot;,\r\n            &amp;quot;limit&amp;quot;: 10,\r\n            &amp;quot;table_name&amp;quot;: &amp;quot;documents&amp;quot;,\r\n            &amp;quot;key_column&amp;quot;: &amp;quot;doc_id&amp;quot;,\r\n            &amp;quot;text_column&amp;quot;: &amp;quot;content&amp;quot;,\r\n            &amp;quot;query_text_input&amp;quot;: &amp;quot;alloydb search&amp;quot;\r\n        }$json$::jsonb\r\n    ]\r\n);&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42fb92ad0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;(Note: The unified &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ai.hybrid_search&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; API isn't limited to just two components. It also natively supports external search sources via the new FDW).&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Database-native orchestration&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;hybrid_search()&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; function executes the entire workflow in a single query plan, minimizing overhead and helping ensure transaction consistency. It uses:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Dynamic CTE generation:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The function constructs dynamic SQL, creating Common Table Expressions (CTEs) for each component. Each CTE is responsible for calculating the positional rank (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ROW_NUMBER()&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) of its results.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Kernel-level fusion:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; All ranked component results are immediately combined using a &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;FULL OUTER JOIN&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; based on the document ID.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Final RRF score:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The combined ranks are used to calculate the final, unified score using the RRF formula:&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_bAjNPuH.max-1000x1000.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By taking this approach, brittle score calculations and application-side joins become unnecessary. While Reciprocal Rank Fusion (RRF) is our current ranking algorithm, we intend to introduce additional merging and ranking options down the road.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_534j9B2.max-1000x1000.jpg"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The result: Performance and operational simplicity&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Moving from complex external workflows to native SQL functions provides immediate, measurable benefits:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Metric&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-step application workflow (simulated by manual SQL)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB AI-native hybrid_search() UDF&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Code complexity&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Complex score normalization functions, service calls, and application joins.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Zero external logic; a single, declarative SQL call.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Maintenance&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Constant tuning of normalization formulas.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;No tuning required; RRF is rank-based and distribution-agnostic.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Enhancing FTS performance: The RUM extension&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AlloyDB AI's &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;hybrid_search()&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; function is designed to deliver comprehensive search relevance by combining high-performance vector search with FTS. While AlloyDB's native FTS capabilities are powerful, reliance on the standard PostgreSQL GIN index for the text component can present a bottleneck for advanced operations.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The challenge with GIN indexes is that they do not store the positional information of words. This limitation forces costly table scans to re-analyze content for:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Relevance ranking:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Calculating search result scores based on word proximity and frequency&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Phrase searching:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Finding exact word order, which requires positional information.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Introducing the RUM extension for low-latency FTS&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;RUM extension&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; is a powerful index access method based on GIN that directly resolves these performance issues.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;RUM's core advantage:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Unlike the GIN index (which maps &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;word -&amp;gt; [docID]&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;), the RUM index stores the positional information of each word directly within the index (e.g., &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;word -&amp;gt; [(docID1, [pos])], ...&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Performance benefits:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; This allows RUM to perform complex operations like ranking and phrase matching predominantly within the index itself, avoiding expensive heap scans. RUM provides significantly faster relevance ranking and efficient phrase and proximity searches.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Great for hybrid search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; RUM is a critical complement to vector search (like ScaNN) within a hybrid search framework, as GIN’s potential latency makes it less suitable for a real-time hybrid approach.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The RUM extension is an excellent choice for ranking-heavy or high-concurrency search applications. By storing word positions directly in the index, RUM cuts out the need to re-scan table pages during ranking, giving you fast and sorted results. However, it comes with some trade-offs: slower index builds and a larger disk footprint. If your workload prioritizes quick queries and precise ranking over write speeds and storage density, RUM is an investment that's well worth it.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Measurable gains and integration&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Adopting RUM directly bolsters the performance of FTS in AlloyDB, showing considerable gains for raw FTS queries, as well as gains on hybrid search that includes FTS queries. Below are performance statistics showing the speedup between using RUM compared to GIN on various &lt;/span&gt;&lt;a href="https://github.com/beir-cellar/beir" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BEIR&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; datasets.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_rFbNdqQ.max-1000x1000.png"
        
          alt="4"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="7eyx0"&gt;(Performance based on tests using the &lt;a href="https://github.com/beir-cellar/beir"&gt;BEIR Natural Questions benchmark&lt;/a&gt;.)&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;RUM integrates with the specialized &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;&amp;lt;=&amp;gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; distance operator for search and ranking, which is supported for use in hybrid search SQL calls, delivering the best of both worlds. The following &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/next26/alloydb-ai-hybrid-search" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;codelab &lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;shows how to configure and use the RUM index with the Hybrid Search UDF. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Introducing BM25 Index: The modern ranking standard&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We recently introduced the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/create-bm25-index"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BM25 index&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to CloudSQL and  AlloyDB in preview, bringing the industry-standard relevance ranking via the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;pg_textsearch &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;extension&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. BM25 uses &lt;/span&gt;&lt;a href="https://en.wikipedia.org/wiki/Tf%E2%80%93idf" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;TF-IDF&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which accounts for term frequency saturation and document length normalization, delivering significantly higher precision and quality for keyword-based search queries. By leveraging native BM25 indexes within AlloyDB AI, you can achieve superior ranking accuracy out of the box, so that exact-term matches and critical keywords score well without having to implement external search engines or complex custom scoring logic.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Install pg_textsearch extension\r\nCREATE EXTENSION pg_textsearch;\r\n\r\n-- Create the native BM25 index on the content column\r\nCREATE INDEX idx_docs_bm25 \r\nON cymbal_products \r\nUSING bm25 (product_description) \r\nWITH (text_config=&amp;#x27;english&amp;#x27;);\r\n\r\n-- Full text search query\r\nSELECT product_name, product_description &amp;lt;@&amp;gt; &amp;#x27;cherry tree&amp;#x27; AS bm25_score \r\nFROM cymbal_products\r\nORDER BY bm25_score \r\nLIMIT 5;&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f9bba50&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;External search: Extending versatility with foreign data wrappers&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To further extend search capabilities, we also introduced the external search Foreign Data Wrapper (&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/reference/extensions#fdw-extensions"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;FDW&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;) in AlloyDB AI. This lets you execute full-text searches against specialized external clusters, starting with Elasticsearch, OpenSearch, and Solr, and provides several key architectural advantages:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Optimized retrieval:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Leverage ranking algorithms and a richer feature set from dedicated search backends without leaving the AlloyDB environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Unified SQL interface:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Interact with external data, perform joins, and meld results using standard PostgreSQL SQL without losing the expressiveness of advanced FTS queries.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Strong portability:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Maintain existing search infrastructure while benefiting from the simplified hybrid architecture offered by AlloyDB AI.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The following codelabs are e&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;nd-to-end code guides walking through hybrid search with &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/next26/alloydb-elastic?hl=en#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Elasticsearch&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/alloydb-solr-fdw?hl=en#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Solr&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; integrations.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;A unified architecture for AI-powered search&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The true challenge of modern search is not technical in nature, but in achieving architectural simplicity and sustained operational stability. AlloyDB AI addresses this with a unified platform built on three core innovations:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Simplifying hybrid search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; AlloyDB AI’s hybrid search function, powered by RRF, transforms a fragile, multi-step application workflow into a single, high-performance SQL call. This native implementation eliminates the need for complex score normalization and application-side joins, drastically lowering the operational and engineering cost while delivering consistently accurate and fast hybrid relevant results.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Optimizing FTS performance:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; To ensure the FTS component of Hybrid Search meets diverse application demands, AlloyDB AI offers distinct full-text search options. The RUM extension optimizes for low-latency performance by storing positional data directly in the index to enable faster relevance ranking and efficient phrase searches&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; — &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;important when you need fast query speeds. Alternatively, the BM25 index provides industry-standard relevance ranking, and is the preferred choice when prioritizing keyword scoring precision. You can select between these options depending on whether your focus is search latency or ranking precision.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enhancing versatility through external search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The addition of the external search through FDW extends AlloyDB’s reach to specialized search backends like Elasticsearch. This lets you leverage the superior scale and advanced retrieval features of dedicated search clusters while maintaining a familiar PostgreSQL interface. By integrating these external results directly into the hybrid search framework, AlloyDB AI ensures that you can combine even the most massive text repositories with vector-based semantic insights.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By consolidating the complexity of scoring, joining, and re-ranking within the database kernel, and offering a dual-path approach to FTS — leveraging RUM for optimized, low-latency, internal FTS or external search for specialized, scalable backends — AlloyDB AI delivers a robust and self-contained search foundation. This coexistence is essential to the overall hybrid search story, providing the flexibility to choose the optimal path based on workload, data volume, and existing infrastructure. Now, you can focus on building intelligent application features, confident that your search architecture is both highly performant and easy to maintain.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Watch it in action&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Watch how AlloyDB is the &lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=vYB0P9rPXt8" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ultimate hybrid search engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Learn about &lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=-JxQb-kjFHk" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BM25 support in AlloyDB&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started &lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to bring more speed and cost-efficiency to your AI workloads?&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New to AlloyDB?&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Discover AlloyDB with a &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/free-trial-cluster"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;30-day free trial&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started with hybrid search&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Set up a &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/full-text-search-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;text-search index&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and select a &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/choose-index-strategy"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;vector search index&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Once these are created, you can find some &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;examples&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to perform various hybrid search queries.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;External search&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Create a foreign data wrapper and foreign table in AlloyDB to query external data from &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/elastic-search"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Elasticsearch&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/solr-search"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Solr&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/opensearch"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OpenSearch&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Tue, 06 Oct 2026 07:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/databases/simplify-ai-search-with-alloydb-hybrid-search-and-rrf/</guid><category>Databases</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>AlloyDB: A unified database engine for hybrid search</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/databases/simplify-ai-search-with-alloydb-hybrid-search-and-rrf/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Kumar Ramamurthy</name><title>Senior Product Manager, Databases</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Ricky Zhou</name><title>Software Engineer, Databases</title><department></department><company></company></author></item><item><title>Introducing Google Cloud Modernize, transforming for (and with) AI</title><link>https://cloud.google.com/blog/products/infrastructure-modernization/google-cloud-modernize-accelerate-transformation-with-ai/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’re announcing&lt;/span&gt; &lt;a href="https://cloud.google.com/solutions/modernize"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Modernize&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, an end-to-end transformation portfolio to help enterprises collapse multi-year roadmaps with the power of AI.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Cloud Modernize&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; brings together Google Cloud’s proven migration and modernization tools, including Migration Center, Google Cloud VMware Engine, &lt;/span&gt;&lt;a href="https://cloud.google.com/solutions/mainframe-modernization?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Mainframe Modernization&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and our new &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-agentic-migration?_gl=1*gf0ap9*_ga*MzIwOTk1MDYuMTc5MDI2ODQyMQ..*_ga_4LYFWVHBEB*czE3OTAzNDY4NzEkbzQkZzEkdDE3OTAzNDY4NzMkajU4JGwwJGgw&amp;amp;e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;EKS-to-GKE Migration Agent&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, into a single portfolio. With it, enterprise teams now have purpose-built agentic capabilities across infrastructure assessment, platform modernization, and application modernization. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Central to this portfolio is &lt;/span&gt;&lt;a href="https://console.cloud.google.com/modernization-hub"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Modernization Hub&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a newly launched in-console experience where developers and architects can analyze source code, map dependencies, and accelerate modernization for Java, .NET, and mainframe applications. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Agentic infrastructure assessment&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Every transformation begins with an accurate assessment. With new Gemini capabilities in&lt;/span&gt; &lt;a href="https://console.cloud.google.com/migration"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Migration Center&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Agentic Quick Estimator (GA)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; converts VMware inventory exports (such as RVTools) and infrastructure inputs into total cost of ownership (TCO) projections for your Compute Engine environment.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Using an interactive chat interface, teams can test real-time modeling assumptions, such as evaluating multi-region footprints or comparing BYOL licensing against pay-as-you-go, to discover contextual cost optimizations. This condenses weeks of spreadsheet modeling into a defensible business case in minutes.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_7EZHDX9.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="n9zlq"&gt;Migration Center’s Quick TCO Estimator and agentic chat interface&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For specialized support, we also provide a comprehensive modernization assessment at no cost through our&lt;/span&gt; &lt;a href="https://cloud.google.com/solutions/cloud-migration-program"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Rapid Migration &amp;amp; Modernization Program (RaMP)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Platform modernization&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With the rise of real-time AI agents querying backend systems, workloads increasingly require high-throughput infrastructure that removes I/O bottlenecks. Once you’ve defined your target environment using our assessments, there are several new purpose-built compute options for mission-critical workloads:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span&gt;&lt;strong style="vertical-align: baseline;"&gt;SAP S/4HANA at scale (&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/memory-optimized-machines"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;X5 Series&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;, GA):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Delivers single-node 43 TiB memory configurations that remove the previous 29 TiB ceiling, allowing enterprise ERP estates to run without distributed partitioning overhead.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Core-optimized database performance (&lt;/strong&gt;&lt;a href="https://cloud.google.com/blog/products/compute/compute-engine-m4n-vms?e=48754805"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;M4N Series, GA&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Delivers 26.57 GiB RAM per vCPU paired with Hyperdisk Extreme. This prevents organizations from overprovisioning compute cores to meet memory requirements, reducing software licensing costs by more than 20% for Oracle and other core-licensed databases.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span&gt;&lt;strong style="vertical-align: baseline;"&gt;Ultra-low latency data engines (&lt;/strong&gt;&lt;a href="https://cloud.google.com/blog/products/compute/storage-optimized-z4d-vm-and-bare-metal-instances"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Z4D, GA&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; and Z4M, Preview):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Deliver up to 84,000 GiB and 168,000 GiB of high-speed local NVMe SSD respectively, 400 Gbps networking for both Z4D and Z4M, and RDMA support for Z4M. This throughput reduces I/O wait times and prevents query timeouts when real-time AI agents query vector stores, operational databases, and large-scale data pipelines.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Cloud elasticity for VMware environments&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For organizations operating VMware estates that need the elasticity of the cloud but aren’t ready to re-architect their environment, the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Cloud self-managed VMware solution &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;provides administrative control on Bare Metal Z3 shapes with VMware Cloud Foundation (VCF) 9.1. Native global VPC links connect VMware estates directly to Compute Engine, &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine?utm_source=gemini"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GKE), BigQuery, and Gemini Enterprise, allowing teams to ground autonomous agents in operational data without code changes.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Automated container transitions to GKE&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The new &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;EKS-to-GKE Agentic Migration&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (Public Preview) automates transitions from AWS Elastic Kubernetes Service to&lt;/span&gt; &lt;a href="https://cloud.google.com/kubernetes-engine?utm_source=gemini"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; thanks to: &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;An automated pipeline:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Manages discovery, Kubernetes manifest translations, storage and network mappings across clouds.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enterprise-grade security:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Built-in Human-in-the-Loop (HITL) approval gates and in-memory credential security maintain strict GitOps compliance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;A fast-track to modern runtimes:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Quickly moves workloads to GKE to take advantage of low-latency model serving, autoscaling, and multi-agent orchestration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/customers/netease-games"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;NetEase Games&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; demonstrated the value of this platform approach by containerizing services on GKE, reducing infrastructure scaling times from hours to five minutes during peak launches while cutting server costs by 40%:&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"By integrating diverse computing instances and powerful orchestration, we have transformed our infrastructure into a competitive advantage, ensuring NetEase remains a leader in the global gaming market."&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;  - &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Deng Ding&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Director of Site Reliability Engineering, NetEase Games&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Application modernization with Modernization Hub&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Landing on modern infrastructure enables teams to unlock legacy business logic and modernize their core applications. &lt;/span&gt;&lt;a href="https://console.cloud.google.com/modernization-hub"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Modernization Hub&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; centralizes several modernization tools directly inside the Google Cloud console. &lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;.NET and Java modernization&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Modernization Hub integrates the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Cloud&lt;/strong&gt; &lt;a href="https://docs.cloud.google.com/migration-center/docs/app-modernization-assessment?utm_source=gemini"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;App Modernization CLI&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; (CodMod)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, which uses Gemini to analyze large source code repositories, understand legacy application architectures,maps hidden dependencies, identifies modernization challenges and generates modernization recommendations. This enables customers to migrate legacy .NET Framework applications to modern .NET Core  running on Linux containers, reducing OS licensing overhead, while also making these applications and data accessible to modern AI agent workflows.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/Modernization_Hub_Take3.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="n9zlq"&gt;Modernization Hub’s in-console interface&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Mainframe modernization&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For customers running mainframes, we provide specialized solutions to help accelerate and de-risk end-to-end application transformation to Google Cloud:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mainframe-assessment-tool/docs/overview?utm_source=gemini"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Mainframe Assessment Tool&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Parses legacy mainframe codebases, extracts business rules, maps application and data dependencies to generate cloud-ready target application specifications and power agentic modernization workflows.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mainframe-dual-run/docs/dual-run-overview?utm_source=gemini"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Dual Run&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Substantially reduces cutover risk by replaying live production transaction streams simultaneously across the mainframe and the new cloud applications, verifying functional equivalence before going live. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mainframe-connector/docs/overview?utm_source=gemini"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Mainframe Connector&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Copies mainframe data directly into Google Cloud services (BigQuery, AlloyDB, GCS and others), , unlocking legacy data and supports hybrid architectures.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=KxCIByjwVJg&amp;amp;utm_source=gemini" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Intesa Sanpaolo&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; used Google Cloud mainframe modernization solutions to accelerate their core banking transformation off the mainframe and onto Google Cloud:&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“To confidently move forward with Mainframe Modernization, we will need to provide confidence and assure the bank's leadership and internal control units as well as get approval from the regulators. One of the enablers for this is Google Cloud Dual Run, which is gradually providing the evidences to build such confidence to all three groups.”&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; - &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Claudio Balbo&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Head of IT Architecture, Intesa Sanpaolo&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;A proven track record and partner ecosystem&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/customers/deutsche-boerse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Deutsche Börse Group&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; transitioned its mission-critical SAP S/4HANA environment and DAX® index calculations to Google Cloud, cutting disaster recovery times from hours to minutes, reducing aggregation latency by over 50%, and shortening development cycles from months to days.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To deliver these capabilities at enterprise scale, we are also collaborating with our global partner ecosystem to integrate Google Cloud Modernize with enterprise delivery frameworks, including &lt;/span&gt;&lt;a href="https://www.cognizant.com/us/en" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cognizant&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“Google’s modernization offering heralds the next frontier of AI-native modernization — enabling enterprises to autonomously decode legacy complexity and accelerate into cloud-first, intelligent architectures. Coupled with Cognizant's AI-led, governed delivery engine, we amplify this shift through agentic business processing, unlocking faster, outcome-driven value at scale.” &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;- &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Nishant Upadhyaya&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, Global Practice Head – Digital Engineering, Cognizant&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Accelerate your transformation now&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Modernizing your infrastructure and applications is the baseline requirement for enterprise AI agility. Google Cloud Modernize equips teams to navigate each phase safely and efficiently as they pursue AI readiness and adoption.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Get started:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Join the webinar:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Register for our &lt;/span&gt;&lt;a href="https://cloudonair.withgoogle.com/events/adaptive-enterprise-accelerating-app-mod-infra" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;upcoming session on November 17&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to see live demos of Modernization Hub and our agentic migration tools.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Explore the console:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Log into the Google Cloud Modernize console to access &lt;/span&gt;&lt;a href="https://console.cloud.google.com/modernization-hub"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Modernization Hub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.   &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Start planning:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Visit the Google Cloud Modernize&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/solutions/modernize"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;webpage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or request a &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/migration-assessment-offer"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;modernization assessment&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; at no cost today.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Mon, 05 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/infrastructure-modernization/google-cloud-modernize-accelerate-transformation-with-ai/</guid><category>Application Modernization</category><category>Cloud Migration</category><category>Infrastructure Modernization</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing Google Cloud Modernize, transforming for (and with) AI</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/infrastructure-modernization/google-cloud-modernize-accelerate-transformation-with-ai/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Souvik Choudhury</name><title>Senior Director, Product Management, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Tom Nikl</name><title>Cloud Modernization &amp; Migration Team, Google Cloud</title><department></department><company></company></author></item><item><title>AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer</title><link>https://cloud.google.com/blog/products/containers-kubernetes/ai21-trains-its-models-on-ai-hypercomputer/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;&lt;strong&gt;Editor’s note&lt;/strong&gt;: AI21 Labs is a leading global AI lab with a long track record of building foundation models, most notably the Jamba family, and today focuses on specialized LLMs and agent optimization technology. By adopting Google Cloud AI Hypercomputer, AI21 cut high-priority job wait times from 72 hours to 12 and manual scheduling interventions from 20 per week to zero. &lt;/span&gt;&lt;/p&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At &lt;/span&gt;&lt;a href="https://www.ai21.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AI21&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, we build foundation models and agent optimization products that help enterprises run agents at frontier quality, efficiently. Our language models, including the &lt;/span&gt;&lt;a href="https://www.ai21.com/jamba/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Jamba&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; family, and our agent optimization product suite run demanding production workloads, including our own. We chose &lt;/span&gt;&lt;a href="https://cloud.google.com/ai-infrastructure"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud AI Hypercomputer&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to support them at scale.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To keep our model training runs highly utilized, we needed a performant, scalable environment codesigned across infrastructure, orchestration, and consumption models. Our model training runs on one of our shared &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GKE) clusters, pooling thousands of Google Cloud A3 (powered by NVIDIA H100 Tensor Core GPUs) and A3 Ultra (powered by NVIDIA H200 Tensor Core GPUs) instances, so any team can draw on the full capacity of the fleet rather than being boxed into its own slice. The cluster also trains models and agent-optimization workloads beyond the Jamba family. That approach keeps utilization high, and it makes scheduling hard. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The gridlock of high-utilization clusters&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Prior to leveraging GKE for orchestration, we used to negotiate capacity by hand in Slack. If you needed capacity for a training run, you posted in #gpu-resources and hoped for the best.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;That worked fine when the cluster had headroom. It stopped working once utilization pinned near 100%, which is where you want a reserved compute fleet to sit.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Over time, every request became a negotiation. Team leads spent their time refereeing compute disputes. Our high-priority jobs — the large, multi-node training runs that need half or more of the cluster at once and serve as the critical path for model projects — could sit blocked for up to 72 hours waiting for enough contiguous capacity to open up. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_5g4QmVa.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="cwh8v"&gt;Figure 1: A typical day on #gpu-resources — manual requests, ad-hoc coordination, and researchers waiting on replies.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Scarcity created two distinct problems, and it took us a while to see them as separate. The first was contention: determining who gets compute access next, which we resolved through negotiation. The second was fragmentation: capacity that was technically free but scattered in pieces too small for a large job to use, a bin-packing problem no amount of negotiation could fix. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image2_vk4hCkH.max-1000x1000.png"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="cwh8v"&gt;Figure 2: GPU fragmentation. 8 free GPUs scattered as 1+1+4+2 across four nodes can’t fit an 8-GPU workload.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sometimes we had plenty of capacity free on paper, but it was scattered across different machines in chunks too small for a larger job to actually land. Without all-or-nothing admission, the cluster could reach a deadlock, with machines holding resources without doing useful work until someone stepped in manually. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;It was clear the status quo wasn’t working and we needed something better.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Choosing the right scheduler&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We looked at a few open-source batch schedulers, including Apache YuniKorn, Volcano, and &lt;/span&gt;&lt;a href="https://kueue.sigs.k8s.io/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Kueue&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. YuniKorn didn’t cover all our use cases. And while Volcano had more features, integrating it with our environment would have required replacing core Kubernetes scheduler components. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Kueue won on simplicity and integration. It worked with standard Kubernetes, didn’t require replacing core components, and didn’t force us to rewrite our job specs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Pairing Kueue with AI Hypercomputer’s flexible, open operations through GKE also contributed to our success. We rely on GKE because it gives us the right level of control for compute-intensive AI work — close access to GPU hardware and drivers, without the overhead of managing raw instances ourselves. Kueue’s native integration with GKE, including with Google Cloud capacity types like&lt;/span&gt; &lt;a href="https://cloud.google.com/spot-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spot VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and&lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/introducing-dynamic-workload-scheduler"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Dynamic Workload Scheduler&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, meant that once our reserved capacity filled up, the same scheduling logic could reach out to elastic capacity automatically instead of leaving jobs stuck.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Closing the loop with open source&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Adopting Kueue turned into something bigger than a simple process change. Because we partner with Google Cloud, we have a direct line to a Technical Account Manager, who saw an opportunity to make AI21 a design partner for the Kueue team. This way, we wouldn’t just be a user, but a source of real production requirements that could help shape where the tool went next. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;One of the first things to come out of that partnership was a requirements document we shared with the Kueue team, describing behavior we needed that didn’t exist yet: fair admission ordering across teams for multi-node jobs, without the preemption that usually comes bundled with fairness. The Kueue team built it.&lt;/span&gt; &lt;a href="https://kueue.sigs.k8s.io/docs/concepts/admission_fair_sharing/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Admission Fair Sharing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (AFS) reorders the admission queue to favor teams that have historically used less capacity without disrupting jobs already running. Around the same time, we also turned on&lt;/span&gt; &lt;a href="https://kueue.sigs.k8s.io/docs/concepts/topology_aware_scheduling/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Topology Aware Scheduling&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. This existing feature made Kueue aware of our physical cluster layout, so it can refuse to admit jobs that won’t fit on a single node rather than placing them with nowhere to run. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For us, that’s what the best-case open-source feedback loop looks like: Real requirements surfaced through production use, fed directly back into Kueue.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Same fleet, less friction&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Slack &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;#gpu-resources channel is archived now. Every workload, whether a debug pod, a multi-node training run, or an inference deployment, gets queued, prioritized, and scheduled automatically. The results were immediate. Manual interventions dropped from about 20 a week to zero. High-priority jobs that used to wait up to 72 hours now wait 12. Fragmentation across the cluster fell from 15% to 8%, and the “zombie job” problem — workloads partially admitted with nowhere to run — is gone.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;None of this changed our total cost. We run our reserved fleet at close to 100% utilization on purpose, so raw spend was never the variable we were optimizing. What changed is where our people’s time goes. Team leads aren’t refereeing compute disputes anymore, and researchers aren’t waiting on replies in Slack. That frees everyone to run more experiments and iterate faster.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Fri, 02 Oct 2026 19:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/containers-kubernetes/ai21-trains-its-models-on-ai-hypercomputer/</guid><category>AI Hypercomputer</category><category>GKE</category><category>AI infrastructure</category><category>Customers</category><category>Containers &amp; Kubernetes</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>AI21 achieves an 83% reduction in time-to-start for AI workloads with AI Hypercomputer</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/containers-kubernetes/ai21-trains-its-models-on-ai-hypercomputer/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Barak Peleg</name><title>VP, Technology and Architecture, AI21</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Asaf Ben-Tovim</name><title>DevOps Engineer, AI21</title><department></department><company></company></author></item><item><title>Announcing Spanner queues: Transactional messaging for agentic workloads and beyond</title><link>https://cloud.google.com/blog/products/databases/spanner-queues-provide-native-transactional-messaging/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI agents don't just answer queries — they can autonomously issue refunds, manage inventory, execute multi-step handoffs, and orchestrate sub-agents, to name but a few complex agentic workflows. That frequently requires the agents to maintain internal state in an operational database while dispatching asynchronous actions through a separate messaging or event queue system. And unfortunately for the teams building these applications, managing two systems with disjointed commit points destroys transactional consistency in agentic systems.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we are excited to announce the general availability of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/queues/queues-overview"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner queues&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: native transactional messaging embedded directly within Spanner. Designed specifically for reliable agentic execution, with Spanner queues, creating a message is simply another write in your transaction. An agent's state change and intended downstream actions commit together atomically or fail completely.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Compare this to traditional approaches: When a database state update succeeds, but the action dispatch fails, your AI agent decides to act but doesn’t execute. If the message dispatch succeeds but the state transaction rolls back, your agent executes an action based on an invalid state. In asynchronous multi-agent coordination, retries, speculation, and race conditions amplify these failures, forcing developers to build complex outbox patterns, idempotency layers, and reconciliation workers — a heavy reliability tax on agentic architecture.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Key capabilities for agentic architectures&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Spanner queues introduces several core capabilities to support these complex asynchronous workflows without introducing infrastructure overhead.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Atomic decide-and-act enqueue.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Within a single Spanner read-write transaction, agents can update internal memory or state tables and enqueue tasks to peer agents simultaneously. Backed by Spanner's strict serializability and global external consistency, state changes and execution intent commit as a single atomic unit.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Scheduled execution and delays.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Queue messages can be dispatched immediately upon commit or scheduled for future delivery. Agentic patterns such as delayed retries, scheduled agent check-ins, or SLA escalation timers can be enqueued transactionally alongside memory updates without requiring external cron schedulers or polling infrastructure.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Streaming SQL pull for agent workers.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Autonomous agents consume tasks dynamically using streaming SQL reads. Agent runtimes stream incoming tasks, process them as capacity becomes available, and acknowledge task completion within a transaction to guarantee end-to-end task execution state reliability.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Episodic memory persistence and handoffs.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Agent memory updates — including long-term episodic summaries, reflective state transitions, and context handoffs across sub-agents — can be persisted asynchronously and transactionally via queues. This guarantees that an agent's internal memory remains fully synchronized with its execution history without blocking real-time interactive turns.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Why Spanner queues are essential for autonomous agents&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Transactional exactly-once agent execution: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;When an agent evaluates tool call results, state modification, reasoning persistence, and downstream tool invocation tasks are saved in one transaction. We guarantee at-least-once delivery and at-most-once ACK, enabling you to achieve exactly-once processing.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Robust multi-agent orchestration and handoffs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; In multi-agent systems (A2A), handing off state from a primary agent to a specialist agent is represented as a durable, transactionally committed message. Specialist agent workers receive deliverable messages while maintaining a fully auditable lineage of agent interactions.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;First-class timeouts and human-in-the-loop workflows:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Agent workflows frequently require pausing for human approvals or scheduled follow-ups. Spanner queues handles timeout management natively: A single transaction records the pending approval state and schedules an automated escalation message, resolving whichever triggers first.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;SQL-native observability for agent queues:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Inspecting in-flight agent workloads, monitoring task backlogs, or auditing agent execution history can be accomplished with standard SQL queries over queue tables, avoiding opaque message-store black boxes.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Beyond agentic workflows&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Beyond AI agents, Spanner queues serves as a flexible, multi-purpose messaging platform across a variety of traditional event-driven architectures. Whether powering real-time activity feeds in social applications, delivering live updates in news publishing, orchestrating order processing and inventory workflows in retail, or handling high-throughput asynchronous task processing and transaction notifications in financial services, Spanner queues provides a robust foundation for asynchronous message delivery and transactional event-driven workflows within your primary database.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Under the hood: Transactional mechanics of Spanner queues&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because Spanner queues are represented as first-class relational structures in Spanner, you define, inspect, and manage queues using familiar GoogleSQL.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;1. Defining a queue and enqueuing atomically&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When an agent decides to approve a customer refund, it updates the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Orders&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; table and dispatches an execution task to the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;OrderAgentTasks&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; queue within a single ACID transaction:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;-- Assume parent table:\r\n-- CREATE TABLE Orders (\r\n--   OrderId STRING(64) NOT NULL, ...\r\n-- )\r\n-- PRIMARY KEY (OrderId);\r\n\r\n-- Define the transactional queue table\r\nCREATE QUEUE OrderAgentTasks (\r\n  OrderId STRING(64) NOT NULL,\r\n  TaskId STRING(64) NOT NULL,\r\n  TaskType STRING(64) NOT NULL,\r\n  Payload JSON NOT NULL\r\n) PRIMARY KEY (OrderId, TaskId, TaskType), INTERLEAVE IN Orders;\r\n\r\n-- Inside a Read-Write Transaction:\r\n-- 1. Update business state atomically\r\n\r\n-- BEGIN TRANSACTION;\r\n\r\nUPDATE Orders\r\nSET Status = \&amp;#x27;REFUND_APPROVED\&amp;#x27;,\r\n    UpdatedAt = PENDING_COMMIT_TIMESTAMP()\r\nWHERE OrderId = @orderId;\r\n\r\n-- 2. Enqueue the asynchronous agent action in the same transaction\r\nINSERT INTO OrderAgentTasks (OrderId, TaskId, TaskType, Payload)\r\nVALUES (\r\n  @orderId,\r\n  @taskId,\r\n  \&amp;#x27;EXECUTE_REFUND\&amp;#x27;,\r\n  JSON \&amp;#x27;{&amp;quot;action&amp;quot;: &amp;quot;execute_refund&amp;quot;, &amp;quot;amount&amp;quot;: 49.99}\&amp;#x27;\r\n);\r\n\r\n-- COMMIT;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cd2e010&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This ensures that the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;EXECUTE_REFUND&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; task exists if and only if the order status successfully transitioned to &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;REFUND_APPROVED&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;2. Temporal scheduling and atomic cancellation&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For workflows that depend on time — such as waiting up to 72 hours for a manager's approval before escalating — agents populate the system &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;DeliverTime&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; column to defer message visibility:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;-- Schedule an automated escalation check-in 72 hours in the future\r\nINSERT INTO OrderAgentTasks (OrderId, TaskId, TaskType, Payload, DeliverTime)\r\nVALUES (\r\n  @orderId,\r\n  @escalationTaskId,\r\n  \&amp;#x27;ESCALATE_UNAPPROVED_ORDER\&amp;#x27;,\r\n  JSON \&amp;#x27;{&amp;quot;action&amp;quot;: &amp;quot;escalate_to_supervisor&amp;quot;}\&amp;#x27;,\r\n  TIMESTAMP_ADD(CURRENT_TIMESTAMP(), INTERVAL 72 HOUR)\r\n);&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cd2edd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If the manager approves the request after four hours, your application doesn't have to deal with phantom escalation alerts firing  days later. In a single transaction, you update the order status and cancel the pending escalation task using a standard SQL &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;DELETE&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- BEGIN TRANSACTION;\r\n\r\nUPDATE Orders\r\nSET Status = &amp;#x27;MANAGER_APPROVED&amp;#x27;,\r\n    ApprovedBy = @managerId\r\nWHERE OrderId = @orderId;\r\n\r\n-- Atomically cancel the pending delayed escalation task. `ASSERT_ROWS_MODIFIED 1` will\r\n-- act as a safeguard and cause a statement level error.\r\n\r\n-- A statement level error can be captured and the transaction can be\r\n-- user-aborted/cancelled in case the queue entry was already deleted.\r\n-- Otherwise the transaction will complete successfully regardless if the \r\n-- queue entry still exists and you&amp;#x27;ll only know if a queue entry was deleted by\r\n-- checking the number of rows affected by the DELETE.\r\nDELETE FROM OrderAgentTasks\r\nWHERE OrderId = @orderId\r\n  AND TaskId = @escalationTaskId\r\n  AND TaskType = &amp;#x27;ESCALATE_UNAPPROVED_ORDER&amp;#x27;,\r\nASSERT_ROWS_MODIFIED 1;\r\n\r\n-- COMMIT;&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cd2de90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;3. Streaming consumption, lease renewal, and atomic acknowledgment&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Downstream agent workers consume tasks using the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;RECEIVE_&amp;lt;QueueName&amp;gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; table-valued function (TVF) over a streaming SQL connection (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ExecuteStreamingSql&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;). Spanner automatically manages message leases, returning a unique &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;SpannerLeaseToken&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and expiration timestamp with each leased task:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Step 1: Consume tasks via a long-lived streaming SQL query\r\nSELECT\r\n  OrderId,\r\n  TaskId,\r\n  TaskType,\r\n  Payload,\r\n  SpannerLeaseToken,\r\n  SpannerLeaseExpirationTimestamp\r\nFROM RECEIVE_OrderAgentTasks(max_duration =&amp;gt; &amp;#x27;20m&amp;#x27;);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cd2c790&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because AI agent tasks often involve multi-turn LLM reasoning or external API calls that take longer than default lease windows, workers can actively extend their lease using the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;RENEWLEASE_&amp;lt;QueueName&amp;gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; function:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;-- Step 2: Extend lease ownership during long-running agent execution\r\nSELECT *\r\nFROM RENEWLEASE_OrderAgentTasks(lease_tokens =&amp;gt; [@leaseToken]);&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cd2cfd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When the agent finishes executing its external tool (passing &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;TaskId&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; as the external API's idempotency key), it opens a read-write transaction to record the final state and acknowledge the message by deleting it with &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ASSERT_ROWS_MODIFIED 1&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Step 3: Atomically checkpoint agent results and ACK the message\r\n-- BEGIN TRANSACTION;\r\n\r\nUPDATE Orders\r\nSET RefundTransactionId = @externalRefundId,\r\n    Status = &amp;#x27;REFUND_COMPLETED&amp;#x27;\r\nWHERE OrderId = @orderId;\r\n\r\nDELETE FROM OrderAgentTasks\r\nWHERE OrderId = @orderId\r\n  AND TaskId = @taskId\r\n  AND TaskType = @taskType\r\nASSERT_ROWS_MODIFIED 1;\r\n\r\n-- COMMIT;&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cd2ea10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Using &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ASSERT_ROWS_MODIFIED 1&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; protects your system against lease-expiration races. If a worker stalled due to a network pause and its lease expired, another worker may have already processed and deleted the task. When the stalled worker resumes and attempts to execute the statement, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ASSERT_ROWS_MODIFIED 1&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;  detects that the queue row is already gone and throws a statement-level error. Catching that error and aborting that transaction prevents stale workers from overwriting newer database state.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Spanner change streams vs. Spanner queues&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Spanner change streams capture database data changes (inserts, updates, and deletes) in near real-time for downstream integration and auditing.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;While both change streams and queues allow applications to react to data changes, Spanner change streams are designed for continuous change data capture (CDC) and data streaming to downstream analytics or storage. In contrast, Spanner queues are explicitly designed for transactional task orchestration, supporting native message leases, scheduled deliveries, SQL-based pulling, and atomic acknowledgments within read-write transactions.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Spanner provides a unified foundation for agentic data, combining relational, hybrid search, graph, and key-value capabilities under strict global consistency. Spanner queues completes the agentic loop by enabling agents to transition seamlessly from reasoning over data to executing transactional actions within a single unified platform. Spanner queues are now generally available. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Sign up for the &lt;/span&gt;&lt;a href="https://console.cloud.google.com/freetrial?redirectPath=%2Fspanner&amp;amp;facet_utm_source=google&amp;amp;facet_utm_campaign=%28organic%29&amp;amp;facet_utm_medium=organic&amp;amp;facet_url=https%3A%2F%2Fcloud.google.com%2Fblog%2Fproducts%2Fspanner%2Ftry-cloud-spanner-databases"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner 90-day free trial&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and read up our &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner/docs/queues/queues-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;public documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to start building resilient, exactly-once agentic workloads by creating a queue table in your Spanner database today.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Fri, 02 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/databases/spanner-queues-provide-native-transactional-messaging/</guid><category>Spanner</category><category>Databases</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_bE1RXzx.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Announcing Spanner queues: Transactional messaging for agentic workloads and beyond</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_bE1RXzx.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/databases/spanner-queues-provide-native-transactional-messaging/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Nitin Sagar</name><title>Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Matthew Mucklo</name><title>Engineering Manager &amp; Lead Cloud Spanner Data Streaming</title><department></department><company></company></author></item><item><title>GKE CPU startup boost: Accelerate app starts without over-provisioning</title><link>https://cloud.google.com/blog/products/containers-kubernetes/gke-cpu-startup-boost-faster-pod-starts-lower-costs/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether you’re launching microservices in response to sudden traffic spikes, deploying new software releases, or scaling up application replicas, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;pod startup time&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; is critical to maintaining a fast, responsive user experience for applications running on Google Kubernetes Engine (GKE).&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_mhC0eeP.max-1000x1000.png"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="0udcy"&gt;Yet, platform engineers and developers face a persistent dilemma: Applications often demand significantly more CPU power during startup than they do during steady-state operations. Sizing CPU requests for normal, steady-state usage leads to CPU throttling during launch, which can result in sluggish cold starts and readiness probe timeouts. On the flip side, over-provisioning baseline CPU requests to satisfy short-lived startup bursts wastes valuable compute resources, inflating infrastructure bills.&lt;/p&gt;&lt;p data-block-key="ciu5o"&gt;Today, we are excited to announce &lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/boost-application-startup"&gt;&lt;b&gt;CPU startup boost&lt;/b&gt;&lt;/a&gt; for GKE in preview. Integrated directly into GKE's Vertical Pod Autoscaler (VPA), CPU startup boost dynamically elevates a container's CPU allocation during initialization and seamlessly scales it back to baseline steady-state levels once the application is ready - &lt;b&gt;all without restarting your containers.&lt;/b&gt;&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Why modern applications need extra CPU at boot time&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When a new container launches, it may perform intensive initialization tasks before it begins serving user requests. Depending on your tech stack, the following startup workloads require substantial CPU cycles:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Java JVM applications&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Frameworks like Spring Boot require high CPU burst capacity for class loading, classpath scanning, instantiating dependency injection containers, and running Just-in-Time (JIT) compilation.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Node.js servers&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Apps parse JavaScript files, build complex module dependency trees (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;require&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;/&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;import&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), and execute V8 engine optimization and JIT compilation passes during initial execution.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Python and AI/ML microservices&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: These services spend initial cycles importing heavy libraries (such as PyTorch, NumPy, or LangChain), compiling &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.pyc&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; bytecode, establishing ORM database schemas, and pre-loading cache structures.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you size CPU requests strictly for steady-state performance, these initialization workloads experience CPU throttling on launch, delaying readiness probes. To prevent slow cold starts, teams frequently overprovision CPU requests. However, once the application stabilizes, those extra CPU resources sit idle, increasing your cloud spend without adding value.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_zZExLw4.max-1000x1000.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;h3 data-block-key="0udcy"&gt;&lt;b&gt;How CPU startup boost can help&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="8k1a3"&gt;CPU Startup Boost solves this by giving your workloads temporary vCPU "boosts" during launch, and automatically returning them to baseline once initialization completes.&lt;/p&gt;&lt;p data-block-key="c9jmj"&gt;Key benefits:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="96uhr"&gt;&lt;b&gt;Faster cold starts&lt;/b&gt;: Reduce application initialization times by up to &lt;b&gt;2x&lt;/b&gt;, accelerating auto-scaling responsiveness during unexpected traffic surges.&lt;/li&gt;&lt;li data-block-key="f852p"&gt;&lt;b&gt;Optimized cloud spend&lt;/b&gt;: Right-size steady-state CPU requests to fit actual runtime needs rather than paying for idle startup headroom.&lt;/li&gt;&lt;li data-block-key="5opq6"&gt;&lt;b&gt;Zero pod restarts&lt;/b&gt;: Dynamic resource resizing happens live inside the running container.&lt;/li&gt;&lt;li data-block-key="3r67d"&gt;&lt;b&gt;Flexible policy controls&lt;/b&gt;: Apply simple pod-level multiplier factors (e.g., 2x CPU during startup) or define granular, container-specific rules for complex multi-container pods.&lt;/li&gt;&lt;/ul&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Under the hood: Kubernetes In-place Pod Resize&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Historically, changing a pod's resource requests or limits required deleting and recreating the pod. This disruptive process triggered container restarts, cache invalidation, and node rescheduling overhead.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To fix that, CPU startup boost builds on Kubernetes &lt;/span&gt;&lt;a href="https://kubernetes.io/docs/tasks/configure-pod-container/resize-container-resources/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;In-place Pod Resize (IPPR)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Tracked under &lt;/span&gt;&lt;a href="https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/1287-in-place-pod-resize" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;KEP-1287&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, IPPR introduced dynamic, in-place resource mutation. Introduced as Alpha in Kubernetes 1.27, promoted to Beta in v1.33, and graduating to General Availability (GA) in v1.35, IPPR allows the Kubernetes control plane and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kubelet&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to update container CPU and memory requests on running pods &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;without restarting the container process&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;GKE leverages IPPR within the VPA  to apply startup CPU boosts at pod admission and smoothly step them down post-readiness.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;How CPU startup boost works (pod lifecycle overview)&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;CPU startup boost operates across three distinct phases:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Admission phase&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: When you deploy a pod, the VPA admission webhook intercepts the creation request. It calculates the elevated CPU request based on your policy (e.g., 2x multiplier or +2 vCPUs) and injects the boosted CPU request along with tracking annotations into the pod spec.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Startup phase&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The pod is scheduled and initialized with the boosted CPU allocation. Your application completes class loading, JIT compilation, or module parsing at top speed without experiencing CPU throttling.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Unboosting phase&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: As soon as the pod's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;readinessProbe&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; passes (plus any configured &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;durationSeconds&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; cooldown delay), the VPA updater issues an in-place resize request. The CPU request steps back down to your baseline level while the container continues running uninterrupted.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Prerequisites and availability&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;CPU startup boost is available today in preview on GKE:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GKE version&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Version &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;1.36.0-gke.4447000&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or later on Standard and Autopilot clusters.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cluster Modes&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Enabled natively on GKE Autopilot (VPA is active by default). On GKE Standard, simply ensure Vertical Pod Autoscaling (VPA) is enabled.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Workload Support&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Works with standard Kubernetes controllers, including Deployments and StatefulSets.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started: Configuring CPU startup boost&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Configuring CPU startup boost is as simple as adding a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;startupBoost&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; section to your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;VerticalPodAutoscaler&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; manifest.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Example 1: Pod-level boost with fixed steady-state (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;updateMode: "Off"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;)&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you want to use VPA purely for startup boost while keeping steady-state CPU requests locked to your manifest definitions, set &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;updateMode&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;"Off"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;apiVersion: &amp;quot;autoscaling.k8s.io/v1&amp;quot;\r\nkind: VerticalPodAutoscaler\r\nmetadata:\r\n  name: java-app-startup-boost\r\n  namespace: default\r\nspec:\r\n  targetRef:\r\n    apiVersion: &amp;quot;apps/v1&amp;quot;\r\n    kind: Deployment\r\n    name: java-app\r\n  updatePolicy:\r\n    updateMode: &amp;quot;Off&amp;quot;\r\n  startupBoost:\r\n    cpu:\r\n      type: &amp;quot;Factor&amp;quot;\r\n      factor: 2\r\n      durationSeconds: 10&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f457bd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;In this example, GKE doubles the container's CPU request during launch and holds the boosted allocation for 10 seconds after readiness probes pass before scaling back to baseline.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Example 2: Combining startup boost with continuous VPA auto-scaling&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you want GKE to boost CPU during launch &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;and&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; continuously optimize steady-state resources post-startup, set &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;updateMode&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;"InPlaceOrRecreate"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;apiVersion: &amp;quot;autoscaling.k8s.io/v1&amp;quot;\r\nkind: VerticalPodAutoscaler\r\nmetadata:\r\n  name: nodejs-app-vpa\r\n  namespace: default\r\nspec:\r\n  targetRef:\r\n    apiVersion: &amp;quot;apps/v1&amp;quot;\r\n    kind: Deployment\r\n    name: nodejs-service\r\n  updatePolicy:\r\n    updateMode: &amp;quot;InPlaceOrRecreate&amp;quot;\r\n  startupBoost:\r\n    cpu:\r\n      type: &amp;quot;Factor&amp;quot;\r\n      factor: 3\r\n      durationSeconds: 15&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f4573d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Example 3: Granular container-level boost&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For pods running sidecars (such as logging agents or service mesh proxies) that do not require extra CPU on boot, target specific app containers:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;apiVersion: &amp;quot;autoscaling.k8s.io/v1&amp;quot;\r\nkind: VerticalPodAutoscaler\r\nmetadata:\r\n  name: app-container-boost\r\nspec:\r\n  targetRef:\r\n    apiVersion: &amp;quot;apps/v1&amp;quot;\r\n    kind: Deployment\r\n    name: API-gateway\r\n  updatePolicy:\r\n    updateMode: &amp;quot;Off&amp;quot;\r\n  resourcePolicy:\r\n    containerPolicies:\r\n    - containerName: &amp;quot;web-app&amp;quot;\r\n      mode: &amp;quot;Off&amp;quot;\r\n      startupBoost:\r\n        cpu:\r\n          type: &amp;quot;Quantity&amp;quot;\r\n          quantity: &amp;quot;2&amp;quot;\r\n          durationSeconds: 5&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f455f90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Verifying startup boost in your cluster&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can verify that GKE applied and downscaled the startup boost using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kubectl&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;1. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Inspect pod annotations&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Check for the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;vpaCpuStartupBoost&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; tracking annotation:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;kubectl get pod &amp;lt;POD_NAME&amp;gt; -o yaml&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f455290&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Look for annotations indicating the original baseline and boosted CPU requests.&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;2. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Monitor in-place resize events&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Confirm that GKE downscaled the CPU request back to baseline after readiness:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;kubectl get events --field-selector reason=InPlaceResizedByVPA&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f456450&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An event with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;reason=InPlaceResizedByVPA&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; confirms successful in-place downscaling post-startup.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Best practices for production workloads&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pairing with Horizontal Pod Autoscaler (HPA)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: When using HPA alongside CPU startup boost, ensure a robust &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;readinessProbe&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; is defined and keep &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;durationSeconds&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; short (e.g., &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;0s&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;–&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;10s&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). This prevents HPA from falsely interpreting initialization CPU spikes as high steady-state load.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Handle traffic spikes with &lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/configure-capacity-buffer"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;GKE capacity buffers API&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: By reducing the startup tax at scale, you can achieve higher workload density on fewer nodes to improve overall utilization. Consider adopting the GKE capacity buffers API to absorb sudden traffic surges with minimal operational overhead while maintaining strict SLOs. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GKE Autopilot resource ratios&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: On GKE Autopilot, remember that pods must maintain valid CPU-to-memory ratios. Ensure baseline memory allocations accommodate the boosted CPU ratio during startup.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started today&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;CPU startup boost gives GKE users the best of both worlds: lightning-fast cold starts for CPU-intensive workloads like Java, Node.js, and Python, paired with maximum resource efficiency and lower cloud costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to accelerate your GKE workloads?&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explore the &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine/docs/how-to/cpu-startup-boost"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE CPU startup boost documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Learn more about &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine/docs/how-to/vertical-pod-autoscaling"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Vertical Pod Autoscaling on GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Try CPU startup boost on your GKE Standard or Autopilot clusters running GKE &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;1.36.0-gke.4447000&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or later!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 02 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/containers-kubernetes/gke-cpu-startup-boost-faster-pod-starts-lower-costs/</guid><category>Compute</category><category>GKE</category><category>Developers &amp; Practitioners</category><category>Containers &amp; Kubernetes</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>GKE CPU startup boost: Accelerate app starts without over-provisioning</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/containers-kubernetes/gke-cpu-startup-boost-faster-pod-starts-lower-costs/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Abdel Sghiouar</name><title>Cloud Developer Advocate</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jakub Pawełczak</name><title>Cloud Product Manager</title><department></department><company></company></author></item><item><title>What’s new with Google Cloud</title><link>https://cloud.google.com/blog/topics/inside-google-cloud/whats-new-google-cloud/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="kgod7"&gt;Want to know the latest from Google Cloud? Find it here in one handy location. Check back regularly for our newest updates, announcements, resources, events, learning opportunities, and more. &lt;/p&gt;&lt;hr/&gt;&lt;p data-block-key="ru1z9"&gt;&lt;b&gt;Tip&lt;/b&gt;: Not sure where to find what you’re looking for on the Google Cloud blog? Start here: &lt;a href="https://cloud.google.com/blog/topics/inside-google-cloud/complete-list-google-cloud-blog-links-2021"&gt;Google Cloud blog 101: Full list of topics, links, and resources&lt;/a&gt;.&lt;/p&gt;&lt;hr/&gt;&lt;p data-block-key="b0lnw"&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: []&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;Sept 28 - Oct 2&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Claude Sonnet 5.5 is now available on Google Cloud. &lt;/strong&gt;Built for focused coding and knowledge work, it delivers stronger performance for feature development, and creating presentation-ready documents, while offering more intelligence with  a lower cost per task for most work at faster speed.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Deploy Sonnet 5.5 directly from Model Garden with Google Cloud’s enterprise security, scale, and governance built in. &lt;a href="https://cloud.google.com/products/model-garden/claude"&gt;Try it now&lt;/a&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mastering Storage Management with Storage Intelligence&lt;br/&gt;&lt;/strong&gt;In our new blog, explore practical strategies to modernize storage operations using Storage Intelligence Advisor and enhanced Batch Operations. See how customers like Shipt and Palo Alto Networks are using these features to manage their estates at scale. Today, Storage Intelligence is used by 25 of the top 50 GCS customers. 2 in 3 customers have Storage Intelligence enabled for over 90% of their storage footprint.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="18" href="https://cloud.google.com/blog/products/storage-data-transfer/storage-intelligence-advisor-and-batch-operations-updates" rel="noopener" target="_blank"&gt;Explore best practices on our blog&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cut Gen AI costs by 70% with Apigee X dynamic routing&lt;br/&gt;&lt;/strong&gt;Generative AI deployments face a harsh cost-performance trade-off: static model selection either wastes budget or compromises quality. Discover how to build an Intelligent AI Gateway using Apigee X to dynamically evaluate prompt complexity in real-time. This new blueprint automatically routes simple queries to lightweight models (like Gemini 3.5 Flash Lite) and complex reasoning tasks to advanced models (like Gemini 3.7 Flash), achieving over 70% cost savings without sacrificing output quality. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="22" href="https://goo.gle/4ycN5Pf" rel="noreferrer noopener" target="_blank"&gt;Read the guide to optimize your Gen AI architecture&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Agent Clinic launches new hands-on session on production agent evaluations&lt;br/&gt;&lt;/strong&gt;A new episode of the AI Agent Clinic is now available showing organizations how to move past manual "vibe checks" to build robust, production-grade agent evaluations. In this hands-on session, Google Cloud engineer Dani Zamora and Matthew Feroz (Merge) take DocsHound, an open-source LangGraph agent, and build an end-to-end eval pipeline in under an hour. Learn how to standardize multi-turn traces across any framework using OpenTelemetry and OpenInference, pair LLM judges with deterministic checkers, and catch silent quality regressions before shipping. &lt;br/&gt;&lt;br/&gt;&lt;a href="https://youtu.be/wPdoZRbvaF4?si=9JrwgwVLzOepNmqq&amp;amp;t=1" rel="noopener" target="_blank"&gt;Watch here&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Drive AI impact: Join Google's Agentic Data Cloud event on Nov 4&lt;br/&gt;&lt;/strong&gt;Every enterprise plans to adopt agentic AI within two years, but data bottlenecks hold them back—AI accesses just 45% of enterprise data today (&lt;em&gt;MIT, 2026&lt;/em&gt;). Google's Agentic Data Cloud provides the trustworthy foundation and real-time context AI needs. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="28" href="https://cloudonair.withgoogle.com/events/trusted-agents-real-outcomes-agentic-data-cloud" rel="noreferrer noopener" target="_blank"&gt;Join our product leaders November 4th&lt;/a&gt; to explore the latest innovations across databases, analytics, business intelligence, and storage, and get direct answers in our executive Q&amp;amp;A with Google Data Cloud VP &amp;amp; GM Andi Gutmans. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="29" href="https://cloudonair.withgoogle.com/events/trusted-agents-real-outcomes-agentic-data-cloud" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register today!&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Sept 21 - Sept 25&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Master MCP tool authorization and agent governance with Apigee&lt;br/&gt;&lt;/strong&gt;While the Model Context Protocol (MCP) solves interoperability for autonomous AI agents, chained actions like CRM edits or database queries quickly expose systems to unauthorized execution. Join our technical deep dive on Thursday, October 1, 2026, at 5:00 PM CEST featuring Christophe from Google Cloud. Learn how positioning Apigee between MCP clients and enterprise backends enables fine-grained authorization (FGA), complete audit trails, and policy evaluation via emerging standards like OpenID AuthZEN.&lt;br/&gt;&lt;br/&gt;Language and accessibility note: This session will be hosted in French, but non-French speakers can follow along seamlessly by turning on Google Meet live translated captions to read in English, Spanish, German, Portuguese, or Italian.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="46" href="https://rsvp.withgoogle.com/events/apigee-emea-office-hours-2024/sessions#:~:text=Gouvernance%20des%20Agents%20%3A%20Ma%C3%AEtriser%20l%27autorisation%20des%20tools%20MCP%20avec%20Google%20Apigee" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the October 1 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Apigee Trace Viewer Tutorial: Capturing &amp;amp; Analyzing Proxy Traces&lt;br/&gt;&lt;/strong&gt;Streamlining API proxy debugging just got easier with a new tutorial by Apigee Customer Engineer Tyler Ayers. The guide covers end-to-end instructions for capturing debug traces in both Google Cloud Apigee X (or Hybrid) and the local Apigee Emulator, extracting trace JSON data via the web UI or automated REST APIs, and analyzing execution flows, variable mutations, and latency bottlenecks using the open source Apigee Trace Viewer. &lt;br/&gt;&lt;br/&gt;&lt;strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="49" href="https://goo.gle/4746upt" rel="noreferrer noopener" target="_blank"&gt;Read the Apigee Trace Viewer guide today.&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automate Apigee proxy testing locally&lt;br/&gt;&lt;/strong&gt;Catching errors early saves time and money. A new tutorial by Apigee customer engineer Tyler Ayers shows how to use the Apigee Local Emulator for automated testing. Learn to run tests locally, integrate them into CI/CD pipelines, and deploy on Google Cloud Run for shared sandboxes. This approach provides instant feedback and zero cloud costs, helping teams speed up deployment cycles.&lt;br/&gt;&lt;br/&gt;&lt;strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="53" href="https://goo.gle/4hAQYHL" rel="noreferrer noopener" target="_blank"&gt;Read the tutorial&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scale your enterprise multi-agent systems with Apigee&lt;br/&gt;&lt;/strong&gt;Deploying multi-agent architectures in production introduces critical hurdles around security, operational control, and runtime expenses. Discover how Apigee API Hub provides a central discovery surface to eliminate agent sprawl across tools, Model Context Protocol (MCP) servers, and enterprise APIs. Learn how to turn existing backend services into secure MCP tools using Agent Gateway guardrails, while applying semantic caching and intelligent model routing to keep compounding token costs predictable.&lt;br/&gt;&lt;br/&gt;Join Google Cloud &lt;strong style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;in Chicago in Oct.15 for The AI Evolution. &lt;/strong&gt;&lt;strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="60" href="https://goo.gle/45e67I0" rel="noreferrer noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;Reserve your seat for Chicago&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Automate Apigee proxy testing with the Apigee Local Emulator&lt;br/&gt;&lt;/strong&gt;Waiting on remote deployments to validate API proxy logic slows down release cycles and increases infrastructure overhead. Join Nigel Walters on Thursday, October 8, 2026, at 5:00 PM CEST for a Community TechTalk on shift-left testing for Apigee. Discover how to use the Apigee Local Emulator and apigee-emulator-service to run sub-second assertion suites on local machines, automate CI/CD checks in GitHub Actions, and deploy ephemeral preview sandboxes on Google Cloud Run.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="64" href="https://goo.gle/acttsessions" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the October 8 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Now in Public Preview: AI-assisted EKS-to-GKE migrations with deterministic guardrails&lt;br/&gt;&lt;/strong&gt;Migrating complex Kubernetes estates from AWS EKS to GKE is traditionally high-friction and error-prone. Now in Public Preview, &lt;strong&gt;GKE Agentic Migration &lt;/strong&gt;is an open-source agent plugin that replaces ad-hoc LLM prompting with an AI-assisted migration workflow protected by deterministic guardrails.&lt;br/&gt;&lt;br/&gt;Running locally in your development harness, it indexes source IaC, maps cloud-specific primitives (such as Karpenter to Custom Compute Classes), and validates configurations offline—delivering reviewable pull requests and data-migration runbooks with zero live cluster mutations.&lt;br/&gt;&lt;br/&gt;Learn more in the &lt;strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="70" href="https://cloud.google.com/blog/products/containers-kubernetes/gke-agentic-migration?e=48754805" rel="noreferrer noopener" target="_blank"&gt;announcement blog&lt;/a&gt;&lt;/strong&gt; and &lt;strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="71" href="https://github.com/gke-labs/gke-agentic-migration." rel="noreferrer noopener" target="_blank"&gt;try the plugin on GitHub&lt;/a&gt;&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Opus 5.5 is now available on Google Cloud.&lt;/strong&gt; Built for everyday complex tasks, it delivers stronger agentic coding, research, and analysis while handling long-running work at a lower cost per token. Google Cloud continues to provide enterprise customers with broad model choice to build, deploy, and scale their AI agents securely. &lt;br/&gt;&lt;br/&gt;&lt;strong&gt;Try it &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="74" href="https://console.cloud.google.com/agent-platform/publishers/anthropic/model-garden/claude-opus-5-5" rel="noreferrer noopener" target="_blank"&gt;here&lt;/a&gt;.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Import Delta Lake tables with Dataflow Job Builder!&lt;br/&gt;&lt;/strong&gt;Migrating to borderless Lakehouse just got a lot easier. You can now import Delta Lake tables stored in Cloud Storage using &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="85" href="https://docs.cloud.google.com/dataflow/docs/guides/job-builder" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Dataflow Job Builder&lt;/strong&gt;&lt;/a&gt;, a no-code/low-code interface for authoring Dataflow pipelines. Because Dataflow is a fully managed service, you are spared the overhead of provisioning and managing virtual machines. For step-by-step guidance, check out the documentation &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="86" href="https://docs.cloud.google.com/dataflow/docs/guides/delta-lake-df-lakehouse-integration" rel="noreferrer noopener" target="_blank"&gt;here&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Sept 14 - Sept 18&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Storage Intelligence Advisor for Google Cloud Storage is now GA&lt;br/&gt;&lt;/strong&gt;Google Cloud Storage customers can now manage cloud storage more effectively with &lt;strong&gt;Storage Intelligence Advisor&lt;/strong&gt;, delivering curated metrics, automated anomaly detection, and actionable recommendations right out of the box, with zero setup required.&lt;br/&gt;&lt;br/&gt;Advisor baselines activity across your projects and automatically detects four key anomalies: surges in operations, unexpected rises in cross-region egress, and spikes in errors. Each finding includes deep drill-down visibility into the resources driving the change, alongside prescriptive steps to remediate issues before they impact performance or cost.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="136" href="https://docs.cloud.google.com/storage/docs/storage-intelligence/advisor-overview" rel="noopener" target="_blank"&gt;Learn more to get started with Storage Intelligence Advisor&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build private WebSockets from Apigee X to Cloud Run&lt;br/&gt;&lt;/strong&gt;Real-time AI agents and streaming architectures often require persistent, bidirectional connections. A new implementation guide by Apigee Customer Engineer Joel Gauci demonstrates how to establish private southbound connectivity between Apigee X and Cloud Run. Using Private Service Connect (PSC) and a Regional Internal Application Load Balancer, teams can enforce API governance and security policies at the edge while keeping backend services completely isolated from the public internet.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="132" href="https://goo.gle/4h4ABlh" rel="noreferrer noopener" target="_blank"&gt;Explore the step-by-step guide and open-source code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Connecting Gemini Enterprise Agent Runtime to Apigee with Private Service Connect&lt;/strong&gt; &lt;br/&gt;Deploying autonomous AI agents often presents security, compliance, and cost challenges. A new reference guide details how to build an end-to-end, private architecture between Gemini Enterprise Agent Runtime and Apigee. This design helps protect internal backends and manage token quotas. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="128" href="https://goo.gle/4h4PAMd" rel="noreferrer noopener" target="_blank"&gt;Read the full community guide and deploy the code&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Discover what’s new and next in Apigee&lt;br/&gt;&lt;/strong&gt;As enterprise architectures adapt to generative AI and autonomous workflows, Apigee is expanding its proven platform capabilities to support modern AI gateway use cases alongside traditional API management. Join our session on Thursday, September 24, featuring Apigee Product Manager Geir Sjurseth. Get an inside look at recent product releases, explore architectural patterns for securing models and agents, and bring your questions for the live Q&amp;amp;A.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="124" href="https://goo.gle/4y4j44A" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the September 24 Apigee product update&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Managed Service for Apache Kafka supports clusters with public Internet access!&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/networking-kafka#connect-clients-to-a-public-cluster"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Kafka public clusters&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, you can now produce and consume messages from clients outside your VPC—including your local machine, for faster, frictionless testing. Public clusters unlock use cases like IoT devices, retail storefronts, and telco network towers. Enable public access on new or existing clusters via the Google Cloud console, gcloud CLI, or REST API. &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/create-cluster"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spin up your first public cluster&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or reach out to kafka-hotline@google.com with questions.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Stream data directly into Bigtable using Bigtable subscriptions, now in Preview!&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;You can write Pub/Sub messages to a Bigtable table with zero ETL with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/pubsub/docs/bigtable-subscriptions"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Bigtable subscriptions&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. No pipelines, no code, delivered by the serverless, zero-ops experience you already know with Pub/Sub. Power your AI workloads, from model telemetry to real-time context engineering, without the overhead of managing complicated ETL pipelines. Built to be dependable, with native support for dead-letter topics. &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/pubsub/docs/bigtable-subscriptions"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Try the feature today&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;!&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Sept 7 - Sept 10&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Why Your Voice Agent Needs Session Auditing&lt;br/&gt;&lt;/strong&gt;Moving voice agents to production demands robust quality monitoring. This guide dives deep into the inner workings of the Agent Development Kit (ADK) responsible for audio session auditing. Learn how the ADK's &lt;code&gt;save_live_blob&lt;/code&gt; feature intercepts, buffers, and stores raw audio chunks during active Gemini Live sessions. We explore building an automated post-processing pipeline to seamlessly stitch these fragments into cohesive, playable audio files. Discover how to leverage these vital audio audit trails to monitor real-world interactions, diagnose failures, and ensure enterprise-grade reliability. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="107" href="https://discuss.google.dev/t/why-your-voice-agent-needs-session-auditing-and-how-to-build-it/390882" rel="noreferrer noopener" target="_blank"&gt;Read the full guide here&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AlloyDB Omni Red Hat RPM Orchestrator now Generally Available&lt;br/&gt;&lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="139" href="https://docs.cloud.google.com/alloydb/omni/docs/redhat-orchestrator-overview" rel="noreferrer noopener" target="_blank"&gt;AlloyDB Omni Red Hat RPM orchestrator&lt;/a&gt; is now Generally Available. The AlloyDB Omni Red Hat RPM orchestrator offers a new way to manage PostgreSQL-compatible workloads on bare metal or VM platforms, combining the high performance of AlloyDB, access to generative AI features and Gemini models to build AI agents and applications, and full automation. The orchestrator simplifies cluster provisioning and lifecycle management by allowing you to define reference architecture specifications, customizable by adjusting instance parameters, node configurations, and networking options — discover all details in &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="140" href="https://cloud.google.com/blog/products/databases/alloydb-omni-rpm-orchestrator-is-generally-available" rel="noreferrer noopener" target="_blank"&gt;full blog post&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Aug 31 - Sept 4&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Automate VM guest software lifecycle with VM Extension Manager, now GA&lt;br/&gt;&lt;/strong&gt;Google Cloud VM Extension Manager is now generally available, eliminating the need for custom startup scripts to manage guest OS extensions across Compute Engine fleets. Define declarative, project-wide policies that enforce desired software states across all regions and zones. Benefit from continuous drift detection with automatic self-healing, multi-zone phased rollouts with automated rollbacks on failure, and centralized fleet health visibility integrated with Cloud Monitoring.&lt;br/&gt;&lt;br/&gt;Explore &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="18" href="https://docs.cloud.google.com/compute/docs/vm-extensions/about-global-policies" rel="noreferrer noopener" target="_blank"&gt;VM Extension Manager documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Assess Apigee migrations without a target environment&lt;br/&gt;&lt;/strong&gt;Planning a migration to Apigee X or Hybrid? You can now assess your legacy Apigee Edge SaaS or OPDK environment earlier in your planning cycle. Using the updated --skip-target-validation flag in the Apigee Migration Assessment Tool, teams can generate a full inventory and establish scope baselines before target infrastructure or IAM credentials are provisioned.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="24" href="https://goo.gle/4iKScRI" rel="noreferrer noopener" target="_blank"&gt;Read the guide to learn more.&lt;/a&gt;&lt;br/&gt;&lt;br/&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Claude Fable 5.1 is now available on Agent Platform&lt;/strong&gt;. It brings performance improvements over Fable 5 across reasoning, full-lifecycle coding, multi-tool workflows, and knowledge work.&lt;/p&gt;
&lt;p&gt;Anthropic also announced Enterprise Frontier Safeguards, a solution that gives customers the option to safely deploy Anthropic’s most capable models while storing their data in cloud infrastructure they control.&lt;/p&gt;
&lt;p&gt;We continue to offer enterprise customers options across frontier models to build, deploy, and scale securely on Google Cloud.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Aug 24 - Aug 28&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Grok 4.6 is now available in Preview on Gemini Enterprise.&lt;/strong&gt; xAI's most capable model, built for coding, agentic tasks, and knowledge work, Grok 4.6 joins Grok 4.3 and Grok 4.20 in Model Garden and becomes the flagship of the Grok family. It supports reasoning, function calling, and structured output for multi-step agentic workflows, and accepts text and image input.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="58" href="https://console.cloud.google.com/agent-platform/publishers/xai/model-garden/grok-4.6" rel="noreferrer noopener" target="_blank"&gt;Get started today&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Empowering autonomous agents with advanced security governance&lt;/strong&gt;&lt;br/&gt;AI agents offer incredible productivity gains, but granting them access to read emails, query databases, and trigger APIs introduces critical new security risks. In fact, 79% of tech leaders cite security and governance as their biggest challenge to scaling AI. Traditional tools are no longer enough to handle automated threats like prompt injection and dynamic permissions. Discover how forward-thinking enterprises are using secure-by-default design, agent identity governance, and human-in-the-loop controls to deploy agents with confidence.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="61" href="https://cloud.google.com/blog/topics/ai-infrastructure/state-of-ai-infrastructure-report-agent-governance-and-security?e=48754805" rel="noreferrer noopener" target="_blank"&gt;Read more&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stateful processing is available in BigQuery continuous queries in Preview&lt;br/&gt;&lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="67" href="https://docs.cloud.google.com/bigquery/docs/continuous-queries-introduction#supported_stateful_operations" rel="noreferrer noopener" target="_blank"&gt;Stateful operations&lt;/a&gt; significantly expand what’s possible with BigQuery continuous queries. This feature allows users to leverage functions like JOINs, aggregations, and windowing functions directly in their streaming queries. Now you can calculate metrics over time (for example, a 30-minute average) to power your downstream applications and AI agents with much richer, real-time signals.&lt;/li&gt;
&lt;li&gt;Try out our feature &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="68" href="https://docs.cloud.google.com/bigquery/docs/continuous-query-joins" rel="noreferrer noopener" target="_blank"&gt;here&lt;/a&gt; and share your feedback with bq-continuous-queries-feedback@google.com!&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Synthetic data generator tool is available for Managed Service for Kafka&lt;br/&gt;&lt;/strong&gt;You’ve launched your first Kafka cluster. Now what? The next thing to do is to produce some data to the cluster, but that involves modifying a client application somewhere or spinning up a virtual machine. The synthetic data generator tool, now generally available, can start sending mock data to your cluster in 3 clicks, and will get data streaming into your cluster in less than two minutes. The perfect utility for those moments you just want to test your cluster and new features. Try &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="71" href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/quickstart-synthetic-data" rel="noreferrer noopener" target="_blank"&gt;our quickstart&lt;/a&gt; today!&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dataflow pipeline updates are faster &amp;amp; more flexible&lt;br/&gt;&lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="76" href="https://docs.cloud.google.com/dataflow/docs/guides/upgrade-guide" rel="noreferrer noopener" target="_blank"&gt;Dataflow pipeline updates&lt;/a&gt;&lt;strong&gt; &lt;/strong&gt;can now stop-and-replace pipelines, a major addition to the existing in-place-update feature. The new parallel pipeline option accelerates the migration between the old &amp;amp; new pipeline, resulting in reduced disruption to your business. You can also set a timeout on drains that prevents runaway costs for your pipeliness in the event of stuck processing. This feature is generally available. Try it &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="77" href="https://docs.cloud.google.com/dataflow/docs/guides/updating-a-pipeline" rel="noreferrer noopener" target="_blank"&gt;here&lt;/a&gt;!&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Aug 17 - Aug 21&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Webinar: Agent Identity as the backbone for secure AI innovation&lt;/strong&gt;&lt;br/&gt;An AI agent with a stolen API key looks identical to a legitimate one. As autonomous agents scale across enterprise systems, static credentials and legacy IAM policies can no longer keep up with machine-speed execution. Join Shaun Liu, Product Manager at Google Cloud, on August 27 at 1 PM ET to explore Google Cloud’s vision for unifying agent, human, and nonhuman identity into a workload-centric platform using verifiable cryptographic identities (SPIFFE, ID-JAG, OAuth).&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="24" href="https://www.brighttalk.com/webcast/18282/673389?utm_source=Social" rel="noreferrer noopener" target="_blank"&gt;Register for the webinar now&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Aug 10 - Aug 14&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Diagnosing Apigee Hybrid Cassandra Read Latency for Peak Performance&lt;br/&gt;&lt;/strong&gt;Diagnose real-time Cassandra read latency and resolve API key verification bottlenecks in Apigee Hybrid with this step-by-step troubleshooting guide. Learn how to deploy a debugging client and query performance tables to maintain sub-millisecond response times. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="16" href="https://goo.gle/4bXcW4w" rel="noreferrer noopener" target="_blank"&gt;&lt;em&gt;Read the Apigee Hybrid Cassandra Troubleshooting Guide&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Keep moving with agents! The All Things Agentic Hackathon is officially live.&lt;br/&gt;&lt;/strong&gt;We're challenging builders to build next-generation agents that take on the busy work and handle the heavy lifting in the background using Gemini 3.5 and Google Cloud. Compete for your share of $190,000 in prizes, cash, and Google Cloud credits! Submissions are open from August 3, 2026, to August 31, 2026.&lt;br/&gt;&lt;br/&gt;&lt;a href="allthingsagentichackathon.devpost.com" rel="noopener" target="_blank"&gt;Learn more and register&lt;/a&gt;. &lt;a href="g.dev/cloud/all-things-agentic" rel="noopener" target="_blank"&gt;Sign up&lt;/a&gt; for GEAR to get exclusive updates and your badge. #AllThingsAgenticHackathon&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accelerate PostgreSQL migrations using Gemini in Database Migration Service&lt;br/&gt;&lt;/strong&gt;Enterprise database migrations often stall during the "last mile" of translating legacy stored procedures, triggers, and custom functions from Oracle or SQL Server. Database Migration Service (DMS) now provides AI-assisted code conversion powered by Gemini in Databases. By combining deterministic compiler rules for 1:1 syntax with Gemini contextual synthesis for complex procedural blocks, DMS converts legacy code into native PostgreSQL and AlloyDB with full schema awareness and side-by-side validation.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="21" href="https://cloud.google.com/blog/products/databases/accelerate-postgresql-migrations-with-gemini-in-dms" rel="noreferrer noopener" target="_blank"&gt;Read the full blog post&lt;/a&gt; to learn how to streamline your database code conversion.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compute Flex CUDs now available for G2 and G4 GPU VMs&lt;br/&gt;&lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="28" href="https://cloud.google.com/compute/docs/instances/committed-use-discounts-overview#spend_based" rel="noreferrer noopener" target="_blank"&gt;Compute Flexible Committed Use Discounts (Flex CUDs)&lt;/a&gt; are now available for &lt;strong&gt;G2 (NVIDIA L4) &lt;/strong&gt;and &lt;strong&gt;G4 (NVIDIA RTX Pro 6000) VMs&lt;/strong&gt;. You can now lock in predictable savings while retaining the flexibility to adapt across VM families, migrate between regions, and combine general-purpose compute, GKE, Cloud Run, and G2 &amp;amp; G4 GPU VMs under a single spend commitment. Flex CUDs for G-series VMs let you lock in savings today while preserving the agility to upgrade to latest hardware without disruption!&lt;br/&gt;&lt;br/&gt;Explore&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="29" href="https://cloud.google.com/compute/vm-instance-pricing" rel="noreferrer noopener" target="_blank"&gt; VM instance pricing&lt;/a&gt; or learn more about &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="30" href="https://cloud.google.com/compute/docs/instances/committed-use-discounts-overview#spend_based" rel="noreferrer noopener" target="_blank"&gt;Flex CUDs&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Rapid Bucket accelerates the training and checkpoint performance in PyTorch Ecosystem via GCSFS&lt;br/&gt;&lt;/strong&gt;With the release of GCSFS &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="37" href="https://github.com/fsspec/gcsfs/releases/tag/2026.8.0" rel="noreferrer noopener" target="_blank"&gt;2026.8.0&lt;/a&gt;, organisations can now unlock maximum ROI from their AI/ML infrastructure by eliminating data starvation on GPUs in PyTorch ecosystem when they are using Frameworks like Dask, Pandas, PyTorch , PyTorch Lightning, Hugging Face Datasets, Ray dataetc. By making adaptive concurrent prefetching the default, GCSFS dynamically predicts and background-fetches sequential read patterns—boosting single-file throughput by 5x, and scaling up to 21 GiB/s , saturating the NIC when paired with &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="38" href="https://docs.cloud.google.com/storage/docs/rapid/rapid-bucket" rel="noreferrer noopener" target="_blank"&gt;Rapid Bucket&lt;/a&gt;. Saturating the NIC translates to significantly improved &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="39" href="https://cloud.google.com/blog/products/ai-machine-learning/goodput-metric-as-measure-of-ml-productivity" rel="noreferrer noopener" target="_blank"&gt;accelerator goodput&lt;/a&gt; and reduced training wait times with zero integration friction. Training and checkpoint restore workflows benefit from intelligent memory management that automatically drains the buffer during random reads to completely avoid bandwidth or memory penalties.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Aug 3 - Aug 7&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Navigate data sovereignty and AI innovation with hybrid cloud&lt;/strong&gt;&lt;br/&gt;For enterprises facing strict compliance rules, keeping sensitive data on-premises often means missing out on cutting-edge AI. Data from the 2026 State of AI Infrastructure report reveals that 52% of IT leaders are adopting hybrid cloud strategies to bridge this gap. Our latest blog post explores how Google Distributed Cloud (GDC) helps organizations deploy connected or air-gapped models to run advanced AI entirely within secure environments—mitigating geopolitical risks without sacrificing innovation. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="106" href="https://cloud.google.com/blog/topics/hybrid-cloud/state-of-ai-infrastructure-report-on-hybrid-cloud-and-gdc" rel="noreferrer noopener" target="_blank"&gt;Read more&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;SAP and Google Cloud Launch BDC Connect for BigQuery&lt;br/&gt;&lt;/strong&gt;For years, enterprises have struggled with the cost, risk, and complexity of moving mission-critical SAP data into advanced analytics platforms. The general availability of SAP Business Data Cloud (BDC) Connect for BigQuery marks a turning point. By introducing revolutionary zero-copy, bi-directional data sharing, this new capability seamlessly bridges SAP systems with Google Cloud's powerful data and AI ecosystem. Instead of wrestling with manual data duplication and lost business context, organizations can now eliminate silos, dramatically lower their analytics costs, and rapidly deploy trustworthy, agentic AI solutions grounded in real-time operational reality. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="110" href="https://cloud.google.com/blog/products/sap-google-cloud/sap-and-google-cloud-launch-bdc-connect-for-bigquery?e=48754805" rel="noreferrer noopener" target="_blank"&gt;Read the full announcement to learn how to transform your data strategy&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Google Cloud Cortex Framework version 7 is now generally available!&lt;br/&gt;&lt;/strong&gt;This release helps you modernize your data architecture for AI agent readiness, enabling you to quickly deploy, customize, and extend robust data products while simplifying orchestration and reducing infrastructure overhead. It provides &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="130" href="https://docs.cloud.google.com/cortex/docs/data-product#available_data_products" rel="noreferrer noopener" target="_blank"&gt;data product accelerators&lt;/a&gt; for SAP-sourced data to build trusted, high-quality &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="131" href="https://docs.cloud.google.com/cortex/docs/data-product" rel="noreferrer noopener" target="_blank"&gt;data products&lt;/a&gt; ready for advanced analytics and agentic use cases. The Framework integrates with Google Cloud products including &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="132" href="https://docs.cloud.google.com/bigquery/docs" rel="noreferrer noopener" target="_blank"&gt;BigQuery&lt;/a&gt;, &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="133" href="https://docs.cloud.google.com/dataform/docs" rel="noreferrer noopener" target="_blank"&gt;Dataform&lt;/a&gt;, &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="134" href="https://docs.cloud.google.com/dataplex/docs" rel="noreferrer noopener" target="_blank"&gt;Knowledge Catalog&lt;/a&gt;, and &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="135" href="https://cloud.google.com/products/gemini-enterprise-agent-platform" rel="noreferrer noopener" target="_blank"&gt;Gemini Enterprise Agent Platform&lt;/a&gt;. Learn more in our &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="136" href="https://cloud.google.com/blog/products/sap-google-cloud/cortex-framework-v7-power-ai-agents-with-sap-data-faster?e=48754805" rel="noreferrer noopener" target="_blank"&gt;announcement blog&lt;/a&gt;, &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="137" href="https://docs.cloud.google.com/cortex/docs/overview" rel="noreferrer noopener" target="_blank"&gt;technical documentation&lt;/a&gt;, or try a &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="138" href="https://docs.cloud.google.com/cortex/docs/demo-deployment" rel="noreferrer noopener" target="_blank"&gt;demo deployment&lt;/a&gt; today. &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;From API Management to AI Gateway with Apigee&lt;br/&gt;&lt;/strong&gt;Massive LLM adoption unlocked automation but exposed critical vulnerabilities, from unpredictable token costs to security risks like prompt injection. Without central management, organizations face accelerated technical debt. Learn how to transform Apigee into an enterprise AI Gateway to centralize governance. This architectural roadmap details how to utilize semantic cache to optimize token costs, implement prompt protection policies for security, and productize tools using the emerging MCP standard.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="141" href="https://goo.gle/44PIO7p" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Read the full architectural roadmap on the Apigee Community Hub&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Centrally govern enterprise AI traffic with Apigee AI Gateway&lt;br/&gt;&lt;/strong&gt;Manage, track, and secure model communication across your entire infrastructure from a single pane of glass. In a new video walkthrough, Principal Architect Tyler Ayers demonstrates how Apigee AI Gateway simplifies agentic governance. Learn how to transparently proxy model traffic, log real-time token counts, and apply runtime security quotas without impacting your developer workflow.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="145" href="https://goo.gle/44bBi6q" rel="noreferrer noopener" target="_blank"&gt;Watch the Apigee AI Gateway demo&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Maximize Provisioned Throughput Utilization&lt;br/&gt;&lt;/strong&gt;Sudden traffic micro-spikes can exceed per-second quotas, triggering 429 errors or forcing overflow into shared resource pools. A new architectural guide demonstrates how to build a serverless "shock absorber" using Cloud Run and Google Cloud Tasks. By decoupling request ingestion from execution, this queue-based pattern flattens volatile traffic bursts and smoothly drips requests to Gemini at your exact quota rate, maximizing Provisioned Throughput utilization while eliminating job failures during peak usage. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="149" href="https://medium.com/google-cloud/smoothing-spiky-llm-traffic-maximize-provisioned-throughput-utilization-with-a-queuing-176753d96818" rel="noreferrer noopener" target="_blank"&gt;Read the step-by-step setup guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Eliminate security blindspots in agentic tool interactions&lt;br/&gt;&lt;/strong&gt;Unmonitored agentic tool calls via the Model Context Protocol (MCP) can introduce critical security risks to your enterprise architecture. Join our technical deep dive on Thursday, August 13, to discover how to position Apigee as a centralized security gateway. Featuring the new ParsePayload policy and payload operations groups in API Products, this session demonstrates how to enforce granular tool filtering, manage execution quotas, and scale secure agent ecosystems without impeding developer velocity. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="152" href="https://goo.gle/4y4j44A" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the August 13 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Jul 27 - Jul 31&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Data Cloud and Apigee CDMX: The AI Agent Evolution | August 12, 2026&lt;br/&gt;&lt;/strong&gt;Enterprise AI demands evolution beyond basic conversational assistants. To generate real value, AI models must connect with the organization's core systems and live data sources. Join us this August 12 at &lt;strong&gt;Google CDMX &lt;/strong&gt;for the exclusive event &lt;strong&gt;AI Evolution: Powering Tomorrow's Enterprise&lt;/strong&gt;. Learn how to design an agile and secure ecosystem by unifying the power of Gemini, Apigee, and data agent technologies through practical demonstrations led by Google Cloud engineers.&lt;br/&gt;&lt;br/&gt;Secure your spot for the in-person session in Mexico City &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="34" href="https://goo.gle/3TyS9hg" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register now!&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="48" href="https://vastedge.com/" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Vast Edge&lt;/strong&gt;&lt;/a&gt;, built on GCP, launches the first live recovery interface for cloud backups, enabling IT teams to inspect backup contents in real time. This transforms backups from a blind, log-based process into an interactive platform where teams can &lt;strong&gt;instantly search, preview, and validate the exact data available for restore&lt;/strong&gt;.&lt;br/&gt;&lt;br/&gt;This platform protects Google Workspace, NetSuite, Salesforce, Workday and many SaaS environments, providing complete visibility and enterprise-grade oversight.&lt;br/&gt;&lt;br/&gt;Visit&lt;strong&gt; &lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="49" href="https://vastedge.com/backup-and-disaster-recovery" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Vast Edge Backup &amp;amp; Disaster Recovery&lt;/strong&gt;&lt;/a&gt; and get a free trial of their backup solutions on the GCP Marketplace for&lt;strong&gt; &lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="50" href="https://console.cloud.google.com/marketplace/product/vastedge-public/google-workspace-backup-restore?hl=en" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Google Workspace Backup&lt;/strong&gt;&lt;/a&gt;,&lt;strong&gt; &lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="51" href="https://console.cloud.google.com/marketplace/product/vastedge-public/netsuite-backup-restore?hl=en" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;NetSuite Backup&lt;/strong&gt;&lt;/a&gt;,&lt;strong&gt; &lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="52" href="https://console.cloud.google.com/marketplace/product/vastedge-public/salesforce-backup-restore-vastedge?hl=en" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Salesforce Backup&lt;/strong&gt;&lt;/a&gt;,&lt;strong&gt; &lt;/strong&gt;and&lt;strong&gt; &lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="53" href="https://console.cloud.google.com/marketplace/product/vastedge-public/workday-backup-restore-vastedge?hl=en" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Workday Backup&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;.&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Jul 20 - Jul 24&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Claude Opus 5, Anthropic’s latest model, is now available on Agent Platform.&lt;/strong&gt; It brings performance improvements over Opus 4.8 across coding, long-running agents, and knowledge work.The model is Zero Data Retention (ZDR) compatible. For safety, high-risk workflows — such as penetration testing or exploit generation — it will notify you and fall back to Opus 4.8.We’re excited to continue to offer enterprise customers options across frontier models to build, deploy, and scale AI securely. Try it &lt;a href="https://console.cloud.google.com/agent-platform/publishers/anthropic/model-garden/claude-opus-5"&gt;here&lt;/a&gt;. &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Apigee Northam Roadshow 2026 | The AI Agent Evolution: Powering Tomorrow's Enterprise&lt;br/&gt;&lt;/strong&gt;AI is evolving. As your organization deploys autonomous agents, the integration between APIs and models becomes critical. Join Google Cloud specialists for an exclusive day of deep-dive sessions and live demos. Discover how the unified power of Apigee and the Google Cloud Agent Platform allows you to build, govern, and scale high-performance AI agents with complete control.  Call to Action: &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="93" href="https://goo.gle/4gOIblK" rel="noreferrer noopener" target="_blank"&gt;Register for Sunnyvale&lt;/a&gt; | &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="94" href="https://goo.gle/3TLCPhi" rel="noreferrer noopener" target="_blank"&gt;Register for NYC&lt;/a&gt; | &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="95" href="https://goo.gle/45e67I0" rel="noreferrer noopener" target="_blank"&gt;Register for Chicago&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deploy an Apigee Proxy for MCP Registry Discovery  &lt;br/&gt;&lt;/strong&gt;Learn how to deploy an Apigee X proxy to format Apigee API Hub data into the Model Context Protocol (MCP) Registry format. This tutorial by Tyler Ayers guides developers through cloning the sample repository, deploying using the Apigee Feature Templater (aft), and testing the endpoint to make API data easily discoverable by coding agents. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="99" href="https://goo.gle/3RTus2N" rel="noreferrer noopener" target="_blank"&gt;Read the full community tutorial to get started.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Simplify AI Infrastructure: Getting Started with Apigee AI Gateway&lt;br/&gt;&lt;/strong&gt;Managing a complex AI landscape with multiple backend environments can present significant operational and governance challenges. A new tutorial walks you through how to build a unified API proxy using Apigee AI Gateway. By establishing a single, secure entry point for all model traffic, teams gain access to real-time analytics, comprehensive tracing, and financial operations auditing—completely seamlessly, and with absolutely no modifications required to client environments or user configurations. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="102" href="https://goo.gle/4wI5Por" rel="noreferrer noopener" target="_blank"&gt;Read the step-by-step setup guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Your AI agents are ready. Is your data?&lt;br/&gt;&lt;/strong&gt;The biggest bottleneck to scaling AI isn't the models—it's giving them access to business context. As enterprises move to proactive systems of action, legacy infrastructure often buckles under the nonlinear speed of AI agents. Google Cloud’s new Agentic Data Cloud, built on AI-native infrastructure, solves this by unifying data, AI models, and operational databases. Discover how a borderless Lakehouse and active Knowledge Catalog can empower your AI agents with trusted, real-time context without unnecessary engineering overhead. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="106" href="https://cloud.google.com/blog/topics/ai-infrastructure/state-of-ai-infrastructure-report-and-the-agentic-data-cloud" rel="noopener" target="_blank"&gt;Read more&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Secure and govern your AI at Apigee AI Horizon in London&lt;br/&gt;&lt;/strong&gt;Moving AI from basic prompts to complex agentic workflows requires trust and control. Join us on Tuesday, 1st September 2026 at Google London for our 5th edition of Apigee AI Horizon. Discover how Google Cloud product leaders and architects are using Apigee and Model Armor to secure LLM APIs, implement policy controls, and manage token consumption. Do not miss this one—register soon!&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="110" href="https://goo.gle/4b8XamT" rel="noreferrer noopener" target="_blank"&gt;Secure your spot for AI Horizon London&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Jul 13 - Jul 17&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Resource-Based CUD Sharing is Now Enabled by Default&lt;/strong&gt;&lt;br/&gt;Starting &lt;strong&gt;June 16, 2026&lt;/strong&gt;, the default setting for Google Cloud &lt;strong&gt;Resource-based Committed Use Discount (CUD)&lt;/strong&gt; sharing will change from disabled to &lt;strong&gt;enabled&lt;/strong&gt; for new billing accounts and eligible existing accounts without active CUDs. This update automatically maximizes your savings by pooling underutilized discounts across your resources.&lt;br/&gt;&lt;br/&gt;You retain full control and can adjust your CUD sharing preferences at any time by changing your CUD scope configuration. For instructions, see &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="49" href="https://docs.cloud.google.com/compute/docs/committed-use-discounts/share-resource-cuds-across-projects#turning-on-committed-use-discount-sharing" rel="noreferrer noopener" target="_blank"&gt;Enable CUD sharing&lt;/a&gt; or &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="50" href="https://docs.cloud.google.com/compute/docs/committed-use-discounts/share-resource-cuds-across-projects#turning-off-committed-use-discount-sharing" rel="noreferrer noopener" target="_blank"&gt;Disable CUD sharing&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Webinar for India: Google Cloud for EdTech: Optimizing Traffic and Token Governance at Scale&lt;br/&gt;&lt;/strong&gt;API traffic surges and AI model integration are reshaping the EdTech landscape. Join Satyam Maloo for the webinar&lt;strong&gt; Google Cloud for EdTech: Optimizing Traffic and Token Governance at Scale &lt;/strong&gt;on July 23, 2026. Learn to implement advanced rate limiting, gain granular token visibility, and leverage real-time analytics to govern your platform effectively. Whether you’re scaling for peak academic seasons or integrating complex AI workflows, this session provides the infrastructure blueprint you need.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="53" href="https://goo.gle/4yqrKm0" rel="noreferrer noopener" target="_blank"&gt;Register Now&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Scaling AI Agents: Treat prompts like software artifacts&lt;br/&gt;&lt;/strong&gt;As AI agents move into production, monolithic system prompts often result in configuration drift, merge conflicts, and silent runtime failures. The solution is adopting a &lt;em&gt;Prompts-as-Code&lt;/em&gt; architecture. By breaking prompts into modular skill files and using a build-time transpiler, engineering teams can introduce dependency resolution, static validation, and CI/CD rigor to their agent's control plane. Stop manually editing massive text files and start building deterministic, reliable agent infrastructure.&lt;br/&gt;&lt;br/&gt;Read more &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="57" href="https://developers.googleblog.com/building-scalable-ai-agents-with-modular-prompt-transpilation/" rel="noreferrer noopener" target="_blank"&gt;here&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;Jul 6 - Jul 10&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Webinar: Introducing Google Cloud NGFW Enterprise advanced malware protection - powered by Palo Alto Networks&lt;br/&gt;&lt;/strong&gt;Discover the new Cloud NGFW advanced malware sandbox, arriving in preview later this year. Powered by Palo Alto Networks Advanced Wildfire, it leverages data from 70,000+ customers to help defeat advanced malware. Join us on July 16 at 11 AM EDT to learn how to build a resilient, zero-trust cloud infrastructure that protects your apps and data, wherever they reside.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="18" href="https://www.brighttalk.com/webcast/18282/668861?utm_source=GCBlog" rel="noreferrer noopener" target="_blank"&gt;Register for the webinar now&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Safely run AI-generated code in Cloud Run sandboxes&lt;br/&gt;&lt;/strong&gt;Cloud Run sandboxes, now in public preview, are lightweight, isolated execution boundaries that you can spawn near-instantly &lt;strong&gt;within your existing Cloud Run service instances&lt;/strong&gt;.&lt;br/&gt;&lt;br/&gt;Whether you need to let an LLM run a dynamically generated Python script to calculate business margins or spin up a headless browser to perform web research, Cloud Run sandboxes give you a secure, isolated sandbox to run these tasks without leaving your serverless environment.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="22" href="https://cloud.google.com/blog/topics/developers-practitioners/google-cloud-run-sandboxes-are-in-public-preview" rel="noreferrer noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;Read the blog&lt;/a&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt; to learn more and get started today.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Australia API Horizon: Scaling Enterprise Governed AI Agents&lt;br/&gt;&lt;/strong&gt;The transition from AI chatbots to autonomous agents is the most critical integration point for your business. Join Google Cloud at our upcoming events to explore exclusive deep-dive sessions on architecting for the agentic era.&lt;br/&gt;&lt;br/&gt;Discover how to use Apigee as an intelligent AI Gateway to govern, secure, and scale high-performance architectures. You will learn to seamlessly build AI tools from your existing APIs and maintain control over your entire ecosystem.&lt;br/&gt;&lt;br/&gt;Join us in your preferred city:
&lt;ul&gt;
&lt;li&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="36" href="https://goo.gle/4voh18S" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Sydney:&lt;/strong&gt; July 28, 2026, at Google Sydney, One Darling Island.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="37" href="https://goo.gle/4h2x0FS" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Canberra:&lt;/strong&gt; July 29, 2026, at Hotel Realm.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="38" href="https://goo.gle/4yisb1F" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Melbourne:&lt;/strong&gt; August 4, 2026, at Google Melbourne.&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build highly available, multi-region services on Cloud Run&lt;br/&gt;&lt;/strong&gt;Maintaining uptime for business-critical applications just got a lot easier on Cloud Run. Service health, now Generally Available, automates cross-region failover by leveraging readiness probes for instance-level health checks with a simple, two-click setup. You can configure service health with global external Application Load Balancers for public-facing applications or cross-region internal Application Load Balancers for private networking traffic.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="42" href="https://cloud.google.com/run/docs/configuring/configure-service-health" rel="noreferrer noopener" target="_blank"&gt;Learn how to configure service health for Cloud Run.&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Report: 83% of organizations need infrastructure upgrades for agentic AI&lt;br/&gt;&lt;/strong&gt;The shift from conversational bots to autonomous agents is breaking legacy systems. Our new &lt;em&gt;State of AI Infrastructure&lt;/em&gt; report details how engineering leaders are adapting to these massive new workloads. To eliminate inference bottlenecks, control hidden scaling costs, and manage agent sprawl, the industry is rapidly moving toward fluid compute, centralized governance, and unified, co-designed architectures.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="46" href="https://cloud.google.com/blog/products/compute/state-of-ai-infrastructure-report-overview?e=48754805" rel="noreferrer noopener" target="_blank"&gt;Explore our key infrastructure insights&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stop tinkering, start scaling: the industrialized AI Playbook&lt;br/&gt;&lt;/strong&gt;Did you know that only 5% of custom AI investments actually return measurable business value? The problem isn’t the technology—it’s how organizations are wired to run it.&lt;br/&gt;&lt;br/&gt;In this compelling read, Google Cloud Consulting breaks down the operational blueprint that bridges the stark gap between "cool tech experiments" and real, P&amp;amp;L-impacting enterprise ROI.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="50" href="https://www.google.com/url?q=https%3A%2F%2Fmedium.com%2F%40kjouannigot_73547%2Fscaling-trusted-ai-google-cloud-insights-to-capture-enterprise-roi-aa6c9b308adb" rel="noreferrer noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;Read the full article on Medium&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI Agent Clinic: Slashing App Latency by 80%&lt;br/&gt;&lt;/strong&gt;Prototyping an AI agent is easy, but scaling for live traffic presents unique challenges. In the latest AI Agent Clinic, our technical experts partner with a developer to optimize PlaybackIQ, a live football analysis agent. This session demonstrates how to use OpenTelemetry to trace bottlenecks in the Gemini Enterprise Agent Platform and deploy to Cloud Run for high-concurrency scaling, achieving an 80% reduction in response time. Learn production-grade debugging strategies to optimize your own LLM applications.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="54" href="https://www.google.com/search?q=https://youtu.be/G7olcqETSn8" rel="noreferrer noopener" target="_blank"&gt;Watch the 60-minute teardown&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jun 29 - Jul 3&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Claude Sonnet 5, Anthropic’s latest model, is now available on Agent Platform&lt;/strong&gt;. &lt;br/&gt;This addition serves as a drop-in replacement for Sonnet 4.6, giving organizations expanded choice for task completion across enterprise workflows. It features enhanced reasoning, cleaner code generation, and computer use capabilities for desktop and browser workflows.&lt;br/&gt;&lt;br/&gt;By continuing to rapidly bring frontier models to our platform, Google Cloud offers an uncompromised choice of the industry's best technology to build, test, and scale enterprise-grade AI.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://console.cloud.google.com/agent-platform/publishers/anthropic/model-garden/claude-sonnet-5?hl=en" rel="noreferrer noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;em&gt;Get started today.&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Automate your AI governance with Apigee and YAML&lt;br/&gt;&lt;/strong&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Manual API gateway configurations can quickly slow down your AI engineering velocity. Join the Apigee community on Thursday, July 16, to discover an automated, declarative blueprint for model garden management. Learn how a simple, repeatable YAML pattern lets your AI practitioners instantly spin up secure, policy-backed enterprise configurations  without friction. Bring your questions and connect during our live Q&amp;amp;A session. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://goo.gle/4y4j44A" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the July 16 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Build next-generation AI portals for autonomous agents&lt;br/&gt;&lt;/strong&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Standard developer portals were designed for human developers to subscribe to static APIs. Today, autonomous agents, LLM toolkits, and dynamic runtimes demand a central nervous system for governance. Join our technical deep dive on Thursday, July 23, to explore Apigee's new AI Portals solution. You will see exactly how to deploy full-service, MCP powered hubs to safely manage enterprise self-service for models, tools, and agents. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://goo.gle/4y4j44A" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the July 23 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Protect your infrastructure from advanced cyberattacks at the API layer (Presented in Portuguese)&lt;br/&gt;&lt;/strong&gt;In an era of increasingly sophisticated threats, relying solely on traditional firewalls leaves critical data gaps. Join our technical community TechTalk on Thursday, July 30—conducted in Portuguese—to learn how to proactively mitigate risks directly at the gateway layer. This session demonstrates how to configure and govern essential Apigee security policies to build a robust line of defense, ensuring maximum availability and complete integrity for your enterprise microservices. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://goo.gle/4y4j44A" rel="noreferrer noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;strong&gt;Register for the July 30 Portuguese Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jun 22 - Jun 26&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Accelerate TPU model loading while saving RAM on GKE.&lt;br/&gt;&lt;/strong&gt;Large model cold starts often stall scaling and leave high-value TPUs idle. The open-source &lt;strong&gt;Run:ai Model Streamer&lt;/strong&gt; now natively supports TPUs with Google Cloud Storage in&lt;strong&gt; &lt;/strong&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://github.com/vllm-project/tpu-inference" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;TPU vLLM 0.18.0&lt;/strong&gt;.&lt;/a&gt; This integration accelerates inference pipelines on GKE by streaming tensors directly into CPU memory, bypassing local disk bottlenecks and the "double-buffering" trap. In benchmarks, loading a 480B parameter model was &lt;strong&gt;over 2x faster&lt;/strong&gt; while cutting peak host memory usage by half. &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://discuss.google.dev/t/accelerate-tpu-model-loading-while-saving-ram-on-gke/374835" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Read the full guide and get started today&lt;/strong&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Stop Training Blind: Scaling AI with the New OpenTelemetry-Based TPU AI Telemetry Collector Agent&lt;br/&gt;&lt;/strong&gt;Google Cloud’s new AI Telemetry Collector agent standardizes TPU monitoring using OpenTelemetry. It optimizes enterprise ML workloads by identifying silent failures and providing zero-cost operational metrics without draining host CPU cycles. The agent seamlessly routes telemetry to Google Cloud Monitoring or Prometheus and custom Grafana setups. Pre-installed on Google-optimized Ubuntu images or available via Docker, it tracks memory, network latency, and core utilization to maximize multi-node training efficiency.&lt;br/&gt;&lt;br/&gt;You can read more of this capability by clicking this &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://discuss.google.dev/t/stop-training-blind-scaling-ai-with-the-new-opentelemetry-based-tpu-ai-telemetry-collector-agent/375210" rel="noreferrer noopener" target="_blank"&gt;link&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jun 15 - Jun 19&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Join us for a deep dive into agentic AI control with AppyThings&lt;br/&gt;&lt;/strong&gt;Your integrations aren’t failing—they are evolving. When users interact with AI agents, they no longer arrive directly at your site, resulting in experiences stripped of your context, expertise, and intended experience. Join us on Thursday, June 25, for a community tech talk in partnership with AppyThings to learn how to solve this new gateway challenge. We will explore how MTN laid an integration foundation with the Model Context Protocol (MCP) to deliver accurate, consistent experiences. Our technical experts will demonstrate how to leverage Apigee as a centralized tools management solution to govern agent access. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://goo.gle/3Sfle0y" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the session&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Optimize Spot VM Deployments with Capacity Advisor for Spot, Now in Public Preview&lt;br/&gt;&lt;/strong&gt;Google Compute Engine has launched &lt;strong&gt;Capacity Advisor for Spot&lt;/strong&gt; to Public Preview, now open to all customers. This tool turns Spot capacity discovery into a data-driven process by providing real-time deployment recommendations to maximize obtainability and minimize preemption risks. Query the &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://docs.cloud.google.com/compute/docs/instances/view-vm-availability" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Capacity Advisor API&lt;/strong&gt;&lt;/a&gt; for obtainability and minimum estimated uptimes, or use the new &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://console.cloud.google.com/compute/capacityAdvisor" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Console UI&lt;/strong&gt;&lt;/a&gt; featuring a global availability map, spot price lookups, and historical preemption rate trends to visually find the most cost-efficient compute capacity.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://docs.cloud.google.com/compute/docs/instances/view-vm-availability" rel="noreferrer noopener" target="_blank"&gt;Get started today&lt;/a&gt; to start optimizing your Spot VM deployments!&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Build a multi-tenant agentic AI system&lt;br/&gt;&lt;/strong&gt;When scaling generative AI across different business units, your teams need specialized AI agents with unique operational rules and tools. Our new reference architecture helps you build a centralized multi-tenant platform to prevent fragmented silos, eliminate data exposure risks, and maintain unified compliance. Read the guide to &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://docs.cloud.google.com/architecture/multi-tenant-agentic-ai-system" rel="noreferrer noopener" target="_blank"&gt;design and deploy a multi-tenant agentic AI system&lt;/a&gt; in Google Cloud.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How to Configure Gemini Enterprise to Connect to a Custom MCP Server&lt;br/&gt;&lt;/strong&gt;The Gemini Enterprise MCP Connector was a big announcement at Google Cloud Next because it introduces the ability to connect Gemini Enterprise to MCP servers. This blog &lt;a href="https://medium.com/google-cloud/how-to-configure-gemini-enterprise-to-connect-to-a-custom-mcp-server-2e28adc96420" rel="noopener" target="_blank"&gt;post&lt;/a&gt; provides a step-by-step guide on how to configure your first Custom MCP Server connector using the Google Maps Ground Lite MCP server as an example. Once you understand this flow, you can configure multiple MCP servers with Gemini Enterprise to bring all the context you need.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jun 8 - Jun 12&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Simplify Multi-Cloud Planning with Cloud Location Finder, now Generally Available&lt;/strong&gt; &lt;br/&gt;Cloud Location Finder provides up-to-date data on public regions, zones, and Google Distributed Cloud Connected locations across Google Cloud, AWS, Azure, and OCI. You can now programmatically discover locations based on provider, proximity, territory, and carbon footprint to optimize your global infrastructure strategy for performance, compliance, and sustainability. &lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" data-airgap-id="14" href="https://cloud.google.com/location-finder/docs" rel="noreferrer noopener" target="_blank"&gt;Get started for free today&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jun 1 - Jun 5&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Modeling the physical world with BigQuery Graph&lt;/strong&gt;&lt;br/&gt;Managing complex supply chains requires more than just spreadsheets; it requires a digital replica of the physical world. In this &lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://cloud.google.com/blog/products/data-analytics/modeling-a-digital-twin-using-bigquery-graph" rel="noreferrer noopener" target="_blank"&gt;post&lt;/a&gt;, Guru Rangavittal and Candice Chen explore how BigQuery Graph enables organizations to build a digital twin by turning physical assets into an interconnected map of nodes and edges. By moving beyond traditional relational databases, businesses gain real-time clarity into operations—from executing surgical ingredient recalls to analyzing weather-driven logistics risks. Discover how BigQuery Graph transforms reactive firefighting into proactive, precision modeling, allowing you to see critical connections in seconds and future-proof your supply chain.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Apigee for AI: Govern LLMs and MCP Servers (Presented in Spanish)&lt;br/&gt;&lt;/strong&gt;Learn how to securely transition your AI initiatives from experimental prototypes to enterprise-ready deployments. Join Luis Cuellar on June 18 for a technical deep dive (presented in Spanish) exploring Apigee’s latest AI gateway capabilities. Discover how to centralize governance over Model Context Protocol (MCP) servers, protect Large Language Models (LLMs) with robust API gateway security policies, and manage token-based quotas.&lt;br/&gt;&lt;br/&gt;&lt;a class="colors-hyperlink-primary underline focus-visible outline-offset-0 rounded" href="https://goo.gle/4dyC2Ie" rel="noreferrer noopener" target="_blank"&gt;&lt;strong&gt;Register for the June 18 Spanish Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;May 25 - May 29&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.anthropic.com/news/claude-opus-4-8" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Anthropic’s Claude Opus 4.8&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is now available on &lt;/span&gt;&lt;a href="https://console.cloud.google.com/vertex-ai/publishers/anthropic/model-garden/claude-opus-4-8"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;. &lt;/strong&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;As we continue to expand our platform's model offerings, this addition gives organizations more options for handling complex, multi-stage enterprise workflows. Claude Opus 4.8 brings strong capabilities in agentic coding, allowing developers to manage extensive refactors and tracking dependencies over extended sessions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API Horizon Munich July 6, 2026: Orchestrating the Next Era of AI and APIs &lt;br/&gt;&lt;/strong&gt;Master the orchestration of next-gen AI and digital ecosystems. Join Google Cloud experts and DACH tech leaders on July 6 for an exclusive look at the Apigee roadmap, Agent Management, and Model Context Protocol (MCP). Gain real-world insights and connect with the regional integration community.&lt;strong&gt;&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/4dTxQmo" rel="noopener" target="_blank"&gt;Register now&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Securing AI Agents: The Extended Agent Gateway Pattern&lt;br/&gt;&lt;/strong&gt;Learn how to prevent autonomous AI agents from invoking unauthorized APIs. Join Apigee Specialist Joel Gauci on June 4 for a technical deep dive into the Extended Agent Gateway pattern. This session covers enforcing Fine-Grained Authorization (FGA), implementing secure token exchange, and establishing Model Context Protocol (MCP) governance at the API gateway layer to protect enterprise backend services.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/4fbAsxg" rel="noopener" target="_blank"&gt;&lt;strong&gt;Register for the June 4 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;API-to-Agent Security: Exposing REST APIs to Gemini Enterprise via MCP&lt;br/&gt;&lt;/strong&gt;Connect Gemini Enterprise agents to core data without creating security hazards. Join Google Cloud Specialist Nigel Walters on June 11 to learn how to instantly transform legacy REST APIs into secure Model Context Protocol (MCP) servers. We’ll cover how to safely register tools with Gemini while enforcing gateway-level guardrails like rate limiting and access control policies.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/4nVyjIr" rel="noopener" target="_blank"&gt;&lt;strong&gt;Register for the June 11 Community TechTalk&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;May 18 - May 22&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Chinese Webinar | June 4: AI Command and Control&lt;br/&gt;&lt;/strong&gt;As AI agents move from experimental pilots to core enterprise functions, governance has become a critical next step. Join Google Cloud on June 4th at 10:00 AM (Beijing Time) to learn how to build a secure AI management layer architecture. We'll explore how to develop governed MCP (Model Context Protocol) endpoints, manage tool access to enterprise data, and leverage robust audit logs to operationalize AI. This session also includes a practical demonstration of these governance frameworks on Google Cloud.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/4dx4Lf5" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;Register here&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GCP Announces New Features to Benchmark and Optimize LLMs for On-Device Use Cases&lt;br/&gt;&lt;/strong&gt;Deploying fine-tuned LLMs from GCP to edge devices like smartphones is complex due to fragmented hardware. Google AI Edge Portal bridges this gap, giving GCP developers the ability to test AI performance on 120+ Android devices, representing the full diversity of high, medium, and low tier smartphones on the market today. This week at I/O, we announced brand new &lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/benchmark-llms-on-device-with-ai-edge-portal" rel="noopener" target="_blank"&gt;capabilities&lt;/a&gt; to benchmark and debug LLM performance across these devices. &lt;a href="https://docs.google.com/forms/d/e/1FAIpQLSfTcGPycQve8TLAsfH46pBlXBZe9FrgJAClwbF7DeL1LgVn4Q/viewform" rel="noopener" target="_blank"&gt;Sign-up&lt;/a&gt; to utilize these new features in private preview today.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;May 11 - May 15&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Build Your AI &amp;amp; MCP Control Tower for Universal Governance&lt;br/&gt;&lt;/strong&gt;Master the future of agentic security with Apigee. Join our Community TechTalk on May 21 to discover how Apigee serves as a central "Control Tower" for the Model Context Protocol (MCP). We will explore how new JSON-RPC tool authorization enables fine-grained access policies across your organization, ensuring secure and scalable AI deployments. Whether managing internal tools or external users, learn to govern your agentic ecosystem with absolute precision. This session is designed for global coverage across EMEA and AMER regions.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/4u9slWF" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;Register for the May 21 Community TechTalk&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Apr 27 - May 1&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Master Your Launch: The Apigee Production Go-Live Checklist&lt;br/&gt;&lt;/strong&gt;Ensure a secure launch with the Apigee production guide. Join Nicola Cardace on May 28 to explore security guardrails, including IAM roles, mTLS configurations, and encrypted KVM migrations. Scheduled at 11 AM EDT / 5 PM CEST to support EMEA and AMER teams, this TechTalk provides the technical roadmap you need to flip the switch with absolute confidence.&lt;br/&gt;&lt;br/&gt;&lt;strong style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;a href="https://goo.gle/4elMCTI" rel="noopener" target="_blank"&gt;Register for the May 28 Community TechTalk&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Transforming APIs into Governed Agentic Tools on the Google Cloud Agentic Platform&lt;br/&gt;&lt;/strong&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Turn your APIs into secure, governed agentic tools on the Google Cloud Agentic Platform. Join Specialist Christophe Lalevée on May 7 for a technical deep dive into AI productization. Scheduled at 5 PM CEST / 11 AM EDT to maximize coverage for developers across EMEA and AMER, this session explores the integration and governance frameworks required to scale enterprise-ready AI with confidence.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://goo.gle/3PfWm7M" rel="noopener" target="_blank"&gt;Register for the May 7 Community TechTalk&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/accelerator-optimized-machines#g4-machine-types" rel="noopener" target="_blank"&gt;Fractional G4 VMs&lt;/a&gt; are Generaly Available, providing a highly efficient and cost-effective entry point for AI and graphics workloads. These new configurations, using NVIDIA virtual GPU (vGPU) technology, allow you to leverage the power of the NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs in flexible, smaller increments, so you can right-size your infrastructure to match the specific demands of your applications. By providing more granular access to advanced hardware, fractional G4 VMs let you optimize resource allocation and reduce overhead without sacrificing performance. You can now select from additional GPU slice sizes for your specific needs:
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;1/2 GPU:&lt;/strong&gt; Ideal for more intensive tasks such as LLM inference, robotics sensor simulation, and high-fidelity 3D rendering.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;1/4 GPU:&lt;/strong&gt; Optimized for mainstream workloads, including mid-range creative design, video transcoding, and real-time data visualization.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;1/8 GPU:&lt;/strong&gt; Great for lightweight applications such as remote desktops, productivity tools, and entry-level streaming services.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Transitioning AI from a sandbox prototype to an enterprise-grade system is a major hurdle. A monolithic script won't suffice for widespread deployment. To achieve true scale and reliability with Gemini, organizations must adopt service-oriented micro-agent architectures, establish Zero-Trust security, and implement rigorous EvalOps. Master the "Agentic Maturity Ladder" to ensure your AI &amp;amp; Agentic solutions are robust, secure, and ready for the real world.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://lnkd.in/gHBH8cTv" rel="noopener" target="_blank"&gt;Watch the deep dive&lt;/a&gt; and &lt;a href="https://discuss.google.dev/t/beyond-the-prototype-scaling-production-grade-agents-with-gemini/356140" rel="noopener" target="_blank"&gt;read the developer blog&lt;/a&gt; to learn more.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ML Development in VS Code with Google Cloud Power: Workbench Extension Now Available&lt;br/&gt;&lt;/strong&gt;Data scientists and developers can now combine the local productivity of VS Code with the scalable infrastructure of Google Cloud. The new Google Cloud Workbench Notebooks extension allows you to connect to and run notebooks on managed cloud environments directly within your local IDE. This integration streamlines the ML lifecycle by eliminating context switching and providing high-performance compute for complex workloads in a familiar interface. As part of our commitment to the developer ecosystem, the extension is fully open-sourced to support community-driven innovation.
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Install from Marketplace:&lt;/strong&gt; &lt;a href="https://marketplace.visualstudio.com/items?itemName=GoogleCloudTools.workbench-notebooks" rel="noopener" target="_blank"&gt;GoogleCloudTools.workbench-notebooks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Contribute on GitHub:&lt;/strong&gt; &lt;a href="https://github.com/GoogleCloudPlatform/colab-enterprise-vscode" rel="noopener" target="_blank"&gt;colab-enterprise-vscode&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Apr 20 - Apr 24&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Announcing the 2026 Google Cloud Partners of the Year&lt;br/&gt;&lt;/strong&gt;Google Cloud is honored to celebrate the winners of the 2026 Partner of the Year awards! These awards recognize an exceptional group of partners across AI, Security, Infrastructure, and more, who have demonstrated a commitment to customer success. From global system integrators to specialized startups, these winners are leveraging the power of Google Cloud to solve complex challenges and drive digital transformation worldwide. Join us in congratulating these organizations for their innovation, collaboration, and impactful results over the past year.&lt;br/&gt;&lt;br/&gt;See the &lt;a href="https://cloud.google.com/blog/topics/partners/2026-partners-of-the-year-winners-next26"&gt;2026 Partner Award winners&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Apr 13 - Apr 17&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;We're excited to announce the &lt;strong&gt;Public Preview of Datastream’s metadata integration with Knowledge Catalog&lt;/strong&gt;. This is the first step in our vision to provide a centralized, "single pane of glass" for all Datastream assets. The enhancement automatically synchronizes Streams, Connection Profiles, and Private Connections, eliminating data silos. It enhances discoverability, allowing you to search for Datastream assets using the same interface as BigQuery tables. Centralized governance is also provided, making your real-time data estate more transparent and easier to manage.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Upgrading Apigee OPDK to 4.53 with OS Modernization&lt;br/&gt;&lt;/strong&gt;Modernize your infrastructure using Google’s official, sequential upgrade path. Our Technical expert, Rakesh Talanki outlines how to upgrade Apigee OPDK to v4.53 while migrating to a supported OS (RHEL 8.x/9.x). This guide covers the "build-out" methodology, including multi-data center syncing, to ensure a stable, zero-downtime transition&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/3Oa8uqy" rel="noopener" target="_blank"&gt;Read the guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Cloud Run Worker Pools and CREMA: Powering Serverless AI at Scale&lt;br/&gt;&lt;/strong&gt;Google Cloud has announced the General Availability of &lt;strong&gt;Cloud Run worker pools&lt;/strong&gt;, a new resource type designed specifically for pull-based, non-HTTP workloads. Unlike traditional Cloud Run services that scale based on request traffic, worker pools provide an "always-on" environment for background tasks like processing message queues or running large-scale AI inference. To support this, Google Cloud also open-sourced the &lt;strong&gt;Cloud Run External Metrics Autoscaler (CREMA)&lt;/strong&gt;. Built on KEDA, CREMA enables queue-aware autoscaling for worker pools, allowing them to dynamically scale based on external signals like Pub/Sub backlog or Kafka lag.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Apigee Model Context Protocol (MCP) now Generally Available&lt;br/&gt;&lt;/strong&gt;Expose enterprise APIs as MCP tools for agentic AI applications with the General Availability of MCP in Apigee. This update allows developers to transform APIs into AI-ready tools using OpenAPI Specifications, removing the need for local MCP servers or additional infrastructure. With managed endpoints and semantic search in API hub, you can now provide AI agents with secure, governed access to enterprise data at scale.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/3QfoEQ4" rel="noopener" target="_blank"&gt;&lt;em&gt;Explore the MCP overview&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Apr 6 - Apr 10&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Community TechTalk: Powering Retail Agents with ADK, UCP &amp;amp; Apigee X&lt;br/&gt;&lt;/strong&gt;Move beyond basic chatbots to secure, transactional AI experiences. Join our Community TechTalk on April 16 to learn how Apigee X and Gemini build a "Trust Layer" for AI shopping assistants using UCP standards. We’ll demonstrate how to block prompt injections with Model Armor and implement cost governance via token limits to secure the path from discovery to purchase.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/41ocUgq" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt;Register for the TechTalk&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Implement multimodal capabilities in your AI agents&lt;br/&gt;&lt;/strong&gt;Explore three new reference architectures for building sophisticated multi-agent AI systems that can process and analyze multimodal data. To analyze disparate multimodal data and produce a high-confidence classification, see &lt;a href="https://docs.cloud.google.com/architecture/agentic-ai-classify-multimodal-data" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="vertical-align: baseline;"&gt;Classify multimodal data&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. To create a fluid conversational AI that processes audio and video streams in real time, see&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/architecture/agentic-ai-bidirectional-multimodal-streaming" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="vertical-align: baseline;"&gt;Enable live bidirectional multimodal streaming&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. To consolidate fragmented multimodal data into a searchable knowledge graph, see&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/architecture/agentic-ai-multimodal-graph-rag-resource-orchestration" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="vertical-align: baseline;"&gt;Multimodal GraphRAG resource orchestration&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Automate SecOps workflows with an agentic AI system&lt;br/&gt;&lt;/strong&gt;To accelerate incident response and reduce manual toil for your security team, you need a system that can automate remediation playbooks. Our new reference architecture helps you build an AI agent that orchestrates complex triage and investigation workflows across disparate security tools, such as SIEM, CSPM, and EDR, from a single interface. See the full guide to &lt;a href="https://docs.cloud.google.com/architecture/agentic-ai-orchestrate-security-ops-workflows" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="vertical-align: baseline;"&gt;orchestrate security operations workflows&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Mar 30 - Apr 3&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;ASEAN Webinar | April 30: Mastering Agentic Governance at Scale with GCP&lt;br/&gt;&lt;/strong&gt;As AI agents move from experimental pilots to core enterprise functions, governance is the critical next step. Join Google Cloud experts &lt;strong&gt;Shilpi Puri &amp;amp; Wely Lau&lt;/strong&gt; for a &lt;strong&gt;webinar&lt;/strong&gt; on &lt;strong&gt;April 30th at 11:00 AM SGT&lt;/strong&gt; to learn how to architect a secure AI Management layer. We’ll explore developing governed MCP endpoints, managing tool access to enterprise data, and operationalizing AI with robust audit logs. The session includes a live demo of these frameworks in action on Google Cloud.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/47FX1Wn" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;strong&gt;RSVP here.&lt;/strong&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Mar 23 - Mar 27&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Turn your API sprawl into an agent-ready catalog&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;As organizations scale, APIs often become scattered across multiple gateways, creating "blind spots" that hinder AI adoption. To solve this, we’ve introduced two new capabilities for Apigee API hub: a new integration with API Gateway to automatically centralize API metadata into a single control plane, and a specification boost add-on (now in public preview). This add-on uses AI to enhance your API documentation with the precise examples and error codes that AI agents need to function reliably.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://goo.gle/47dEYqc" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Read the full blog post to get started.&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Webinar | April 16: AI Command &amp;amp; Control&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;As AI agents move from experimental pilots to core enterprise functions, governance is the critical next step. Join Google Cloud expert Satyam Maloo for a webinar on April 16th at 11:00 AM IST to learn how to architect a secure AI Management layer. We’ll explore developing governed MCP endpoints, managing tool access to enterprise data, and operationalizing AI with robust audit logs. The session includes a live demo of these frameworks in action on Google Cloud.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://goo.gle/4t43Vg4" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;RSVP here.&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Modernizing and Decoupling Event Ingestion with Apigee&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;In modern cloud-native architectures, decoupling producers from consumers is critical for building resilient systems. While Google Cloud Pub/Sub provides a scalable backbone, exposing it directly to external clients can introduce security and management overhead. This new guide explores how to leverage Apigee as an intelligent HTTP ingestion point. Learn how to handle security, mediation, and traffic control before messages reach your internal bus using the PublishMessage policy or Pub/Sub API.&lt;/span&gt;&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/3POgsWF" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Read the full guide.&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Mar 16 - Mar 20&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Gemini-powered Assistant in BigQuery Studio Gets Context-Aware Upgrades&lt;br/&gt;&lt;/strong&gt;The Gemini-powered assistant in BigQuery Studio has been transformed into a fully context-aware analytics partner, supporting your entire data lifecycle. The new capabilities include intelligent resource discovery, which uses Dataplex Universal Catalog search to find resources across projects and deep dive into metadata using natural language. You can now automate tasks, such as scheduling production-grade queries directly through the chat interface, and instantly troubleshoot long-running or failed jobs with root cause analysis and cost control auditing.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/use-cloud-assist"&gt;Explore&lt;/a&gt; the full range of what the assistant can do.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Mar 9 - Mar 13&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;div&gt;&lt;strong&gt;Want to use Gemini to develop code and don't know where to start?&lt;/strong&gt;&lt;br/&gt;This &lt;a href="https://medium.com/google-cloud/supercharge-your-spark-development-with-gemini-1540f1cb47d4" rel="noopener" target="_blank"&gt;article&lt;/a&gt; includes a couple of examples of developing code with Gemini prompts; it identified changes that were needed to be made to get the code working. The article also refers to other examples that are available on github. &lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Mar 2 - Mar 6&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Introducing Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model.&lt;/strong&gt; Built for high-volume developer workloads at scale, 3.1 Flash-Lite delivers high quality for its price and model tier. Gemini 3.1 Flash-Lite can tackle tasks at scale, like high-volume translation and content moderation, where cost is a priority. And it can also handle more complex workloads where more in-depth reasoning is needed, like generating user interfaces and dashboards, creating simulations or following instructions.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Starting today, 3.1 Flash-Lite is rolling out in preview to enterprises via &lt;/span&gt;&lt;a href="https://console.cloud.google.com/vertex-ai/studio/multimodal?mode=prompt&amp;amp;model=gemini-3.1-flash-lite-preview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Vertex AI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;developers via the Gemini API in &lt;/span&gt;&lt;a href="https://aistudio.google.com/prompts/new_chat?model=gemini-3.1-flash-lite-preview" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google AI Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;div&gt;
&lt;p&gt;&lt;strong&gt;TechTalk: Implementing Device Authorization Grant (RFC 8628) for Apigee&lt;/strong&gt;&lt;br/&gt;Learn how to authorize "headless" devices like Smart TVs or AI agents that lack keyboards and browsers. Join our Community TechTalk on March 19 (5PM CET / 12PM EDT) to go under the hood of Apigee X/Hybrid. We’ll cover the real-world mechanics of state management, polling, and human-in-the-loop security patterns for devices and autonomous agents.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://goo.gle/4r6o6Zi" rel="noopener" target="_blank"&gt;Register for the TechTalk&lt;/a&gt;&lt;/p&gt;
&lt;/div&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Feb 23 - Feb 27&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Pro-level image generation gets faster and more accessible with Nano Banana 2&lt;br/&gt;&lt;/strong&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Nano Banana 2 is our state-of-the-art image generation and editing model. It delivers Pro-level image generation and editing at the speed you expect from Flash — making the quality, reasoning, and world knowledge you loved about Nano Banana Pro more accessible. Learn more about the model &lt;/span&gt;&lt;a href="https://blog.google/innovation-and-ai/technology/ai/nano-banana-2" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;The Intelligent Path to Compliance: Transforming Regulatory QC with Google Cloud&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Reducing "Refuse to File" (RTF) risks and submission cycle times is critical for life sciences leaders. Google Cloud’s Regulatory Submission Semantic QC Auditor leverages Gemini and RAG architecture to transform Quality Control from a manual burden into an active, intelligent workflow.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By automating semantic cross-referencing, narrative coherence checks, and dynamic guidance-based auditing, this solution ensures rigorous accuracy and auditability. Operating within a secure GxP-ready environment, it empowers teams to detect subtle inconsistencies and generate remediation plans without sacrificing data privacy. &lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://discuss.google.dev/t/the-intelligent-path-to-compliance-transforming-regulatory-quality-control-with-google-cloud/335276" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Learn more&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Stop typing, start interacting! &lt;strong&gt;The Gemini Live Agent Challenge is here&lt;/strong&gt;. Build immersive agents that can help you see, hear, and speak using Gemini and Google Cloud. Compete for your share of $80,000+ in prizes and a trip to Google Cloud Next '26!&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Submissions are open from February 16, 2026 to March 16, 2026. Learn more and register at &lt;/span&gt;&lt;a href="http://geminiliveagentchallenge.devpost.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;geminiliveagentchallenge.devpost.com&lt;/span&gt;&lt;/a&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Feb 9 - Feb 13&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Introducing Gemini 3.1 Pro on Google Cloud. &lt;/span&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;span style="vertical-align: baseline;"&gt;3.1 Pro is a noticeably smarter, more capable baseline for complex problem-solving. We’re shipping 3.1 Pro at scale, building upon our &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/gemini-3-is-available-for-enterprise?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;goal&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to help you transform your business for the agentic future. Learn more about the model’s capabilities &lt;/span&gt;&lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Gemini 3.1 Pro is available starting today in preview in &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Vertex AI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Developers can access the model in preview via the Gemini API in &lt;/span&gt;&lt;a href="https://aistudio.google.com/prompts/new_chat?model=gemini-3.1-pro-preview" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google AI Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://developer.android.com/studio" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Android Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://antigravity.google/blog/gemini-3-1-in-google-antigravity" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Antigravity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://geminicli.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini CLI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Automate Storage Compatibility with GKE Dynamic Default Storage Classes&lt;br/&gt;&lt;/strong&gt;Managing storage across mixed-generation VM clusters in GKE just got easier. With the new &lt;strong&gt;Dynamic Default Storage Class&lt;/strong&gt;, Google Kubernetes Engine automatically selects between Persistent Disk (PD) and Hyperdisk based on a node's specific hardware compatibility. This abstraction eliminates the need for complex scheduling rules and manual pairing, ensuring your volumes "just work" regardless of the underlying infrastructure. By defining both variants in a single class, you reduce operational overhead while maintaining peak performance and cost-efficiency across your entire cluster.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/hyperdisk#automated_disk_type_selection" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;Explore automated disk type selection&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Community TechTalk: AI-Powered Apigee Development with strofa.io&lt;br/&gt;&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;Join the Apigee community on February 26&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; for a deep dive into&lt;/span&gt; &lt;a href="https://www.google.com/search?q=http://strofa.io" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;strofa.io&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Guest speaker Denis Kalitviansky will demonstrate how this new AI-powered tool automates and orchestrates Apigee development, from local emulators to large-scale hybrid environments. Discover how to scale your API management and streamline team collaboration using the latest in AI-driven automation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://goo.gle/3Oerns3" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Register now to reserve your spot.&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jan 26 - Jan 30&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Simplify API Governance with Native OpenAPI v3 Support&lt;br/&gt;&lt;/span&gt;&lt;/strong&gt;Eliminate integration debt and accelerate deployment velocity with the General Availability of OpenAPI v3 (OASv3) support for API Gateway and Cloud Endpoints. You no longer need to downgrade modern specifications to OASv2. Instead, you can now define API contracts and enforce critical policies—including telemetry, quotas, and security—using native Google-specific extensions directly within your OASv3 files. This update ensures your APIs are secure by design while remaining fully compatible with the modern developer ecosystem and Google Cloud’s AI services.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/49Wx58Z" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Get started with OpenAPI v3 on API Gateway and Cloud Endpoints.&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Accelerate API Testing with the New Open Source API Tester&lt;br/&gt;&lt;/span&gt;&lt;/strong&gt;Start validating your APIs with API Tester, a simple, YAML-based Test Driven Development (TDD) framework. Designed for the Apigee community, this tool allows you to write human-readable tests, run them instantly via a web client or CLI, and perform deep unit testing on Apigee proxies. With native support for JSONPath assertions and Apigee shared flows, you can verify everything from payload data to internal variables like &lt;code style="vertical-align: baseline;"&gt;proxy.basepath&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; without leaving your terminal.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://goo.gle/4q5WDGK" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Explore the API Tester guide and start testing your proxies today.&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Secure Sensitive Data with Kubernetes Secrets in Apigee hybrid&lt;br/&gt;&lt;/span&gt;&lt;/strong&gt;Enhance security in Apigee hybrid by accessing Kubernetes Secrets directly within your API proxies. This hybrid-exclusive feature keeps sensitive credentials within your cluster boundary and prevents replication to the management plane. It supports strict separation of duties: operators manage secrets via &lt;code style="vertical-align: baseline;"&gt;kubectl&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, while developers reference them as secure flow variables—ideal for high-compliance and GitOps workflows.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://goo.gle/4qEVffo" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Implement Kubernetes Secrets in your hybrid proxies.&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;See the Console in a Whole New Light: Dark Mode is Now Generally Available in Google Cloud&lt;br/&gt;&lt;/span&gt;&lt;/strong&gt;Elevate your cloud management workflow with Dark Mode, now generally available in the Google Cloud console. We have delivered a modern, cohesive, and accessible experience reimagined for maximum comfort and productivity—especially during extended working hours and low-light environments. Dark Mode can be enabled automatically based on your operating system's preference, or manually through the Settings  -&amp;gt; Appearance menu.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://docs.cloud.google.com/docs/get-started/console-appearance" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Switch to Dark Mode today to enjoy a modern, comfortable, and productive environment!&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Apigee X Networking: PSC or VPC Peering?&lt;br/&gt;&lt;/span&gt;&lt;/strong&gt;Deciding how to connect Apigee X? Watch this video to compare Private Service Connect and VPC Peering. We break down northbound and southbound routing, IP consumption, and how to reach targets on-prem or in the cloud. Learn to simplify your architecture and avoid common networking "gotchas" for a smoother deployment.&lt;br/&gt;&lt;br/&gt;&lt;a href="https://goo.gle/4bWBGdV" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Watch the video.&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 data-draftjs-conductor-fragment='{"blocks":[{"key":"865rk","text":"Week of Dec 16 - Dec 20","type":"header-three","depth":0,"inlineStyleRanges":[],"entityRanges":[],"data":{}}],"entityMap":{}}'&gt;Jan 19 - Jan 23&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Bridge the Gap: Excel-to-API Conversion in Apigee Portals&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Give your customers more ways to connect! This new article by Tyler Ayers explores how to extend the Apigee Integrated Portal to support direct Excel file uploads. By leveraging SheetJS and custom portal scripts, you can enable users to upload spreadsheets, preview data, and submit it directly to your APIs, all without writing a single line of integration code themselves. It’s a powerful way to simplify onboarding for those who aren't yet API-ready.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://goo.gle/3Nq3Pjo" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Learn how to build it&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Elevate your applications with Firestore’s new advanced query engine&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We have fundamentally reimagined Firestore with pipeline operations for Enterprise edition. Experience a powerful new engine featuring over a hundred new query features, index-less queries, new index types, and observability tooling to improve query performance. Seamlessly migrate using built-in tools and leverage Firestore’s existing differentiated serverless foundation, virtually unlimited scale, and industry-leading SLA. Join a community of 600K developers to craft expressive applications that maximize the benefits of rich queryability, real-time listen queries, robust offline caching, and cutting-edge AI-assistive coding integrations.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/new-firestore-query-engine-enables-pipelines?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Learn more about Firestore pipeline operations.&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 02 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/inside-google-cloud/whats-new-google-cloud/</guid><category>Google Cloud</category><category>Inside Google Cloud</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/whats_new_2026_CfhxFWX.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>What’s new with Google Cloud</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/whats_new_2026_CfhxFWX.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/inside-google-cloud/whats-new-google-cloud/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Google Cloud Content &amp; Editorial </name><title></title><department></department><company></company></author></item><item><title>How to implement long-term AI agent memory in AlloyDB and Memorystore for Valkey</title><link>https://cloud.google.com/blog/products/databases/implementing-long-term-ai-agent-memory-in-alloydb-and-memorystore/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Enterprise AI agents need persistent memory to execute complex, multi-day workflows and long-horizon tasks. In this blog, we examine how a 2-tier memory architecture using &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/memorystore/docs/valkey"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; for short-term buffer memory and &lt;/span&gt;&lt;a href="https://cloud.google.com/alloydb/ai"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB AI&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; for long-term persistent memory can help reduce token spend by up to 70%, while maintaining critical data and enterprise guardrails.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Imagine building a personalized travel agent designed to help users book vacations. The user interacts with the agent many times over the course of several days, asking questions that range from brainstorming itineraries to actual purchase intent. To provide a truly seamless experience, this agent must remember flight preferences (e.g. “I only want non-stop flights”), hotel budgets, and dietary restrictions (e.g. “I need Gluten Free dining options”) established in previous sessions. More importantly, it has to hold onto these core facts even when the conversation gets deep into the weeds of sightseeing recommendations and itinerary planning. The agent must ensure that all the follow-up questions and exciting details about places to visit doesn’t cause it to forget or overwrite the user's fundamental requirements and decisions. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Stateful workflows inherently conflict with LLM’s stateless nature&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI agents are being used for multi-turn conversations and long-running workflows, but large language models (LLMs) remain stateless across sessions. When a user returns to an agent days later, the model starts with an empty context window. Unless your application is built to reconstruct past context using long-term agent memory, your users have to explain their goals and context all over again, resulting in a frustrating and fragmented experience.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With million-token context windows now the norm, a common shortcut to this problem is "context stuffing" -  dumping everything you can fit into the prompt at every turn, from raw chat histories to tool execution logs. In fact, this was a common pattern in the early days of AI model usage. But this shortcut quickly creates issues at scale: token costs multiply with every message, response times can drag out past 30+ seconds for otherwise simple prompts, and the model starts suffering from "&lt;/span&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;lost in the middle&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;" degradation, overlooking critical instructions buried in mountains of prompt text.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Another common workaround is to use rolling summaries; asking an LLM to periodically compress older messages into a summary paragraph. While this trims prompt size, LLM-based summarization is inherently lossy. After a few rounds of compression, subtle but important details get filtered out as background noise. A few turns later, your agent quietly breaks the exact constraints you set earlier. Clearly, you  need a more scalable approach for keeping a memory from previous conversations or multi-turn tasks, but without degrading the experience or creating new bottlenecks.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Implementing a 2-tier memory architecture&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To build reliable, cost-effective enterprise agents that respect your guardrails and constraints, we recommend a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;2-tier memory architecture&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Short-term session buffer&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Caches active conversation turns in memory using a token-bounded sliding window, allowing the context window size to remain stable across multiple rounds, even across devices, while keeping the latest messages fresh. This tier requires sub-millisecond, high-throughput lookups on every turn, making &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/memorystore/docs/valkey"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; well suited for maintaining active session state. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Long-term persistent memory&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Stores important facts, user preferences, and episodic facts across sessions. This tier requires transactional integrity, data governance, and hybrid retrieval across relational data and vectors - capabilities provided natively by &lt;/span&gt;&lt;a href="https://cloud.google.com/alloydb/ai"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB AI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While it’s possible to store both long-term memory and active session buffers directly in a relational database, a two-tier architecture with an in-memory cache provides better performance and scalability. Active session buffers are highly ephemeral and require sub-millisecond, high-throughput updates on every single conversation turn. Handling these rapid-fire writes in an in-memory key-value cache prevents write amplification and table bloat in your relational database, which would otherwise require frequent row deletions and intensive vacuuming. This is analogous to adding a caching tier in front of your database to offload high-frequency lookups for hot rows. The division of labor keeps your primary database lean and responsive, allowing it to focus its resources on what it is designed for: transactional consistency, complex hybrid vector search, and long-term analytical query execution. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By pairing short-term caching in Memorystore for Valkey with native AlloyDB AI capabilities, you can run entity extraction, memory compaction, cross-session memory, and hybrid retrieval directly inside the database tier, keeping active prompts lean, fast, and cost-efficient. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Get started with a 30-day AlloyDB free trial instance&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42ef33750&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;image&amp;#x27;, None)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Understanding the four memory types&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To organize long-term agent state effectively, we further divide agent memory into four complementary types across the short-term and long-term storage tiers we described above:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Memory type&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;What it stores&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Storage layer&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Lifespan&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Buffer (Short-term)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Recent raw conversation turns&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Active session&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Summary memory&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Compressed history of older turns, commonly referred to as “compaction”&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Multi-turn window&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Episodic memory&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Past actions, events, and tool outputs&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AlloyDB for PostgreSQL (Hybrid retrieval with structured SQL + full-text search + vector)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Permanent&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Entity &amp;amp; rule memory&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;User preferences, constraints, and vetoes&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AlloyDB for PostgreSQL (Hybrid retrieval with structured SQL + full-text search + vector)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Permanent&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By isolating short-term conversation context from structured long-term rules, your agent retrieves relevant context on demand without filling token windows with raw interaction logs.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The tiered memory architectural blueprint&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The diagram below outlines the read and write paths connecting the application orchestration layer, the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; short-term buffer, and the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB AI&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; long-term repository:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_dkY04WA.max-1000x1000.jpg"
        
          alt="image1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The system operates across two coordinated execution paths:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The read path&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: When a user asks a question, the application fetches the active sliding window from &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, runs in-database query normalization using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ai.generate&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, and queries &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ai.hybrid_search&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to retrieve scoped entity rules and relevant episodic facts.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The write path&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: After generating the response, the turn is immediately cached in &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. An asynchronous background queue worker extracts structured entities from the exchange and writes them directly into &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, where transactional auto-embeddings immediately compute and store vector representations in the database.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Business impact and ROI&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In our benchmark testing across multi-turn development dialogues (45+ turns with heavy tool executions), separating short-term caching from long-term persistence delivered measurable cost and performance improvements compared to naive context stuffing:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Metric / dimension&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Naive context stuffing&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Tiered memory (AlloyDB + Valkey)&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Net business impact&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Active prompt size (turn 45)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;747,033 tokens&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;83,262 tokens&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;88.9% smaller prompt&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Turn 45 response latency&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;33.5 seconds&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;6.7 seconds&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;80.0% faster response&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Per-turn response wait time&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;33.5 seconds&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;4.2s – 6.7s&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;36% to 80% reduction&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Cumulative session tokens&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;17.9M tokens&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;4.09M tokens&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;72.0% token &amp;amp; cost savings&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rule &amp;amp; constraint recall&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Degrades over turns&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Does not degrade &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;ACID-preserved recall&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Note: The metrics above reflect internal benchmark results from a simulated developer workload that mimics a real-world enterprise AI pair-programming assistant interacting with a developer over multiple sessions, projects, and context switches. Actual savings and latencies vary based on prompt structure, query frequency, and data volume.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These results demonstrate how tiered memory changes agent unit economics: instead of an escalating cost curve on every additional turn, prompt sizes remain bounded, reducing ongoing LLM API expenses while keeping response times fast.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Core AlloyDB AI technical advantages&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB AI&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; reduces the operational overhead of implementing persistent agent memory by embedding core AI functions directly into the database engine:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Transactional in-database auto-embeddings (&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/generate-manage-auto-embeddings-for-tables"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;ai.initialize_embeddings&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: AlloyDB automatically generates vector embeddings for text columns using a native integration with &lt;/span&gt;&lt;a href="https://cloud.google.com/products/gemini-enterprise-agent-platform?e=48754805"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Platform&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (formerly Vertex AI), generating up to 3,000 embeddings per second. Using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;incremental_refresh_mode =&amp;gt; 'transactional'&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, AlloyDB keeps embeddings up to date as source data changes within the same transaction, removing the need for custom embedding pipelines, external schedulers, or complex retry logic.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;In-database generative AI functions (&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/evaluate-semantic-queries-ai-operators"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;ai.generate&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: AlloyDB allows you to execute foundation models, such as Gemini, directly from SQL queries. You can use this for in-database query decomposition - breaking compound user questions into single-aspect sub-queries and resolving relative time phrases (like "last session") into explicit identifiers - without making separate roundtrips from your application.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Native hybrid search with built-in Reciprocal Rank Fusion (&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;ai.hybrid_search&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: AlloyDB provides a built-in SQL function that executes Reciprocal Rank Fusion (RRF) directly inside the engine. It combines vector cosine similarity (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;&amp;lt;=&amp;gt;&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; over HNSW or ScaNN indexes) with PostgreSQL full-text search (using BM25, RUM, or GIN) in a single database call, blending semantic matching with exact keyword retrieval while supporting metadata filter pushdown (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;filter_condition&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) for improved performance and deterministic scope isolation.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Direct Agent Platform integration with IAM credentials&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: AlloyDB connects directly to &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Agent Platform&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; foundation models over Google Cloud's private network using &lt;/span&gt;&lt;a href="https://cloud.google.com/products/iam"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud IAM&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; service account roles and database authentication, avoiding the need to store, rotate, or pass API keys in application code.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Unified operational, vector, and governance engine&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: AlloyDB consolidates relational business data, vector embeddings, full-text indexes, and enterprise permissions in a single ACID-compliant PostgreSQL database, avoiding data drift and integration complexity across separate operational and vector databases.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Key implementation patterns&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Below are the core database patterns used to configure the 2-tier memory architecture. For complete, runnable Python and SQL scripts, refer to the companion &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/alloydb-agentic-tiered-memory" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB Agent Memory Codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;1. Setting up schema and auto-embeddings&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, install the necessary extensions and define the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;agent_entities&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; table with structured metadata, a generated &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tsvector&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; column for full-text search, and a vector embedding column.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- 1. Enable AI extensions\r\nCREATE EXTENSION IF NOT EXISTS google_ml_integration CASCADE;\r\nCREATE EXTENSION IF NOT EXISTS vector CASCADE;\r\nCREATE EXTENSION IF NOT EXISTS rum CASCADE;\r\n\r\n-- 2. Create structured entity &amp;amp; preference table\r\nCREATE TABLE IF NOT EXISTS agent_entities (\r\n    entity_id UUID PRIMARY KEY DEFAULT gen_random_uuid(),\r\n    user_id TEXT NOT NULL,\r\n    entity_name TEXT NOT NULL,\r\n    project_id TEXT DEFAULT &amp;#x27;global&amp;#x27;,\r\n    session_id TEXT,\r\n    scope TEXT NOT NULL DEFAULT &amp;#x27;session&amp;#x27;,\r\n    summary TEXT NOT NULL,\r\n    summary_embedding VECTOR(768),\r\n    summary_tsv TSVECTOR GENERATED ALWAYS AS (\r\n        to_tsvector(&amp;#x27;english&amp;#x27;, entity_name || &amp;#x27; &amp;#x27; || summary)\r\n    ) STORED,\r\n    updated_at TIMESTAMPTZ DEFAULT NOW(),\r\n    UNIQUE (user_id, entity_name)\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f454c90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then, generate embeddings using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ai.initialize_embeddings&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. Using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;transactional&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; mode ensures the embeddings are kept up to date as source data changes.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Register transactional in-database auto-embedding\r\nCALL ai.initialize_embeddings(\r\n    model_id =&amp;gt; &amp;#x27;text-embedding-005&amp;#x27;,\r\n    table_name =&amp;gt; &amp;#x27;agent_entities&amp;#x27;,\r\n    content_column =&amp;gt; &amp;#x27;summary&amp;#x27;,\r\n    embedding_column =&amp;gt; &amp;#x27;summary_embedding&amp;#x27;,\r\n    incremental_refresh_mode =&amp;gt; &amp;#x27;transactional&amp;#x27;,\r\n    batch_size =&amp;gt; 10\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f454490&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Finally, create the HNSW vector index and the RUM full-text search index to ensure your hybrid searches are fast and efficient.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;-- Create vector and full-text indexes\r\nCREATE INDEX IF NOT EXISTS agent_entities_embedding_idx\r\n    ON agent_entities USING hnsw (summary_embedding vector_cosine_ops);\r\n\r\nCREATE INDEX IF NOT EXISTS agent_entities_tsv_idx\r\n    ON agent_entities USING rum (summary_tsv rum_tsvector_ops);&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42ef32d10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;2. Querying long-term memory with native hybrid search&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;On the read path, retrieve relevant long-term entities using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ai.hybrid_search&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. This native SQL function executes Reciprocal Rank Fusion (RRF) directly in AlloyDB, seamlessly reranking and combining vector similarity search and full-text keyword search results in a single database query:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;def retrieve_hybrid_entities(\r\n    conn,\r\n    user_id: str,\r\n    project_id: str,\r\n    query_text: str,\r\n    query_vector_literal: str,\r\n    limit: int = 3\r\n) -&amp;gt; dict[str, Any]:\r\n    &amp;quot;&amp;quot;&amp;quot;Queries long-term memory using AlloyDB native hybrid_search (RRF).&amp;quot;&amp;quot;&amp;quot;\r\n    filter_cond = f&amp;quot;user_id = \&amp;#x27;{user_id}\&amp;#x27; AND (project_id = \&amp;#x27;{project_id}\&amp;#x27; OR scope = \&amp;#x27;global\&amp;#x27;)&amp;quot;\r\n    \r\n    search_inputs = [\r\n        json.dumps({\r\n            &amp;quot;data_type&amp;quot;: &amp;quot;vector&amp;quot;,\r\n            &amp;quot;weight&amp;quot;: 0.4,\r\n            &amp;quot;table_name&amp;quot;: &amp;quot;agent_entities&amp;quot;,\r\n            &amp;quot;key_column&amp;quot;: &amp;quot;entity_name&amp;quot;,\r\n            &amp;quot;vec_column&amp;quot;: &amp;quot;summary_embedding&amp;quot;,\r\n            &amp;quot;distance_operator&amp;quot;: &amp;quot;&amp;lt;=&amp;gt;&amp;quot;,\r\n            &amp;quot;limit&amp;quot;: 10,\r\n            &amp;quot;query_vector&amp;quot;: query_vector_literal,\r\n            &amp;quot;filter_condition&amp;quot;: filter_cond\r\n        }),\r\n        json.dumps({\r\n            &amp;quot;data_type&amp;quot;: &amp;quot;text&amp;quot;,\r\n            &amp;quot;weight&amp;quot;: 0.6,\r\n            &amp;quot;table_name&amp;quot;: &amp;quot;agent_entities&amp;quot;,\r\n            &amp;quot;key_column&amp;quot;: &amp;quot;entity_name&amp;quot;,\r\n            &amp;quot;text_column&amp;quot;: &amp;quot;summary_tsv&amp;quot;,\r\n            &amp;quot;limit&amp;quot;: 10,\r\n            &amp;quot;ranking_function&amp;quot;: &amp;quot;ts_rank&amp;quot;,\r\n            &amp;quot;query_text_input&amp;quot;: query_text,\r\n            &amp;quot;filter_condition&amp;quot;: filter_cond\r\n        })\r\n    ]\r\n\r\n    query_sql = &amp;quot;&amp;quot;&amp;quot;\r\n        SELECT e.entity_name, e.summary, e.project_id, e.scope, e.updated_at\r\n        FROM ai.hybrid_search(\r\n          search_inputs =&amp;gt; %s::JSONB[],\r\n          include_json_output =&amp;gt; false\r\n        ) h\r\n        JOIN agent_entities e ON e.entity_name = h.id\r\n        WHERE e.user_id = %s\r\n        LIMIT %s;\r\n    &amp;quot;&amp;quot;&amp;quot;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42ef33e90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;3. In-database memory compaction with AI functions&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To manage long-term storage growth without writing custom pruning scripts, you can run automated extraction and compaction queries directly in &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; to identify the important parts which can benefit from being stored. This pattern uses a SQL common table expression (CTE) with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ai.generate&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to consolidate older episodic entries into a high-density summary:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Consolidate older episodic logs into a permanent summary\r\nWITH old_events AS (\r\n    SELECT document AS content\r\n    FROM langchain_pg_embedding\r\n    WHERE cmetadata-&amp;gt;&amp;gt;&amp;#x27;user_id&amp;#x27; = &amp;#x27;user_dev_42&amp;#x27;\r\n      AND created_at &amp;lt; NOW() - INTERVAL &amp;#x27;30 days&amp;#x27;\r\n    ORDER BY created_at ASC\r\n    LIMIT 50\r\n),\r\nconsolidated AS (\r\n    SELECT ai.generate(\r\n        &amp;#x27;Summarize these historical events into a dense memory paragraph:\\n&amp;#x27; || string_agg(content, E&amp;#x27;\\n&amp;#x27;),\r\n        model_id =&amp;gt; &amp;#x27;gemini-3.5-flash&amp;#x27;\r\n    ) AS summary_text\r\n    FROM old_events\r\n)\r\nINSERT INTO agent_entities (user_id, entity_name, summary, updated_at)\r\nSELECT &amp;#x27;user_dev_42&amp;#x27;, &amp;#x27;longterm_session_summary&amp;#x27;, summary_text, NOW()\r\nFROM consolidated\r\nWHERE summary_text IS NOT NULL\r\nON CONFLICT (user_id, entity_name)\r\nDO UPDATE SET summary = EXCLUDED.summary, updated_at = NOW();&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42ef31cd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For example, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;a long interaction regarding the complexities of changing flights with kids can result in a summary of “I prefer non-stop flights”.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;4. Connecting the 2-tier memory architecture to an ADK agent&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With the core 2-tier memory architecture configured, you can now extend the default ADK Memory provider to use this 2-tier memory architecture as shown in the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/alloydb-agentic-tiered-memory#14" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;accompanying Codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (e.g. &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ADKTieredMemoryProvider&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). To attach the long-term memory to an ADK agent, you simply provide it as a tool (e.g. &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;longterm_memory_tool&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) like this: &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from adk_memory_provider import ADKTieredMemoryProvider\r\n\r\nagent = Agent(\r\n    name=&amp;quot;adk_memory_agent&amp;quot;, model=GEMINI_MODEL, instruction=&amp;quot;Initial instruction&amp;quot;,\r\n    tools=[long_term_memory_tool, run_command_tool],\r\n    before_tool_callback=guardrail.create_before_tool_callback(USER_ID, &amp;quot;CloudRetail&amp;quot;)\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42ef33590&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise governance and multi-tenant security&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Running agent memory in enterprise production environments requires strict security boundaries and access controls. First of all, we need to maintain &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;scope and multi-tenant isolation.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; By indexing &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;user_id&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;project_id&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;scope&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; columns in &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;agent_entities&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and enforcing PostgreSQL Row-Level Security (RLS), you can isolate memory stores across departments, teams, and individual users within the same database cluster. &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/parameterized-secure-views-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Parameterized Secure Views&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (PSV) offer another layer of deterministic application-level security, helping you protect against malicious prompts and overly-broad SQL queries.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition, we need to provide &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;automated memory lifecycle management &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;as shown in the implementation pattern above. Combining scheduled SQL compaction queries with time-based partition pruning helps maintain predictable database footprint and query latencies over time. See the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/alloydb-agentic-tiered-memory" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;accompanying Codelab &lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;for more details on this approach.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Summary and next steps&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Decoupling active context windows from persistent storage is a practical approach to building production-ready AI agents. By pairing &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Memorystore for Valkey&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; for sub-millisecond session caching with &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;AlloyDB AI&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; for transactional long-term storage, you can achieve substantial token cost savings and faster response times while maintaining strict business rules throughout long-horizon tasks and many-turn agentic experiences.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To get started:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Step through the complete hands-on tutorial in the companion &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/alloydb-agentic-tiered-memory" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB Agent Memory Codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to deploy the working 2-tier memory architecture.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Learn more about database-side machine learning features in the &lt;/span&gt;&lt;a href="https://cloud.google.com/alloydb/docs/ai/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB AI documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explore guides on &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/generate-manage-auto-embeddings-for-tables"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;generating auto vector embeddings&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/alloydb/docs/ai/run-hybrid-vector-similarity-search"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;running hybrid vector search&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 02 Oct 2026 07:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/databases/implementing-long-term-ai-agent-memory-in-alloydb-and-memorystore/</guid><category>Databases</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How to implement long-term AI agent memory in AlloyDB and Memorystore for Valkey</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/databases/implementing-long-term-ai-agent-memory-in-alloydb-and-memorystore/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Itai Rosenblatt</name><title>Senior Engineering Manager, AlloyDB AI, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Paul Ramsey</name><title>Product Manager, AlloyDB AI &amp; Cloud SQL, Google</title><department></department><company></company></author></item><item><title>Enabling Cloud Storage end-to-end checksums for improved data integrity and durability</title><link>https://cloud.google.com/blog/products/storage-data-transfer/enabling-end-to-end-checksums-in-cloud-storage/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At Google Cloud, we know that you count on us to maintain the durability and integrity of your data at all times, both at rest and in transit. And now we’re making it easier for developers to take advantage of native data integrity features in Cloud Storage, by enabling end-to-end checksumming by default in all the Cloud Storage SDKs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Like in any disk-based storage system, bits can flip anywhere in their journey, from the application all the way down to the disk. Cloud Storage has always let clients provide a checksum of the object data being uploaded, and receive a checksum of the data being downloaded. Also since its inception, Cloud Storage stores a checksum for every object in its metadata, regardless of how the object was uploaded into Cloud Storage.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;But until recently, ensuring end-to-end data integrity required extra work on the part of developers to calculate and provide checksums to Cloud Storage. Cloud Storage always calculates the crc32 (32-bit cyclic redundancy check) of data it receives and ensures data stored on disk matches this checksum. When a client request includes the object’s checksum, Cloud Storage ensures that this checksum also matches. However, when an upload request doesn’t include a checksum, that upload is vulnerable to a bit flip while the data is in-flight, prior to the server-side checksum computation. Not all customers and clients enable client-side checksums by default, leaving data in this phase unprotected. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To address this gap, the latest version of all Cloud Storage SDKs now internally checksums data being uploaded and passes this checksum to Cloud Storage, if it’s not provided by the application. The SDKs also support verifying the object’s checksum when an object is being downloaded.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Finally, there are many use-cases where applications download select ranges of objects instead of the full object. When using Cloud Storage SDKs with our gRPC API to perform a range read, the SDKs take advantage of gRPC’s built-in end-to-end range checksum, using it to verify the data it receives.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We highly recommend &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/storage/docs/data-validation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;updating to our latest SDK versions&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to take advantage of these important integrity features.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;And now, let’s peek under the hood&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ensuring continuous “chain-of-custody” between the data and its associated checksum from your application down to the disk platter, with no gap where a bit flip could go unnoticed, is quite challenging. And it’s critical to get this right: at our current scale of hundreds of thousands of Cloud Storage frontends, bit flips aren’t theoretical and do happen from time to time.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For instance, consider this simple example: when Cloud Storage receives your data in its frontend, this data gets encrypted with per-object encryption keys. This involves a data copy: the plaintext data is passed through an encryptor into a new memory buffer containing ciphertext. Extremely rarely, a bit in the source or destination memory buffer flips during this process. However, at our scale, extremely rare things happen routinely. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this situation, we maintain chain-of-custody by reversing the whole process: after encrypting the data (1), we calculate a checksum that protects the ciphertext. Then we decrypt the ciphertext (2), and if the resulting plaintext doesn’t match the original (3), we throw everything away and start over. This adds up to a lot of extra CPU time spent on encryption and checksumming, but it’s a necessary step to ensure data integrity.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_XOxHkrR.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Another challenge is how data gets broken up and aggregated as it passes through layers of our stack. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;As data gets uploaded to Cloud Storage, it gets split up into chunks, each of which has its own checksum&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. To manage data efficiently at scale, Cloud Storage groups thousands of chunks together into a storage unit we call a shard file. These gigabyte-sized files are how Cloud Storage ultimately delivers data to our cluster-level storage system, &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/a-peek-behind-colossus-googles-file-system?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Colossus&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Internally, Colossus uses Reed Solomon encodings to spread data across many disks and protect against the failures of individual disks, machines, and racks. This requires chopping up the shard file data into blocks, each of which is again protected by a checksum.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To maintain chain-of-custody of the data as it goes through all these transformations, we take advantage of some nifty properties of cyclic redundancy checks (CRCs), for example, concatenation. When you have two data buffers that each have their own CRC, you can cheaply compute the CRC of the two concatenated buffers without having to re-checksum the data. This comes in handy in many situations, such as when concatenating chunks together into shard files: Colossus can cheaply determine the CRC of the entire shard file from its constituent chunks and store that in its metadata.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ultimately, the data lands on disks managed by our “D” file server (our network attached disks). D stores inline checksums for each range of data within a Colossus block. Whenever data is read from the disk, it is verified at several layers: The Colossus client verifies the data it reads against D’s inline checksums, and the Cloud Storage frontend reads data chunk-by-chunk, verifying each chunk against its checksum before sending it to the client. These chunk-level checksums are what enable our gRPC protocol to provide a checksum for a range read that can be verified by our SDKs, all without losing chain-of-custody.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_wfC90pO.max-1000x1000.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="94cbo"&gt;Chain of Custody: Maintaining Data integrity across Data transformations&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;The above image shows the data integrity handoff across multiple layers under the hood of Google Cloud Storage. On reads, checksums are verified inline at multiple layers to prevent silent corruptions.&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; font-style: italic; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Client passes full object checksum to Cloud Storage Frontends.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; font-style: italic; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Data is split into chunks and individual chunk level checksums are computed.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; font-style: italic; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Shard level checksums are computed based on concatenated chunk level CRCs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Shards are stored across disk blocks with another level of block level inline checksums. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here on the Cloud Storage team, we remain dedicated to maintaining the highest standards of data integrity for our customers. By making end-to-end checksumming the default in our SDKs and maintaining chain-of-custody throughout our internal storage stack, data remains exactly as intended from the moment of upload to the final download. This continuous vigilance reflects our commitment to protecting your data at any scale. To take full advantage of these protections, we recommend updating to the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/storage/docs/data-validation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;latest&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; version of our &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/storage/docs/reference/libraries"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;SDKs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/storage-data-transfer/enabling-end-to-end-checksums-in-cloud-storage/</guid><category>Storage &amp; Data Transfer</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Enabling Cloud Storage end-to-end checksums for improved data integrity and durability</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/storage-data-transfer/enabling-end-to-end-checksums-in-cloud-storage/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Denis Serenyi</name><title>Distinguished Software Engineer</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Yamini Allu</name><title>Software Engineer</title><department></department><company></company></author></item><item><title>Accelerating analytics: PayPal’s journey with Managed Service for Apache Spark</title><link>https://cloud.google.com/blog/products/data-analytics/paypals-journey-with-managed-service-for-apache-spark/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In a data-driven world, PayPal’s ability to deliver timely and actionable insights is central to staying ahead. At PayPal, data powers everything from fraud detection to user experience enhancements. Data is also central to unleashing the potential of agentic solutions and experiences. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Over time, though, our analytics environment had become a complex ecosystem of various technologies and solutions assembled on-premise to address growing demands. While this approach supported our needs at the time, it began presenting new challenges to scale and maintain.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Navigating a challenging analytics landscape&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Due to expedited growth and acquisitions, our data analytics platform gradually turned into an uneven landscape. Each new platform or integration addressed a specific business need, but together, they increased operational overhead and introduced performance blockages. Scalability became increasingly difficult, and time-to-insight slowed as processes grew more complex. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Complexity breeds stagnation&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;PayPal’s &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/paypals-historic-data-migration-is-the-foundation-for-its-gen-ai-innovation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;legacy data analytics platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; was powerful—handling petabytes daily—but it was also increasingly rigid following rapid growth. Scaling up during peak retail events or global launches meant months of planning, slow manual provisioning of hardware, and too often, a compromise between speed and cost.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As PayPal continued to scale globally, we recognized the need for a streamlined, unified infrastructure to drive data efficiency and accelerate innovation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The solution: Unified, cloud-native analytics&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To overcome these obstacles, we migrated our analytics workloads from legacy Hadoop on-premise platforms to Google’s &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Key reasons for this choice included:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Rapid provisioning and elastic scaling: Managed Spark enabled us to deploy clusters in minutes and scale based on processing needs, eliminating lengthy setup and idle resource costs. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Unified infrastructure: Standardizing on Apache Spark created consistency across teams while leveraging Managed Service for Apache Spark and other managed services reduced operational complexity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Seamless integration: Native hooks into &lt;/span&gt;&lt;a href="https://cloud.google.com/storage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GCS), &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and other Google Cloud services streamlined end-to-end data movement.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This move enabled PayPal to modernize our data processing capabilities, leveraging the flexibility, scalability, and reliability of cloud-native solutions. By consolidating previously disparate workflows and batch jobs that run on multiple platforms onto a single cloud-based analytics platform, we reduced data silos and built a unified data foundation that provides faster, richer insights. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This empowered developers and application teams to focus on delivering business value rather than being limited by infrastructure. Crucially, this shift was about more than re-platforming. We fostered a new culture of experimentation, enabling teams to test, tune, and deploy analytics workloads quickly in response to changing business needs.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The results: Faster insights, lower overhead&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The impact of our modernized Google Cloud-based ecosystem leveraging Managed Spark has been profound:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Processing times for core analytics workloads improved by 25%, enabling near real-time insights for key business operations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;SLA adherence rose substantially by 30%, even during traffic surges such as seasonal sales events.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Operational costs dropped as we consolidated tooling and reduced manual maintenance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;But perhaps most importantly, our engineers now spend less time firefighting and more time innovating, rapidly prototyping new analytics capabilities that deliver value to customers and partners.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Transitioning from a fragmented environment to a cohesive, cloud-native platform has fundamentally strengthened PayPal’s analytics capabilities. As business needs evolve, investing in a scalable, unified data foundation ensures that we can deliver insights with speed, precision, and impact—driving continued innovation for customers worldwide. Our journey with Managed Service for Apache  Spark is an important step in building that modern analytics foundation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Learn more about how you can get started with &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; to build your &lt;/span&gt;&lt;a href="https://cloud.google.com/data-cloud"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Agentic Data Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; today.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/paypals-journey-with-managed-service-for-apache-spark/</guid><category>Financial Services</category><category>Data Analytics</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/paypal-apache-spark.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Accelerating analytics: PayPal’s journey with Managed Service for Apache Spark</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/paypal-apache-spark.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/paypals-journey-with-managed-service-for-apache-spark/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vinod Ganesan</name><title>Director, Analytics Reliability Engineering, PayPal</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Raghu Agani</name><title>Sr. Manager, Big Data Platforms Engineering, PayPal</title><department></department><company></company></author></item><item><title>Democratizing Managed Lustre with lower cost and frictionless development</title><link>https://cloud.google.com/blog/topics/developers-practitioners/democratizing-managed-lustre-with-lower-cost-and-frictionless-development/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This is the first of a two-part series exploring how Google Cloud is bringing the foundational values of a high-performance parallel filesystem–TB/s throughput, sub-ms latency at high client scale, and POSIX support–to a broader set of use cases and users.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Historically, due to the cost and special purpose nature of parallel filesystems, colder data had to be stored outside of the filesystem and AI developers have had to maintain separate, slower environments for writing code, compiling libraries, and managing repositories. This fragmentation increases the toil of manual data staging, dataset copying, and managing disjointed namespaces.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud Managed Lustre is solving these problems through our &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;6 cents/GB*month Dynamic Tier&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;and by optimizing Managed Lustre performance for a range of development tasks and workloads – making Managed Lustre a “One-Stop Shop” for high-performance AI and HPC workloads.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Lower Cost: More Lustre for Less with the Dynamic Tier&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Managed Lustre Dynamic Tier provides &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;sub-ms latency for hot data, which allows you to store all of your data in a single namespace, and costs only 6 cents/GB*month&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Throughput, capacity scale and client scale:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Throughput scales linearly with capacity up to 80 PB, while sub-ms latency for hot data remains stable as you scale to tens of thousands of clients.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Single-flat fee:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Predictable pricing. No independent charges for disk media types, data movement within the namespace, or metadata IOPS.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Read Latencies:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Sub-ms latencies for High-Performance Cache (SSD).  The Capacity Pool (“HDD”) is built on Google Cloud Hyperdisk throughput, which has an &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/disks/hd-types/hyperdisk-throughput"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;average read latency of 10 to 30 ms&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Recommended workloads for Dynamic Tier&lt;/strong&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-Epoch Training and/or Training with Optimized Fetch Sizes: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Hot data is promoted to the High Performance Cache (SSD) after the first run. Larger data prefetch will allow you to take advantage of the Dynamic Tier cost structure and gain from low-latency SSD.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Write-Heavy Checkpointing:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Bursty checkpoint writes land directly in the High Performance Cache. Older checkpoints are transparently demoted to the Capacity Pool (HDD).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Rapid Checkpoint Restore:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;  New checkpoints are written to the High Performance Cache, enabling low-latency checkpoint restores.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Interactive Snappiness for Developers:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Low-latency tasks like git cloning, compiling libraries, or running notebooks benefit from a local-disk feel (~300µs average read latencies) on the same shared workspace hosting large training sets.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Frictionless development: Lustre as a one-stop shop for developer’s workloads&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition to Managed Lustre’s scalability for large AI and HPC workloads (checkpoint/restart/data-loading), it also meets the demands for interactive work, meaning developers can start on Managed Lustre and stay on Managed Lustre throughout the entire workload lifecycle:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Unified Foundation &amp;amp; Interactive Performance&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Consolidates the AI and HPC lifecycle into a single namespace, providing a "local disk" feel for interactive work (Read more about the &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-lustre"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;latency benefits of Managed Lustre experienced by Salesforce and others&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Latency:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; ~300µs average read latency—delivering up to 4x better responsiveness than alternative distributed file systems.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerated Setup:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Untar the Linux kernel in ~2 minutes (4.7x faster than alternative file solutions), run a 20-worker parallel git clone of Python in ~40 seconds, compile Python in ~200s.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-f.max-1000x1000.jpg"
        
          alt="democratizing-lustre-with-lower-cost-and-frictionless-development-02-managed-lustre-performance-values"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-f.max-1000x1000_xLOfL6w.jpg"
        
          alt="democratizing-lustre-with-lower-cost-and-frictionless-development-01-performance-comparision"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;High-Concurrency Broadcast &amp;amp; Cluster Startup&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managed Lustre maximizes GPU ROI by preventing storage bottlenecks during cluster initialization. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When thousands of worker nodes attempt to read the exact same file simultaneously (such as a shared model checkpoint, base weights, or container layer), traditional distributed file systems can choke on localized hotspotting, leaving high-cost GPU clusters idle for minutes.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Improves Aggregate Throughput for a large number of clients reading the same file:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Demonstrates a 67% improvement over alternative file solutions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Parallel Loading:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Imports libraries like PyTorch across 4,000+ processes in under 60 seconds.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-f.max-1000x1000_9kShfM1.jpg"
        
          alt="democratizing-lustre-with-lower-cost-and-frictionless-development-03-performance-advantage"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Run One-Stop Shop Workflows for Yourself&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is the code for the tests we’ve run, so that you can perform your own testing.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Low latency for interactive access&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We used &lt;/span&gt;&lt;a href="https://github.com/axboe/fio" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;fio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to emulate small, low-concurrency reads and writes:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;sup&gt;&lt;span style="vertical-align: baseline;"&gt;1 &lt;/span&gt;&lt;/sup&gt;&lt;span style="color: #5f6368; font-size: 16px; font-style: italic; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Storage system specs: 500 MBps per TiB tier of Managed Lustre, 108,000 GiB capacity. Zonal Filestore at 102,400 GiB capacity. Average throughput of 36.7 GB/s to 2,048 client VMs reading the same 40 GiB file.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Read workload\r\nfio --ioengine=libaio --filesize=100M --ramp_time=2s \\\r\n    --runtime=2m --time_based --numjobs=1 --direct=1 --verify=0 --randrepeat=0 \\\r\n    --group_reporting --directory=~/LUSTRE_MOUNT \\\r\n    --name=randread --blocksize=4k --iodepth=1 --readwrite=randread \\\r\n    --buffer_compress_percentage=50\r\n\r\n# Write workload\r\nfio --ioengine=libaio --filesize=100M --ramp_time=2s \\\r\n    --runtime=2m --time_based --numjobs=1 --direct=1 --verify=0 --randrepeat=0 \\\r\n    --group_reporting --directory=~/LUSTRE_MOUNT \\\r\n    --name=randwrite --blocksize=4k --iodepth=1 --readwrite=randwrite \\\r\n    --buffer_compress_percentage=50&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f9f33d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerated setup&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;How to run Linux untar&lt;/strong&gt;&lt;/h4&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Download a kernel tarball\r\nwget -P /tmp https://cdn.kernel.org/pub/linux/kernel/v5.x/linux-5.18.9.tar.xz\r\n\r\n# Extract the archive to the Lustre mount\r\nmkdir ~/LUSTRE_MOUNT/kernel\r\ntar -C ~/LUSTRE_MOUNT/kernel -xf /tmp/linux-5.18.9.tar.xz&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f9f1ed0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the above use case, you will want to take care to avoid the metadata performance tax that can come from running as root (Namely, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tar&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; issues&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt; chown&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt; chmod&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; calls to make extracted files’ owner+permissions match the ones recorded in the archive.).  If you still wish to run as root (and have verified that this approach is compatible with your setup), you may &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;specify&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;`--no-same-owner --no-same-permissions`&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;in order to ensure that extracted files maintain&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; root&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; as owner and have root's default file permissions. In other words, it makes extraction as root behave like extraction as non-root (by ignoring the owner+permissions in the archive).&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;How to run Python gitclone&lt;/strong&gt;&lt;/h4&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;git config --global checkout.workers 20\r\nmkdir ~/LUSTRE_MOUNT/python\r\ngit clone https://github.com/python/cpython.git ~/LUSTRE_MOUNT/python&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f9f01d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;How to run Python compile&lt;/strong&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;pushd ~/LUSTRE_MOUNT/python\r\n./configure &amp;gt; /dev/null\r\nmake &amp;gt; /dev/null\r\npopd&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f9f0310&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;High scale distribution&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Aggregate throughput for distributing one large file to many nodes&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Run the below on each client VM:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Start fio in server mode\r\nfio --server&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f9f1890&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="iadne"&gt;Run the below on a selected client VM:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;# Create a 40 GiB file\r\nfio --name=job1 \\\r\n    --ioengine=libaio \\\r\n    --direct=1 \\\r\n    --buffer_compress_percentage=50 \\\r\n    --blocksize=4m \\\r\n    --iodepth=32 \\\r\n    --filesize=40g \\\r\n    --readwrite=write \\\r\n    --filename ~/LUSTRE_MOUNT/40gb_test\r\n\r\n# Create an fio job file for the read workload\r\ncat &amp;lt;&amp;lt;&amp;#x27;EOF&amp;#x27; &amp;gt; /tmp/read.fio\r\n[job1]\r\nfilename=${HOME}/LUSTRE_MOUNT/40gb_test\r\nrw=read\r\nbs=4m\r\nexitall_on_error=1\r\nEOF\r\n\r\n# Run the read workload using all client VMs in ~/hostfile\r\nfio --client ~/hostfile /tmp/read.fio&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43c3059d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Parallel loading of libraries across many processes&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Run the below on a selected client VM:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Install PyTorch in a virtual env\r\npython3 -m venv ~/LUSTRE_MOUNT/env\r\nsource ~/LUSTRE_MOUNT/env/bin/activate\r\npip3 install --upgrade pip\r\npip3 install torch torchvision torchaudio \r\ndeactivate\r\n\r\n# Import PyTorch on all client VMs in ~/hostfile, 4 processes per host\r\nmpirun --allow-run-as-root --oversubscribe --hostfile ~/hostfile -N 4 \\\r\n  bash -c \&amp;#x27;source ~/LUSTRE_MOUNT/env/bin/activate &amp;amp;&amp;amp; python3 -c &amp;quot;import torch&amp;quot;\&amp;#x27;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42fd3e210&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h1&gt;&lt;strong style="vertical-align: baseline;"&gt;Looking ahead and next steps&lt;/strong&gt;&lt;/h1&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By eliminating the manual data staging tax and lowering entry costs with the Dynamic Tier, Google Cloud Managed Lustre is evolving from an elite, single-purpose engine into a highly versatile, unified storage fabric for the entire AI lifecycle.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the second part of this series, we will focus on &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;upcoming object integration features&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Stay tuned!&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Next steps&lt;/strong&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Run the benchmarks yourself (if you haven’t already): Deploy a Google Cloud Managed Lustre instance using the &lt;/span&gt;&lt;a href="https://console.cloud.google.com/"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud console&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and run tests provided above to benchmark your own workloads.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explore the Dynamic Tier: Read the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-lustre/docs/performance-tiers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Managed Lustre Documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to learn more.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Stay tuned for Part 2: In the next installment of this series, we will dive deep into upcoming object integration features and how they further simplify AI and HPC storage.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Get started with centralizing your development-to-training lifecycle on Google Cloud Managed Lustre!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 13:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/democratizing-managed-lustre-with-lower-cost-and-frictionless-development/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-fri.max-600x600_y4nYtDt.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Democratizing Managed Lustre with lower cost and frictionless development</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-fri.max-600x600_y4nYtDt.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/democratizing-managed-lustre-with-lower-cost-and-frictionless-development/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Barak Epstein </name><title>Senior Product Manager, Google Cloud Managed Lustre</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Yuval Ehrental</name><title>Software Engineer, Google Cloud Managed Lustre</title><department></department><company></company></author></item><item><title>The future of browser-based security: Leveraging browser data for proactive defense</title><link>https://cloud.google.com/blog/products/chrome-enterprise/the-future-of-browser-based-security-leveraging-browser-data-for-proactive-defense/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="us4w0"&gt;The browser has changed significantly. Rather than just a window to the web, it serves as an AI workspace and central operating environment for the modern enterprise. With knowledge workers spending over 56% of their workday in the browser, it is a key gateway for daily work, complex workflows, and direct interaction with autonomous AI agents (&lt;a href="https://services.google.com/fh/files/misc/omdiareport.pdf" target="_blank"&gt;Omdia&lt;/a&gt;, 2026). To keep pace with threat actors, organizations need a browser strategy that places browser telemetry at the center of their security architecture.&lt;/p&gt;&lt;p data-block-key="dua8n"&gt;&lt;b&gt;The rise of shadow AI and agentic risk&lt;/b&gt;&lt;/p&gt;&lt;p data-block-key="872us"&gt;As enterprises adopt AI, new vulnerabilities have emerged. Shadow AI —the unauthorized use of public generative AI tools—can expose sensitive corporate IP through prompt sharing and autonomous agent actions. Ninety-two percent of organizations express concern around potential data leakage through these channels. Leaving this unmonitored creates significant risk (&lt;a href="https://services.google.com/fh/files/misc/omdiareport.pdf" target="_blank"&gt;Omdia&lt;/a&gt;, 2026).&lt;/p&gt;&lt;p data-block-key="11klt"&gt;Legacy security stacks, including traditional endpoint detection and response (EDR) and perimeter firewalls, are fundamentally blind to in-browser interactions. They completely miss high-risk threat vectors unique to AI-driven workflows, such as:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="aov7u"&gt;Malicious extensions that "read" sensitive financial data or "write" keyloggers onto sign-in pages.&lt;/li&gt;&lt;li data-block-key="43gpg"&gt;"Living off the land" (LOTL) tactics and session hijacking.&lt;/li&gt;&lt;li data-block-key="c3um2"&gt;State-sponsored threat actors leveraging AI to accelerate the attack lifecycle.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="e0cqe"&gt;&lt;b&gt;Chrome Enterprise Premium: Your high-fidelity telemetry engine&lt;/b&gt;&lt;/p&gt;&lt;p data-block-key="b4pap"&gt;Chrome Enterprise Premium addresses this visibility gap by capturing browser telemetry. Rather than relying on external observation after the fact, Chrome Enterprise records signals directly at the point of user interaction.&lt;/p&gt;&lt;p data-block-key="3khcv"&gt;&lt;b&gt;Core Capabilities for Modern Defense&lt;/b&gt;&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="5fp8m"&gt;&lt;b&gt;Real-time signals:&lt;/b&gt; Continuous telemetry for network events, high-risk user behaviors, and suspicious domain access.&lt;/li&gt;&lt;li data-block-key="bm027"&gt;&lt;b&gt;Extension telemetry:&lt;/b&gt; Granular visibility into side-loaded extensions and extension-to-domain communications that traditional EDR might miss.&lt;/li&gt;&lt;li data-block-key="46co0"&gt;&lt;b&gt;GenAI and SaaS app reporting:&lt;/b&gt; A dedicated capability to discover and govern sanctioned versus unsanctioned AI tools across the fleet.&lt;/li&gt;&lt;li data-block-key="1pv2n"&gt;&lt;b&gt;Evidence locker:&lt;/b&gt; The ability to capture files and content that violate DLP policies, providing a crucial trail for forensic analysis, root cause determination, and detection rule refinement.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="1omn6"&gt;&lt;b&gt;Transforming reactive review into proactive mitigation&lt;/b&gt;&lt;/p&gt;&lt;p data-block-key="uec"&gt;Transitioning to proactive defense allows security teams to mitigate threats early. By streaming browser signals into security operations platforms such as Google Security Operations, teams can automate responses and reduce manual investigation time, lowering incident response costs.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Blog_table_2000x1115_1.max-1000x1000.png"
        
          alt="Blog table_2000x1115 (1)"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="us4w0"&gt;&lt;b&gt;Insights from the frontlines: Mandiant case studies&lt;/b&gt;&lt;/p&gt;&lt;p data-block-key="3cbf6"&gt;Mandiant observations reveal the practical impact of browser visibility. These cases show how browser telemetry helps detect attacks that bypass standard controls:&lt;/p&gt;&lt;ol&gt;&lt;li data-block-key="duf5j"&gt;&lt;b&gt;RMM Software Download:&lt;/b&gt; Chrome Enterprise Premium flagged a legitimate Remote Monitoring and Management (RMM) executable because it originated from a newly registered domain—a key indicator of social engineering that network tools often miss.&lt;/li&gt;&lt;li data-block-key="1vaqi"&gt;&lt;b&gt;Credential Harvesting:&lt;/b&gt; Chrome Enterprise Premium evaluated URL risks and navigation parameters at the precise moment of interaction, disrupting a phishing attempt before the user could submit credentials.&lt;/li&gt;&lt;li data-block-key="dl743"&gt;&lt;b&gt;Malvertising:&lt;/b&gt; Integration with Google Threat Intelligence allowed for immediate identification of a malicious ad click, containing the threat before the actor gained hands-on-keyboard access.&lt;/li&gt;&lt;/ol&gt;&lt;p data-block-key="aju3a"&gt;Closing the security gap starts with recognizing that the browser is a key resource for enterprise security. With Google telemetry spanning billions of protected devices, organizations can maintain defensive visibility alongside emerging AI usage to drive a proactive security strategy (&lt;a href="https://safebrowsing.google.com/" target="_blank"&gt;Google Safe Browsing&lt;/a&gt;).&lt;/p&gt;&lt;p data-block-key="15cli"&gt;&lt;b&gt;Ready to transform your security posture?&lt;/b&gt;&lt;/p&gt;&lt;p data-block-key="fknf5"&gt;Avoid leaving a blind spot in your security strategy. Learn more about web defense options by reading our digital paper, &lt;b&gt;"&lt;/b&gt;&lt;a href="https://chromeenterprise.google/engage/strengthening-secops-with-browser-telemetry/" target="_blank"&gt;Securing the Browser: How telemetry brings web defense into the next frontier&lt;/a&gt;.&lt;b&gt;"&lt;/b&gt;&lt;/p&gt;&lt;p data-block-key="4bg5d"&gt;This resource outlines a framework for modern defense, detailing adversary tactics, infection vectors, and threat trends. Read the digital paper to evaluate your organization's browser security strategy.&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 09:02:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/chrome-enterprise/the-future-of-browser-based-security-leveraging-browser-data-for-proactive-defense/</guid><category>Chrome Enterprise</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Browser_Telemetry_Whitepapee_BlogHeader_2436.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>The future of browser-based security: Leveraging browser data for proactive defense</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Browser_Telemetry_Whitepapee_BlogHeader_2436.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/chrome-enterprise/the-future-of-browser-based-security-leveraging-browser-data-for-proactive-defense/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Niamh Cunningham</name><title>Senior Product Manager, Chrome Enterprise</title><department></department><company></company></author></item><item><title>Introducing the Server Side Cloud Swift SDK</title><link>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-server-side-cloud-swift-sdk/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For years, &lt;a href="https://swift.org/" rel="noopener nofollow noreferrer" target="_blank"&gt;Swift&lt;/a&gt; was perceived mainly as a UI language tied to Apple client devices. With Swift 6 and strict concurrency checking, it has matured into a viable systems and cloud language, pairing Rust-like data-race safety with predictable, reference-counted performance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To support this ecosystem, Google engineering has launched the official &lt;/span&gt;&lt;a href="https://github.com/googleapis/google-cloud-swift" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud API Client Libraries for Swift&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Built from the ground up for Swift 6.2+, this new SDK uses the latest non-blocking &lt;/span&gt;&lt;a href="https://github.com/apple/swift-nio" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Swift NIO&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; event loops, HTTP/2 multiplexing, &lt;/span&gt;&lt;a href="https://grpc.io" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gRPC&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; transport, and zero-cost compile-time data race safety.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this article, we'll walk you through all you need to know to get started, and to understand how the Server Side Cloud Swift SDK, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, works.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The rise of server-side Swift and cloud-native concurrency&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Traditional backend development often forces a trade-off between developer ergonomics and resource utilization. While managed runtimes offer rapid development, lower-level systems languages provide finer control over memory and CPU footprint, often at the expense of feature velocity.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Server-side Swift aims to strike a practical balance. Swift pairs a lightweight runtime and Automatic Reference Counting (ARC) with expressive syntax. More importantly, Swift 6 introduces compile-time concurrency checking. &lt;span style="vertical-align: baseline;"&gt;When you share state between async tasks across a cloud microservice, the compiler enforces that types conform to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Sendable&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. Data races are caught in your editor before a binary ever compiles or reaches production.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At the network layer, every request to &lt;/span&gt;&lt;a href="https://cloud.google.com"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; APIs runs over event-driven, non-blocking sockets that scale across multicore &lt;/span&gt;&lt;a href="https://www.kernel.org" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Linux&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; server environments without spawning system threads per connection.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/sdk-architecture.max-1000x1000.png"
        
          alt="sdk-architecture"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Where to use the Swift SDK&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Server Side Cloud Swift SDK is engineered for server, container, and automated DevOps environments.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When you build high-throughput microservices with Swift web frameworks like &lt;/span&gt;&lt;a href="https://hummingbird.codes" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Hummingbird&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;a href="https://vapor.codes" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Vapor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; provides native access to Cloud Storage, AI, Identity and Access Management (IAM), and over one hundred other Google Cloud services. You can containerize your executable on Linux and deploy directly to &lt;/span&gt;&lt;a href="https://cloud.google.com/run"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine (GKE)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://cloud.google.com/compute"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Compute Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; VMs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because the SDK compiles on macOS, and Linux, you can develop the backend in your preferred development environment, and then seamlessly deploy to production. And using Swift on both the frontend and backend allows you to share application-specific types across both.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The SDK also excels at platform engineering and DevOps automation. You can author cross-platform CLI utilities and data rotation scripts that run on your developer laptop or inside CI/CD pipelines. These tools authenticate automatically against Google Cloud using &lt;/span&gt;&lt;a href="https://cloud.google.com/docs/authentication/application-default-credentials"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Application Default Credentials (ADC)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;a href="https://cloud.google.com/iam/docs/workload-identity-federation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Workload Identity Federation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you're building an &lt;/span&gt;&lt;a href="https://developer.apple.com/ios/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;iOS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://www.apple.com/ipados/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;iPadOS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://www.apple.com/apple-vision-pro/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;visionOS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; app for the &lt;/span&gt;&lt;a href="https://www.apple.com/app-store/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Apple App Store&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, you should not embed &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; directly into your client bundle. Shipping Google Cloud service account keys or administrative credentials inside a client binary creates security risks. For direct client-side features, use the &lt;/span&gt;&lt;a href="https://github.com/firebase/firebase-ios-sdk" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Firebase SDK for Apple Platforms&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to handle user authentication, real-time &lt;/span&gt;&lt;a href="https://cloud.google.com/firestore"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Firestore&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; sync, and client-side security rules, or route requests through your own Cloud Run backend API.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Getting started with your IDE and packages&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; treats Linux and macOS as first-class citizens, you can develop on Apple hardware with &lt;/span&gt;&lt;a href="https://developer.apple.com/xcode/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Xcode&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or on Linux workstations with &lt;/span&gt;&lt;a href="https://code.visualstudio.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Visual Studio Code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://swift.org/install/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;swiftly&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To install the official Swift compiler on Linux workstations using the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;swiftly&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; CLI installer, run:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;curl -O https://download.swift.org/swiftly/linux/swiftly-$(uname -m).tar.gz &amp;amp;&amp;amp;\r\ntar zxf swiftly-$(uname -m).tar.gz &amp;amp;&amp;amp;\r\n./swiftly init --quiet-shell-followup &amp;amp;&amp;amp;\r\n. &amp;quot;${SWIFTLY_HOME_DIR:-$HOME/.local/share/swiftly}/env.sh&amp;quot; &amp;amp;&amp;amp;\r\nhash -r&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f494bd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Alternatively, you can download prebuilt toolchain tarballs directly from official &lt;/span&gt;&lt;a href="https://www.swift.org/download/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Swift Downloads&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for Ubuntu, Debian, Fedora, or Amazon Linux. Note that &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; requires &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Swift 6.2 or later&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, so verify your compiler version with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;swift --version&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; after installation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To resolve &lt;/span&gt;&lt;a href="https://swift.org/package-manager/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Swift Package Manager&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; bare repository trust warnings when cloning across Linux filesystems, configure &lt;/span&gt;&lt;a href="https://git-scm.com" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Git&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; before building with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;git config --global safe.bareRepository all&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Add the required packages to your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Package.swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; manifest:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;swift package add-dependency https://github.com/googleapis/swift-google-cloud-language-v2.git --from 0.4.0\r\nswift package add-target-dependency GoogleCloudLanguageV2 CloudBackendService --package swift-google-cloud-language-v2&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f497a90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;On macOS you need to change the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;platforms&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; directive:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;// swift-tools-version: 6.2\r\nimport PackageDescription\r\n\r\nlet package = Package(\r\n  name: &amp;quot;CloudBackendService&amp;quot;,\r\n  // Applied when compiling on Darwin/macOS; ignored by SPM on Linux targets\r\n  platforms: [.macOS(.v15)],\r\n ... ...&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f4940d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="bngnr"&gt;In most environments a default-initialized client can make requests:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import Foundation\r\nimport GoogleCloudLanguageV2\r\n\r\nfunc analyzeTextSentiment(text: String) async throws {\r\n  // Initialize explicit API key credentials\r\n  let client = try LanguageServiceClient()\r\n\r\n  // Configure request using structured builder closure\r\n  let document = Document().with {\r\n    $0.type = .plainText\r\n    $0.source = .content(text)\r\n  }\r\n\r\n  let response = try await client.analyzeSentiment(\r\n    request: AnalyzeSentimentRequest().with { $0.document = document }\r\n  )\r\n\r\n  if let sentiment = response.documentSentiment {\r\n    print(&amp;quot;Document sentiment score: \\(sentiment.score)&amp;quot;)\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f497e90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Notice how &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Document().with { ... }&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; avoids verbose temporary variables or mutating setters by providing a clean, thread-safe configuration closure.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Networking, transport, and authentication&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The repository splits infrastructure primitives into modular packages under &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;packages/&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;swift-google-cloud-auth&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Implements &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/docs/authentication/application-default-credentials"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Application Default Credentials&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; discovery, service account JWT signing, external account exchange for Workload Identity Federation, and API keys.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;swift-google-cloud-wkt&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Provides idiomatic Swift types for Google Protocol Buffer well-known types, including nanosecond-precision &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Timestamp&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; representations that bridge cleanly to Swift's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Date&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;swift-google-cloud-gax&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Handles Google API Extensions such as automated retry loops, exponential backoff, and pagination state machines.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When you initialize any client library without arguments, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Credentials.default()&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; automatically scans your environment (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GOOGLE_APPLICATION_CREDENTIALS&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, quota project variables, or the local Google Cloud CLI configuration) and authenticates connections over gRPC or HTTP/2.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you need to programmatically override credentials with an API key or attach custom access headers, you can pass explicit configuration options. For example, you could modify the previous example to use the following:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import Foundation\r\nimport GoogleCloudAuth\r\nimport GoogleCloudGax\r\nimport GoogleCloudLanguageV2\r\n\r\nfunc analyzeTextSentiment(apiKey: String, text: String) async throws {\r\n  // Initialize explicit API key credentials\r\n  let credentials = try Credentials(configuration: .apiKey(apiKey))\r\n  let client = try LanguageServiceClient(\r\n    ClientOptions().with { $0.credentials = credentials }\r\n  )\r\n\r\n  // Configure request using structured builder closure\r\n  let document = Document().with {\r\n    $0.type = .plainText\r\n    $0.source = .content(text)\r\n  }\r\n\r\n  let response = try await client.analyzeSentiment(\r\n    request: AnalyzeSentimentRequest().with { $0.document = document }\r\n  )\r\n\r\n  if let sentiment = response.documentSentiment {\r\n    print(&amp;quot;Document sentiment score: \\(sentiment.score)&amp;quot;)\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43c24bcd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The autogenerated client ecosystem&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud operates a vast ecosystem of APIs whose schemas update regularly. The teams supporting &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; use code generators to automatically update the client libraries with the latest features and with new APIs. Using code generators produces stable APIs, without disruptive breaking changes. While the releases are on a fixed cadence, please contact Cloud Customer Care if you need a particular feature or API urgently.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether you need to rotate keys in &lt;/span&gt;&lt;a href="https://cloud.google.com/secret-manager"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Secret Manager&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or invoke multimodal inference models via the &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/gemini-v1"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini API&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, the generated SDKs follow consistent naming and async/await signatures.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These generated clients offer more than plain unary RPC wrappers. They also offer wrappers that simplify application development. For example, iterating over long results involves fetching pages of results with one RPC, iterating over the page of results, and then preparing a new request to retrieve the following page. Using the generated clients this becomes an asynchronous iterator. This example shows how to query project secrets using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GoogleCloudSecretManagerV1&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import Foundation\r\nimport GoogleCloudSecretManagerV1\r\n\r\n@main\r\nstruct SecretManagerQuickstart {\r\n  static func main() async throws {\r\n    guard let projectId = CommandLine.arguments.dropFirst().first else {\r\n      print(&amp;quot;Usage: SecretManagerQuickstart &amp;lt;projectId&amp;gt;&amp;quot;)\r\n      exit(1)\r\n    }\r\n\r\n    // Connects using Application Default Credentials automatically\r\n    let client = try SecretManagerServiceClient()\r\n\r\n    let request = ListSecretsRequest().with {\r\n      $0.parent = &amp;quot;projects/\\(projectId)&amp;quot;\r\n    }\r\n\r\n    // Async sequence streams pages of secrets automatically\r\n    print(&amp;quot;Secrets in project \\(projectId):&amp;quot;)\r\n    for try await item in try client.listSecretsByItem(request: request) {\r\n      print(&amp;quot; - \\(item.name)&amp;quot;)\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43c249490&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The pagination response returns an asynchronous sequence (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;AsyncSequence&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). You can iterate over items with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;for try await&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; while the client library fetches subsequent pages in the background over non-blocking NIO channels.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Where to go next&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, server-side Swift developers can write end-to-end cloud infrastructure with compile-time race safety, native async/await ergonomic APIs, and zero OS thread congestion.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To inspect the source code, open issues, or contribute new veneers, visit the official repository at &lt;/span&gt;&lt;a href="https://github.com/googleapis/google-cloud-swift" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;googleapis/google-cloud-swift&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Want to discuss server-side Swift architectures or Cloud Run containerization? Join the &lt;/span&gt;&lt;a href="https://developers.google.com/program" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Developer Program&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to continue the conversation.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 04:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-server-side-cloud-swift-sdk/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/introducing-the-server-side-cloud-swift-sdk-.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing the Server Side Cloud Swift SDK</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/introducing-the-server-side-cloud-swift-sdk-.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-server-side-cloud-swift-sdk/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Karl Weinmeister</name><title>Director, Developer Relations</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Carlos O'Ryan</name><title>Software Engineer</title><department></department><company></company></author></item><item><title>What’s new in AI infrastructure and orchestration in September</title><link>https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We hereby declare September to be scalability month! As the world prepares for a surge of agentic fleets, we are shoring up our AI infrastructure and orchestration offerings to gracefully — and quickly — respond to that demand, all while maintaining workload isolation and security, and keeping costs in check. Read on to learn how these enhancements manifest across Google Cloud’s compute, network, storage, and orchestration offerings, plus new ways customers are using Google Cloud AI infrastructure, and third-party industry validation of our strategy. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology, and tools updates&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Kubernetes Engine updates:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The GKE team is all about improving the scalability of the platform, and in September, those improvements came in many shapes and sizes:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Need an execution runtime with higher density for your agentic workloads? We engineered the new open-source &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/agent-substrate-available-on-gke?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Agent Substrate&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to run millions of sandboxes with 10x higher density than standard container runtimes. Agent Substrate also delivers sub-500ms resume operations at over 500 suspend/resume activations per second with a native zero-trust kernel and network isolation. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; GKE now has &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-adds-native-scale-to-zero-capabilities?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;scale-to-zero&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; capabilities built-in. No need to configure complex components to scale your workloads down, thanks to the HPA with the Autoscaling Metric and support for KEP-2021, which do the job for you, out of the box. &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-adds-native-scale-to-zero-capabilities?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Read the blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to learn more. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Further, the GKE HPA (with the above-mentioned Autoscaling Metric) now lets you scale up and down based on custom PromQL metrics, in addition to standard metrics, allowing you to trigger workloads according to conditions that are meaningful and unique to your business. Read more &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/native-support-for-prometheus-metrics-in-gke?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Yet another scalability feature is &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/pod-snapshots"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Pod snapshots&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which lets you save the running state of your workload, including CPU and GPU memory, and restore it on demand. According to internal tests, GKE Pod snapshots can reduce AI inference start-up by as much as 89%. Learn more &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-pod-snapshots"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New migration tool: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Finally, if you’ve always wanted to migrate your container workloads from AWS EKS to GKE but feared a daunting, high-friction engineering endeavor, we’ve just launched &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-agentic-migration?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE agentic migration&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a purpose-built agent plugin that replaces brittle, ad-hoc prompting with an AI-assisted migration pipeline protected by deterministic guardrails. Designed as a compilation of agent skills and a local Model Context Protocol (MCP) server, it uses AI to translate complex AWS EKS IaC and Kubernetes manifests directly into GKE landing zones. Get started with the &lt;/span&gt;&lt;a href="https://github.com/gke-labs/gke-agentic-migration/blob/main/docs/onboarding-guide.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;onboarding guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Feature updates: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Reinforcement learning (RL) and evaluation workloads are a beast: In a standard agentic RL loop, an LLM policy generates actions like code snippets on GPUs and executes them inside isolated CPU sandboxes to observe a reward signal. However, when scaling up this loop to support tens of thousands of parallel rollouts, infrastructure bottlenecks emerge, for instance idle accelerators, image cardinality, and a saturated control plane. To help, we developed GKE Agent Sandbox optimized for RL, plus an Agent Sandbox RL orchestration SDK and native integrations for popular RL gyms and harnesses. All are now generally available, and you can learn more &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/accelerate-agentic-rl-with-gke-agent-sandbox?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Storage updates: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;AI trains and creates lots of data, and that data has to live somewhere — in block storage systems, file systems, object stores and databases. We announced enhancements to our storage portfolio to help this critical layer roll with the agentic punches:  &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/filestore-agent-volumes?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Filestore agent volumes&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; offer high-performance, elastic, persistent file storage for agentic workloads. Thanks to its tight integration with GKE Agent Substrate and GKE Agent Sandbox, Filestore agent volumes automatically allocates and attaches a dedicated, isolated file workspace to GKE agent sandboxes in milliseconds. Request access to the preview &lt;/span&gt;&lt;a href="http://forms.gle/vYPkcFiZVoTjf7Ah7" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New product:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If you run generative AI and RAG data layers  — think Milvus, Pinecone, Qdrant, Vespa, Redis, and in-memory context caching — you may want to take a look at the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/compute-engine-m4n-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;M4N family of VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, now GA, which offers the highest per-core IOPS and throughput of leading hyperscalers. Paired with Google Cloud's custom &lt;/span&gt;&lt;a href="https://cloud.google.com/titanium"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Titanium&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; offload architecture and paired with&lt;/span&gt;&lt;a href="https://cloud.google.com/compute/docs/disks/hyperdisks"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; Hyperdisk Extreme&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, M4N instances deliver up to 25,000 MiB/s (25 GiB/s) of aggregate host storage performance and up to 1 million IOPS. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New product:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Another new Compute Engine product, &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/storage-optimized-z4d-vm-and-bare-metal-instances"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Z4D, is GA&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and a strong storage solution for AI/ML training and inference workloads. When configured as a bare metal instance, Z4D provides both the high local SSD (LSSD) capacity and low latency required by agentic microVMs, so you can run thousands of isolated sandboxes per host with native performance and efficiency. Learn about &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/storage-optimized-z4d-vm-and-bare-metal-instances?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Z4D machines here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Today’s AI training and inference pipelines create data faster than most storage management systems can keep up, creating challenges for teams trying to understand their storage estates. A new version of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/storage/docs/storage-intelligence/advisor-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Storage Intelligence advisor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; makes it easier to answer the question: "What’s in my buckets?" and quickly identify unexpected changes. Then, enhanced &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/storage/docs/batch-operations/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;batch operations&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; let you automate bulk changes across your buckets — say, move storage classes, mass-delete stale or temporary data, or apply metadata, tagging, retention, or encryption changes. Learn about the latest in Storage Intelligence advisor &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/storage-intelligence-advisor-and-batch-operations-updates/"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New product:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Last but not least, a new version of &lt;/span&gt;&lt;a href="https://cloud.google.com/products/alloydb"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB for PostgreSQL&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; brings together pioneering Google infrastructure — Colossus distributed file system, and Jupiter network — to power a &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/alloydbs-agentic-database-architecture?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;new, no-compromises database architecture for the agentic era&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Wish you could make GPUs and TPUs scattered around the globe behave as a single pool behind a single entry point? &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gpu-and-tpu-utilization-with-multi-cluster-gke-inference-gateway?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;In this blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, we show you how to do just that. At the edge, the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/setup-multicluster-inference-gateway"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;multi-cluster GKE Inference Gateway&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; focuses on global, multi-region traffic distribution and high availability. Beneath that, the &lt;/span&gt;&lt;a href="https://github.com/llm-d/llm-d-router" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LLM-d router&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; handles complex, memory-aware scheduling algorithms to keep utilization high. This architecture is deliberately runtime-, model-, and accelerator-agnostic, and in tests, routing traffic through the multi-cluster GKE Inference Gateway added less than 1% overhead.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Using or planning to use GKE on TPUs for AI model training or inference? Training massive Large Language Models (LLMs) or running high-throughput inference serving represents a significant investment in specialized AI hardware, such as Cloud TPUs and GPUs. To get the most out of every dollar spent, you need to &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/a-tale-of-tpu-observability-on-gke-part-1/397955" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;understand workload lifecycle metrics in GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. This detailed guide explains how to turn opaque cluster behaviors into actionable telemetry.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GKE customer:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn why gaming startup SeaVerse relies on &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/seaverse-chooses-gke-agent-sandbox?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Agent Sandbox for its multi-tenant workloads&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and how the platform helped it decrease its infrastructure costs by 60%. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep-dives&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In case you missed it, we’re also thrilled to share that Google has been named a leader, including achieving the highest score on either product or strategy, in three key analyst reports from Gartner and Forrester. &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;a href="https://cloud.google.com/blog/products/compute/forrester-wave-public-cloud-platforms-q3-2026-report?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google is a leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Highest overall score of any cloud provider!&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;a href="https://cloud.google.com/blog/products/compute/google-named-a-leader-in-2026-gartner-magic-quadrant-for-scps?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google named a Leader in 2026 Gartner® Magic Quadrant™ for Strategic Cloud Platform Services&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Positioned furthest for “Completeness of Vision” of all vendors evaluated.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/2026-gartner-magic-quadrant-for-container-management?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google is a Leader in the 2026 Gartner Magic Quadrant for Container Management&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Positioned highest in “Ability to Execute” of all vendors evaluated.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud achieved a Gold rating in the latest SemiAnalysis ClusterMax 3.0 report, which evaluates the reliability, performance, support, pricing, and security of GPU providers globally. See the&lt;/span&gt;&lt;a href="https://newsletter.semianalysis.com/p/clustermax-30-the-industry-standard" rel="noopener" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;full report for more&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;August 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology, and tools updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/filestore"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Filestore&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/how-colossus-optimizes-data-placement-for-performance?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Colossus&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/filestore-file-service-runs-on-colossus?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog post&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature: &lt;/strong&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gvisor-sandboxes-for-ray-clusters-on-gke?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gVisor sandboxes are now available in distributed Ray clusters on GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the &lt;/span&gt;&lt;a href="https://docs.ray.io/en/master/cluster/kubernetes/examples/ray-sandboxing.html" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Ray sandboxing User Guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/serverless/introducing-cloud-run-instances"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run instances&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70.  &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides, documentation and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in &lt;/span&gt;&lt;a href="https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this Google Developers blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.  &lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Real-time AI systems make a mess of traditional network load balancing techniques.&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.”&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; Things only get worse when the user gets involved. &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.”&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; For a new approach to managing load in the AI era, read &lt;/span&gt;&lt;a href="https://developers.googleblog.com/scaling-real-time-ai-agents-with-session-aware-load-balancing/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Scaling real-time AI agents with session-aware load balancing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/how-to-build-an-elastic-scalable-llm-inference-platform-on-gke-using-fluid-compute/388108" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Documentation: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/perform-host-maintenance-accelerators"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;update accelerator-equipped hosts&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; according to your tolerance for downtime for your training and inference workloads.   &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Documentation: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/instances/use-aci-images"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;create an ACI image&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; using the Google Cloud CLI, console, or SchedMD's Slurm workload manager&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;. &lt;/strong&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/ai-infrastructure/best-practices-for-dynamic-capacity-management?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;three main ways to achieve dynamic capacity management in Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. &lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Business orchestration software provider &lt;/span&gt;&lt;a href="https://www.uipath.com/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;UiPath&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/customers/how-uipath-built-its-high-performance-gpu-platform"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://mirendil.com/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Mirendil&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, an frontier AI lab focused on accelerating AI development, announced that it is &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/startups/mirendil-selects-ai-hypercomputer?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;using AI Hypercomputer&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://replen.it/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Replenit&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/replenit?e=48754805&amp;amp;hl=en"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;full case study&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for more. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.malachyte.com/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Malachyte&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/solving-retails-cold-start-problem-malachytes-recommendation-reinvention?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;July 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology, and tools updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-lustre"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Managed Lustre&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/c4n-network-and-storage-optimized-vms?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;C4N network and storage optimized VMs are now GA&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's &lt;/span&gt;&lt;a href="https://cloud.google.com/titanium?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Titanium&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/planning-large-clusters#clusters-5k-nodes"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Dataplane V2 up to 15K Nodes with Network Policies (GA)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/introducing-co-operative-time-slicing-for-rl-in-llm-d?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Co-operative time-slicing in llm-d&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New AI security tool:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/introducing-k8s-aibom-on-gke-for-automated-ai-bills-of-materials?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Looking to secure your AI supply chain on GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/k8s-aibom" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;k8s-aibom project&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and get involved.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;On July 27, Google announced &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/announcing-day-0-support-for-kimi-k3-on-google-cloud/385392" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Day 0 support for Moonshot AI’s Kimi K3&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/developers-practitioners/autopilot-clusters-with-gke-managed-dranet-gpus-and-tpus"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/codelabs/gke-autopilot-tpus-dranet-gemma#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn to run Ray on TPUs, not GPUs. In &lt;/span&gt;&lt;a href="https://developers.googleblog.com/run-ray-on-tpu-part-1-the-foundations/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Part 1&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (&lt;/span&gt;&lt;a href="https://developers.googleblog.com/run-ray-on-tpu-part-2-ray-ai-libraries/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Part 2&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in &lt;/span&gt;&lt;a href="https://developers.googleblog.com/how-to-use-google-microbenchmarks-for-evaluating-tpu-performance/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Scale your agents without killing your budget. &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/reduce-your-agents-costs-with-gke-agent-sandbox?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Technical blueprint: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/inside-the-optimization-of-mistral-3-large-inference-on-ironwood/385847" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep-dives&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Report: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Google was named a Leader in the inaugural &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/ai-infrastructure/google-is-a-leader-in-gartner-magic-quadrant-for-ai-infra?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gartner&lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;&lt;span style="vertical-align: super;"&gt;Ⓡ&lt;/span&gt;&lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; Magic Quadrant™ for AI Infrastructure&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/content/2026-gartner-mq-ai-infrastructure?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Report:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; We recently surveyed more than 1,400 senior IT leaders for our &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/content/state-of-infrastructure-in-the-agentic-ai-era?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;State of AI Infrastructure report&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/state-of-ai-infrastructure-report-overview?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Read the accompanying blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;June 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology and tool updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Protecting sensitive data used with AI is a critical part of advanced and secure cloud infrastructure. &lt;/span&gt;&lt;a href="https://cloud.google.com/security/products/confidential-computing?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Confidential Computing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; cryptographically protects data in use in hardware-based Trusted Execution Environments (TEEs) with verifiable data integrity, and is &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/verifiable-trust-in-the-ai-era-whats-new-in-confidential-computing?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;now available&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on the accelerator-optimized &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/accelerator-optimized-machines#g4-series"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;G4 machine series&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, featuring &lt;/span&gt;&lt;a href="https://www.nvidia.com/en-us/products/workstations/professional-desktop-gpus/rtx-pro-6000-family/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Get started with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/confidential-computing/confidential-vm/docs/create-a-confidential-vm-instance-with-gpu"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Confidential G4 VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/gpus-confidential-nodes"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Confidential G4 GKE Nodes&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Developer resource: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The new &lt;/span&gt;&lt;a href="https://cloud.google.com/products/tpu/tpu-developer?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;TPU Developer Hub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is the place to go for model builders, optimizers, and developers to learn to unlock the full performance of Google Cloud TPUs. Read more in this &lt;/span&gt;&lt;a href="https://developers.googleblog.com/unlocking-the-power-of-the-tpu-stack-introducing-our-new-developer-hub/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New product: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Scale your AI workloads with the new &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/stop-training-blind-scaling-ai-with-the-new-opentelemetry-based-tpu-ai-telemetry-collector-agent/375210" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OpenTelemetry-Based TPU AI Telemetry Collector Agent&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. For the first time, you can route high-fidelity TPU hardware telemetry to Google Cloud Monitoring, Google Managed Prometheus, or your own self-hosted Grafana stack.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Practitioner guides and how-tos&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/developers-practitioners/experimenting-with-tpus-gke-managed-dranet-and-multi-cluster-inference-gateway?_gl=1*jj3plw*_ga*OTAxNzc0MzU1LjE3ODIyMjAxNDk.*_ga_4LYFWVHBEB*czE3ODI3NTc3NzAkbzkkZzEkdDE3ODI3NTg2MDEkajYwJGwwJGgw&amp;amp;e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provides an overview, or you can get all the technical details in the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/codelabs/gke-inference-gateway-multi-cluster-tpus-dranet#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;hands-on codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;How-to guide:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Did you know you can connect your AI agents to unstructured data in &lt;/span&gt;&lt;a href="https://cloud.google.com/storage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; via Model Context Protocol (MCP)? In &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/developers-practitioners/build-ai-agents-faster-with-gcs-google-cloud-storage-mcp-server"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep-dives&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Report: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;According to an independent benchmark report, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/about-gke-inference-gateway"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Inference Gateway&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/containers-kubernetes/gke-inference-gateway-prefix-caching-accelerates-ai-inference?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecture deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;A closer look at &lt;/span&gt;&lt;a href="https://discuss.google.dev/t/accelerate-tpu-model-loading-while-saving-ram-on-gke/374835" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;the cold start problem, this time for TPUs and GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and how the Run:ai Model Streamer can help change the dynamic. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Customer win:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Leveraging GKE, BigQuery, Cloud SQL, and Gemini Enterprise Agent Platform, &lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=x36QJ-QKRGg" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Pager Health is eliminating operational fragmentation to deliver a simplified, personalized U.S. healthcare experience&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; that transforms lives.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Customer win:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Trustpilot, the customer review platform, built a high-volume streaming pipeline using fine-tuned Gemma models with Dataflow and Gemini Enterprise Agent Platform running on cost-optimized A2 VMs using A100 GPUs, as well as optimized version of vLLM maintained by Gemini Enterprise Agent Platform.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;May 2026&lt;/span&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Product, technology and tool updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product update:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/machine-learning/agent-sandbox"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE Agent Sandbox&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is now generally available.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New open-source project:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://github.com/agent-substrate/substrate" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Substrate&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is a new open-source project aimed at continuing to push the limits of agentic infrastructure density&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New feature:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://ai.google.dev/edge/ai-edge-portal" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google AI Edge Portal&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/benchmark-llms-on-device-with-ai-edge-portal?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Product deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We went &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/storage-data-transfer/cloud-storage-rapid-turbocharges-object-storage-for-ai-analytics?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;into depth about Cloud Storage Rapid&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Research, reports and deep dives&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecture deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Google Global Infrastructure VP Bikash Koley and Engineering Fellow Arjun Singh provide a high-level overview of &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/networking/data-center-and-global-networks-built-for-ai-era"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;the challenges that AI workloads pose to network infrastructure&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and discuss the deep enhancements we’ve made to our data center fabrics, WAN, and global networks to better support them. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Architecture deep dive: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We unveiled a &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/compute/cluster-reliability-for-trillion-parameter-models-on-tpus?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;new cluster-level reliability model&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for developing frontier AI models on TPUs, ditching instance-level reliability &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Customer and partner updates&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Customer win:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Visual media provider &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/infrastructure/how-imgix-processes-8-billion-images-daily-with-g4-vms-powered-by-nvidia-blackwell?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Imgix serves more than 8 billion images and videos from AI Hypercomputer&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; equipped with G4 VMs powered by NVIDIA RTX PRO 6000 Blackwell GPUs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 30 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/</guid><category>AI &amp; Machine Learning</category><category>Containers &amp; Kubernetes</category><category>Compute</category><category>Networking</category><category>Storage &amp; Data Transfer</category><category>AI infrastructure</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Whats_new_in_AI_infrastructure.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>What’s new in AI infrastructure and orchestration in September</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Whats_new_in_AI_infrastructure.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/ai-infrastructure/whats-new-in-ai-infrastructure-this-month/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Alex Barrett</name><title>Editor, Google Cloud blog</title><department></department><company></company></author></item><item><title>Cloud CISO Perspectives: How cybersecurity startups can win CISOs</title><link>https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-how-cybersecurity-startups-can-win-cisos/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="eucpw"&gt;Welcome to the second Cloud CISO Perspectives for September 2026. Today, Alicja Cade and Nick Godfrey, senior directors, Office of the CISO, share their guidance for cybersecurity startups on how to win over the CISOs who will become crucial business partners and customers.&lt;/p&gt;&lt;p data-block-key="99ga0"&gt;As with all Cloud CISO Perspectives, the contents of this newsletter are posted to the &lt;a href="https://cloud.google.com/blog/products/identity-security/"&gt;Google Cloud blog&lt;/a&gt;. If you’re reading this on the website and you’d like to receive the email version, you can &lt;a href="https://cloud.google.com/resources/google-cloud-ciso-newsletter-signup"&gt;subscribe here&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Get vital board insights with Google Cloud&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc42f8cb890&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;Visit the hub&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;https://cloud.google.com/solutions/security/board-of-directors?utm_source=cgc-site&amp;amp;utm_medium=et&amp;amp;utm_campaign=FY26-Q2-GLOBAL-GCP39634-email-dl-dgcsm-CISOP-NL-177159&amp;amp;utm_content=-&amp;amp;utm_term=-&amp;#x27;), (&amp;#x27;image&amp;#x27;, &amp;lt;GAEImage: GCAT-replacement-logo-A&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="hswvv"&gt;&lt;b&gt;How cybersecurity startups can win CISOs&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="998n5"&gt;&lt;i&gt;By Alicja Cade, Senior Director, Financial Services, Office of the CISO, and Nick Godfrey, Senior Director, Office of the CISO&lt;/i&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Alicja_Cade_headshot_2.max-1000x1000.jpg"
        
          alt="Alicja Cade headshot 2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="nj7d4"&gt;Alicja Cade, Senior Director, Financial Services, Office of the CISO&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="0jyqm"&gt;Cybersecurity startups play a crucial role in technology as they build to solve both legacy, existential challenges and the latest problems on the cutting edge. A key part of winning and transforming the cybersecurity field is becoming a strategic partner to CISOs and their security teams.&lt;/p&gt;&lt;p data-block-key="em24t"&gt;Google has supported more than 50 cybersecurity founders over the past four years through our &lt;a href="https://startup.google.com/"&gt;Google for Startups program&lt;/a&gt;, including &lt;a href="https://authologic.com/"&gt;Authologic&lt;/a&gt;, &lt;a href="http://www.bfore.ai/"&gt;BforeAI&lt;/a&gt;, &lt;a href="https://www.build38.com/"&gt;Build38&lt;/a&gt;, &lt;a href="http://cerby.com/"&gt;Cerby&lt;/a&gt;, &lt;a href="https://www.crowdsec.net/"&gt;Crowdsec&lt;/a&gt;, &lt;a href="http://www.riskledger.com/"&gt;Risk Ledger&lt;/a&gt;, and &lt;a href="https://www.mokn.io/"&gt;Mokn&lt;/a&gt;.&lt;/p&gt;&lt;p data-block-key="462ja"&gt;Christian Torres, co-founder and CEO, &lt;a href="https://www.kriptos.io/"&gt;Kriptos&lt;/a&gt;, and a Google for Startups participant, said that building connections between startups and CISOs is crucial to solving critical security challenges.&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/NickGodfrey8975-hi_Tm5UVy8.max-1000x1000.jpg"
        
          alt="NickGodfrey8975-hi"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="sll9w"&gt;Nick Godfrey, Senior Director, Office of the CISO&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="c2abj"&gt;"The Google for Startups program has been the most impactful initiative we've joined as a cybersecurity company. Unlike other accelerator programs, this one speaks our language — the challenges, the ecosystem, and the conversations are 100% aligned with what we do every day at Kriptos. The access to CISOs and security leaders has been invaluable, and the connections we've built through the program are ones we now see regularly across industry events. It's put us exactly where we need to be,” he said.&lt;/p&gt;&lt;p data-block-key="dcpej"&gt;Cybersecurity startup founders face many competing taskmasters as they fight for survival, from demanding capital funders to the relentless pressure of growing their market and networks. CISOs should be key stakeholders for cybersecurity startups so that founders focus on solving thorny challenges in a way that works in the real world.&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-pull_quote"&gt;&lt;div class="uni-pull-quote h-c-page"&gt;
  &lt;section class="h-c-grid"&gt;
    &lt;div class="uni-pull-quote__wrapper h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;
      &lt;div class="uni-pull-quote__inner-wrapper h-c-copy h-c-copy"&gt;
        &lt;q class="uni-pull-quote__text"&gt;Listening to CISOs and understanding the businesses that they serve takes time and effort, and if done right can help deliver better value and create a lasting enterprise foundation and network of allies.&lt;/q&gt;

        
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/section&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here are three top tips from September’s &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/identity-security/meet-the-33-cybersecurity-startups-joining-the-gemini-startup-forum"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Startup Forum for Cybersecurity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, part of the Google for Startups program, where we offered vital guidance, addressed critical domains, and helped foster deep dialogue for the next generation of AI-native cybersecurity startups.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Tip 1: Listen then design and deliver for your customers&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Avoid becoming a round peg in a square hole by combining your problem-solving startup with listening to CISOs who have to protect real systems, networks, and people. Listening to CISOs and understanding the businesses that they serve takes time and effort, and if done right can help deliver better value and create a lasting enterprise foundation and network of allies.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here’s how to develop trusted CISO relationships:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Host diagnostics meetings&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Your meetings with CISOs should focus on mapping their operational bottlenecks and co-authoring collaborative solutions while studying their pain points.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Create "unselling" spaces to build peer trust&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Host intimate, pitch-free roundtable discussions on industry challenges or establish a critique-only advisory board to build genuine relationships with CISOs without the pressure of a sales environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Use neutral networks that don’t include venture capitalists&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Engage with CISOs in low-friction environments by contributing to open-source security projects and participating in academic and geopolitical risk forums where security leaders gather to solve broad industry problems.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Avoid the bait-and-switch pitch&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Never disguise a sales pitch as a research or feedback session, as tricking a CISO into a product demo will permanently destroy their trust.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Center their business context&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Don’t limit your listening to the technical security stack, because the CISO’s primary job is to enable and protect the broader business strategy.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Tip 2: Evaluate AI security to filter out noise&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of just using AI to assemble the product, startups should critically evaluate what makes your approach unique and how you communicate that to potential customers.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Define your moat&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; by investing in proprietary datasets, specialized fine-tuning, and unique orchestration layers that create a true technical moat. Don’t be a wrapper.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Secure the intelligence&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; by proactively designing your models to resist adversarial attacks, prompt injection, and data poisoning. In cybersecurity, model robustness is your ultimate trust signal.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Deliver high-fidelity outcomes&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; by clearly communicating how your AI product reduces cognitive load for defenders, minimizes false positives, and integrates safely into existing operations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Avoid using generic marketing buzzwords like "cognitive," "autonomous," or "revolutionary" without the technical documentation, case studies, and whitepapers to back them up. In a skeptical market, transparency is your best sales tool. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Tip 3: Enthusiastically embrace your sector&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Keep a sharp eye out for common due diligence pitfalls during investment and merger and acquisition cycles. These include ensuring that internal engineering and cybersecurity practices meet external claims, but also evaluating the regulatory context of your business sector as well as the security and reliability of your product and service. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You have to know whether you’re required to abide by data sovereignty, data residency, and other requirements. To avoid this pitfall, engage with broader stakeholders early who know the sector and its nuances well.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;How to keep the conversation going&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Even beyond the crowded field of aspiring cybersecurity companies, startups broadly can benefit immensely by making sure that they listen carefully, evaluate objectively, and take to their sector requirements enthusiastically. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;"Google's Office of the CISO has been a bridge between LetsData and the security leaders we need to reach. Sometimes that bridge is advice on how our offering maps to a CISO's real priorities. Sometimes it is a direct introduction to a CISO who is looking for exactly what we build. For a startup, a warm introduction at that level is priceless,” said Ksenia Iliuk, founder and COO, &lt;/span&gt;&lt;a href="https://letsdata.net/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LetsData&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To learn more about how Google Cloud’s Office of the CISO can help support your organization, check out our &lt;/span&gt;&lt;a href="https://cloud.google.com/solutions/security/leaders"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;CISO Insights hub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Learn something new&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cc9e1d0&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;Watch now&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;https://www.youtube.com/watch?v=Wpo-5ke9uvQ&amp;#x27;), (&amp;#x27;image&amp;#x27;, &amp;lt;GAEImage: Cloud-CISO-Perspectives-logo-A&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="4bd61"&gt;&lt;b&gt;In case you missed it&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="72e5r"&gt;Here are the latest updates, products, services, and resources from our security teams so far this month:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="e425v"&gt;&lt;b&gt;Agentic hacks, real proofs: Inside Google's PageBreak project&lt;/b&gt;: Distinguishing a genuine, exploitable flaw from a convincing hallucination has become a major challenge, often increasing the burden on product teams. PageBreak is an internal AI agent of Google's Product Security team developed to test the security of our first-party web applications and address this challenge. &lt;a href="https://blog.google/security/agentic-hacks-real-proofs-inside-googles-pagebreak-project/" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="45qfl"&gt;&lt;b&gt;Investing together: Wiz Defend and Google Security Operations&lt;/b&gt;: Continuing to deepen the integration between Wiz Defend and Google Security Operations, helping teams work faster wherever they choose to investigate. &lt;a href="https://www.wiz.io/blog/wiz-defend-and-google-security-operations" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="3ekrf"&gt;&lt;b&gt;A unified view of Android security updates for enterprises and OEMs&lt;/b&gt;: We're introducing new libraries that give enterprise partners and OEMs a complete, real-time picture of a device’s security posture. &lt;a href="https://blog.google/security/android-security-state-libraries/" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="ctme1"&gt;&lt;b&gt;Delivering new partner security agents and AI defenses with Gemini Enterprise&lt;/b&gt;: We're expanding our catalog of partner-built security offerings in the Gemini Enterprise ecosystem to help you leverage your full security context. &lt;a href="https://cloud.google.com/blog/products/identity-security/google-cloud-partners-deliver-new-security-agents-and-ai-defenses-with-gemini-enterprise"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="4hp9o"&gt;&lt;b&gt;Google named a Leader in the External Threat Intelligence Service Forrester Wave&lt;/b&gt;: We are proud to announce that Forrester has named Google a Leader in The Forrester Wave™: External Threat Intelligence Service Providers, Q3 2026. &lt;a href="https://cloud.google.com/blog/products/identity-security/google-named-a-leader-in-the-external-threat-intelligence-service-forrester-wave"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="6rdl1"&gt;&lt;b&gt;Wiz named a Leader in the Proactive Security Platforms Forrester Wave&lt;/b&gt;: Forrester’s Proactive Security Platforms evaluation for Q3 2026 rated Wiz with top scores across eight areas, reflecting our commitment to securing the AI era. &lt;a href="https://www.wiz.io/blog/forrester-wave-for-proactive-security-2026" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="c0jei"&gt;&lt;b&gt;Strengthen your CI/CD pipeline with new Secure Source Manager capabilities&lt;/b&gt;: To help you better address software supply chain threats, our Secure Source Manager lets you manage your source and CI/CD systems with unified authentication and authorization mechanisms. &lt;a href="https://cloud.google.com/blog/products/identity-security/strengthen-your-cicd-pipeline-with-new-secure-source-manager-capabilities"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="65fjj"&gt;&lt;b&gt;Building an AI detection engine that understands agent intent&lt;/b&gt;: Analyzing model input and output logs in an AI-native detection pipeline to understand and uncover malicious AI agent behavior. &lt;a href="https://www.wiz.io/blog/building-an-ai-detection-engine-for-agent-intent" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="8url6"&gt;&lt;b&gt;Using AI to discover and fix high-priority exposures across public services and critical infrastructure&lt;/b&gt;: New initiative partners with under-resourced organizations to uncover, remediate exploitable risk at scale. &lt;a href="https://www.wiz.io/blog/scan-for-good-critical-ai-exposures" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="97ope"&gt;Please visit the Google Cloud blog for more security stories &lt;a href="https://cloud.google.com/blog/products/identity-security"&gt;published this month&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Join the Google Cloud CISO Community&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cc9db90&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;Learn more&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;https://rsvp.withgoogle.com/events/google-cloud-ciso-community-interest-form-2026?utm_source=cgc-blog&amp;amp;utm_medium=blog&amp;amp;utm_campaign=FY25-Q1-global-GCP30328-physicalevent-er-dgcsm-parent-CISO-community-2025&amp;amp;utm_content=cisop_&amp;amp;utm_term=-&amp;#x27;), (&amp;#x27;image&amp;#x27;, &amp;lt;GAEImage: GCAT-replacement-logo-A&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="29tyz"&gt;&lt;b&gt;Threat Intelligence news&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="61jq"&gt;&lt;b&gt;ShinyHunters renewed mass exploitation campaign targeting Oracle PeopleSoft&lt;/b&gt;: Mandiant and Google Threat Intelligence Group (GTIG) have identified renewed mass exploitation of CVE-2026-35273 by UNC6240 (ShinyHunters), along with expanded global targeting across multiple sectors. &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/shinyhunters-renewed-mass-exploitation-campaign-targeting-oracle-peoplesoft"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="db7im"&gt;&lt;b&gt;Proactively defend by hardening code pipelines and CI/CD infrastructure&lt;/b&gt;: Check out our actionable blueprint for software and platform architects designed to safeguard the software supply chain against threat vectors that are actively being exploited, third-party risks, and architectural vulnerabilities throughout the entire software development lifecycle. &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/hardening-code-pipelines-and-ci-cd-infrastructure"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="bn9hu"&gt;&lt;b&gt;Infostealer incursion: How stolen credentials breach cloud, code, and AI environments&lt;/b&gt;: Wiz Research analyzes NordStellar data to map the credentials targeted by infostealer families and assess their potential impact across cloud, code, and AI environments. &lt;a href="https://www.wiz.io/blog/infostealer-incursion-cloud-ai-credentials" target="_blank"&gt;&lt;b&gt;Read more&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="b1vg8"&gt;Please visit the Google Cloud blog for more threat intelligence stories &lt;a href="https://cloud.google.com/blog/topics/threat-intelligence/"&gt;published this month&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h3 data-block-key="rcfc5"&gt;&lt;b&gt;Now hear this: Podcasts from Google Cloud&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="a4hsh"&gt;&lt;b&gt;Cloud Security Podcast: Patching browsers with AI, agents, Rust, and your tabs&lt;/b&gt;: Jasika Bawa and Doug Turner of Chrome Security explore how Google Chrome now uses AI agents to autonomously identify and patch security vulnerabilities at an unprecedented scale, significantly accelerating the browser's update cadence. &lt;a href="https://www.youtube.com/watch?v=pCXT8lQqg_U" target="_blank"&gt;&lt;b&gt;Listen here&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="8g5vm"&gt;&lt;b&gt;Cloud Security Podcast: All about Project Atlas, Wiz's AI vulnerability research&lt;/b&gt;: Nir Orfeld, head of vulnerability research, Wiz, discusses how his team uses multi-agent AI systems for discovering high-impact zero-day vulnerabilities in cloud infrastructure. &lt;a href="https://www.youtube.com/watch?v=qRJJ9ekpuVg" target="_blank"&gt;&lt;b&gt;Listen here&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="fqt39"&gt;&lt;b&gt;Cloud Security Podcast: How Google eliminates classes of vulnerabilities at scale&lt;/b&gt;: How do you build the foundations for a secure Google-scale enterprise that stays secure even if an AI is writing the code and nobody has time to review it? Christoph Kern, principal security engineer, Google, explores what secure-by-design really means in the AI era. &lt;a href="https://www.youtube.com/watch?v=43imRRfgLgc" target="_blank"&gt;&lt;b&gt;Listen here&lt;/b&gt;&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="5he96"&gt;To have our Cloud CISO Perspectives post delivered twice a month to your inbox, &lt;a href="https://cloud.google.com/resources/google-cloud-ciso-newsletter-signup"&gt;sign up for our newsletter&lt;/a&gt;. We’ll be back in a few weeks with more security-related updates from Google Cloud.&lt;/p&gt;&lt;/div&gt;</description><pubDate>Wed, 30 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-how-cybersecurity-startups-can-win-cisos/</guid><category>Cloud CISO</category><category>Startups</category><category>Security &amp; Identity</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Cloud_CISO_Perspectives_header_4_Blue.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Cloud CISO Perspectives: How cybersecurity startups can win CISOs</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Cloud_CISO_Perspectives_header_4_Blue.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-how-cybersecurity-startups-can-win-cisos/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Alicja Cade</name><title>Sr. Director, Financial Services, Office of the CISO</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Nick Godfrey</name><title>Senior Director, Office of the CISO</title><department></department><company></company></author></item><item><title>Empower your agents with the Google Cloud CLI remote MCP server</title><link>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’re expanding our ecosystem of managed remote MCP servers by introducing the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud CLI remote MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in preview.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Powered by the popular &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/gcloud"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/bq-cli-reference"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bq (BigQuery)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; command-line tools, this new server gives AI agents immediate, broad access to command-line operations for managing Google Cloud infrastructure and working with advanced BigQuery workflows securely and seamlessly. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Why CLI matters for AI agents&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agents are increasingly performing complex cloud operations, but standardizing how they interact with backend systems remains a challenge. The Google Cloud CLI remote MCP server bridges this gap by packaging the versatility of hundreds of gcloud and bq commands into one single MCP server. This results in two strong benefits for the agent:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Higher-level abstractions:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; CLI commands package complex multi-step workflows, validation checks, and high-level operations into unified commands rather than requiring multi-step API orchestration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Leverages model training:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; LLMs are heavily pre-trained on public command-line documentation, syntaxes, and usage examples, making CLI invocation intuitive and highly accurate for models.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The benefits of putting CLI behind remote MCP&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing cloud infrastructure with AI agents traditionally requires installing and maintaining Google Cloud CLI binaries inside agent execution environments. The Cloud CLI remote MCP server bridges CLI capabilities with MCP benefits by providing an isolated execution sandbox on Google Cloud infrastructure. This solves key infrastructure challenges:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Simplified dependency and runtime management:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; For teams building custom agents, maintaining local CLI versions and dependencies across dev, test, and production environments creates operational overhead. Remote MCP eliminates local installations and runtime maintenance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access for web-based agent endpoints:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Web-hosted agent platforms and web interfaces (such as Gemini Enterprise and other hosted enterprise agent platforms) run in environments where users cannot control or install local packages. Remote MCP enables secure, managed access to Google Cloud CLI operations directly from these surfaces.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise-grade security and governance&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting an AI agent to your infrastructure requires strict, enterprise-ready safeguards. This remote server leverages Google Cloud's standard identity and governance frameworks to keep your environments secure:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero ambient credentials:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The server isolates execution in a network-restricted proxy boundary with no ambient credentials. Authentication and authorization are handled through &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/iam/docs/agent-identity-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Identity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://developers.google.com/identity/protocols/oauth2" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OAuth 2.0&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/iam/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Identity and Access Management (IAM)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Strict policy enforcement:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Every command executed through the remote MCP server is run with the permissions of the authenticated caller identity. Both standard IAM permissions and organization policy service constraints are strictly enforced against downstream target resources.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Advanced protection with Model Armor:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; To minimize the risks associated with AI tool calling, the Cloud CLI remote MCP server integrates with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/model-armor/model-armor-mcp-google-cloud-integration"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model Armor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. You can proactively screen LLM prompts and responses to protect against risks like prompt injection and malicious inputs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud audit logging:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The Cloud CLI remote MCP server can be configured to log every tool invocation to Audit Logs (Data Access logs under cloudcli.googleapis.com/mcp). Security teams can gain full visibility into caller identities, OAuth clients, and IAM authorization decisions (mcp.googleapis.com/tools.call) without exposing sensitive command payloads or personally identifiable information (PII).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting to the Google Cloud CLI Remote MCP Server&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Integrating cloud management into your agents no longer requires packaging Google Cloud CLI binaries, managing local execution runtimes, or maintaining dependencies inside agent container images. Because the Google Cloud CLI remote MCP server implements the standard Model Context Protocol, any MCP-compatible agent platform or orchestration runtime can connect immediately via standard configuration:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;mcpServers&amp;quot;: {\r\n    &amp;quot;google-cloud-cli&amp;quot;: {\r\n      &amp;quot;uri&amp;quot;: &amp;quot;https://cloudcli.googleapis.com/mcp&amp;quot;,\r\n      ...\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fc43cc9fd90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Your agent immediately gains access to execute gcloud and bq commands in a secure, network-isolated cloud sandbox.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Authentication is handled via keyless Agent Identity for hosted Google Cloud platforms, or standard OAuth 2.0 for external runtimes. For authentication options, see the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/mcp/set-up-authentication-mcp-servers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP Authentication Guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Bringing infrastructure management to agents&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;The Cloud CLI remote MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; exposes two powerful tools, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_bq_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, giving your AI agents broad, immediate access to Google Cloud operations through natural language.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Managing cloud infrastructure with &lt;/strong&gt;&lt;code&gt;&lt;strong style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/strong&gt;&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, agents can execute the full breadth of gcloud operations to manage, diagnose, and secure your Google Cloud environment. An example follows:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Observability and incident diagnostics&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An agent streamlines incident diagnostics by automating command execution and reducing context-switching across tools.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_IPA9ALa.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Extending BigQuery operations with the run_bq_command tool&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/use-bigquery-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; already helps organizations analyze and explore data using AI agents, with the introduction of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_bq_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, agents can now tackle advanced BigQuery tasks such as resource allocation, job monitoring, and task scheduling by unlocking the full scope of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/reference/mcp#mcp-tools"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bq CLI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; functionality. Key capabilities include:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Automating scheduled queries&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An agent utilizes BigQuery Data Transfer Service configurations to schedule queries automatically. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Job and resource management&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Gain deep insight into query execution details, including processed data volume, slot usage, and execution plans, as well as managing reservations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access and permissions control:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Data administrators and owners can inspect and update table permissions directly through the agent.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_JA5HWbH.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Pricing and availability&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Google Cloud CLI MCP server is available today in public preview. There is no additional charge to use the MCP server itself. You pay only for the GCP resources you create and any applicable data transfer costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;To get started&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, enable the Cloud CLI Execution API (`&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;cloudcli.googleapis.com&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;`) in your Google Cloud project, grant the required &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;MCP Tool User&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (`&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;roles/mcp.toolUser&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;`) IAM role to your agent or user identity, and configure your MCP client to connect to `cloudcli.googleapis.com/mcp`.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Use the Google Cloud CLI Remote MCP Server Guide&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/sdk/reference/mcp#mcp-tools"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud CLI MCP Reference&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Remote MCP Servers Overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/supported-products"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Full List of Google OneMCP Servers&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/set-up-authentication-mcp-servers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Set up authentication to Google and Google Cloud MCP servers&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/govern/agent-identity-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform &amp;amp; Agent Identity Overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.youtube.com/watch?v=-fb0ycu4kiU" rel="noopener" target="_blank"&gt;&lt;span data-rich-links='{"fple-t":"Automate Google Cloud with Cloud CLI Remote MCP Server","fple-u":"https://www.youtube.com/watch?v=-fb0ycu4kiU","fple-mt":null,"type":"first-party-link"}' style="text-decoration: underline; vertical-align: baseline;"&gt;Automate Google Cloud with Cloud CLI Remote MCP Server&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 30 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</guid><category>Data Analytics</category><category>Developers &amp; Practitioners</category><category>AI &amp; Machine Learning</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Empower your agents with the Google Cloud CLI remote MCP server</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Prosper Nwankpa</name><title>Senior Engineering Manager, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Adam Hwang</name><title>Software Engineering Manager, Google Cloud</title><department></department><company></company></author></item></channel></rss>