<?xml version="1.0" encoding="utf-8"?>

<feed xmlns="http://www.w3.org/2005/Atom" >
  <generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator>
  <link href="https://tedt.org/atom.xml" rel="self" type="application/atom+xml" />
  <link href="https://tedt.org/" rel="alternate" type="text/html" />
  <updated>2026-09-10T11:41:50-07:00</updated>
  <id>https://tedt.org/</id>

  
    <title type="html">Ted Tschopp</title>
  

  
    <subtitle>Technology, theology, and creativity converge in this insightful exploration by Ted Tschopp, who navigates the intricate relationship between these topics in the digital age. From dissecting how technology is reshaping religious practices and ethics to sharing cutting-edge trends in UI/UX design, collaboration tools, and digital storytelling, the author provides a multifaceted analysis. Written with a unique and engaging blend of humor, reflection, and curiosity, this work invites readers to ponder and discuss these connections, offering a contemplative lens on modern innovation and cultural evolution.</subtitle>
  

  
    <author>
        <name>Ted Tschopp</name>
      
        <email>ted@tschopp.org</email>
      
      
        <uri>https://tedt.org/profile/</uri>
      
    </author>
  

  
  
    
    
    <entry>
      <title type="html">OpenAI, Hugging Face &amp;amp; Wikis: Recent Developments</title>
      <link href="https://tedt.org/slides/openai-hugging-face-wiki/" rel="alternate" type="text/html" title="OpenAI, Hugging Face &amp; Wikis: Recent Developments" />
      <published>2026-09-10T00:00:00-07:00</published>
      <updated>2026-09-10T00:00:00-07:00</updated>
      <id>https://tedt.org/slides/openai-hugging-face-wiki</id>
      <content type="html" xml:base="https://tedt.org/slides/openai-hugging-face-wiki/"></content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
          <email>ted@tschopp.org</email>
        
        
          <uri>https://tedt.org/profile/</uri>
        
      </author>

      

      

      
        <summary type="html"></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/slides/decks/openai-hugging-face-wiki/slide-preview.png" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">How Much Energy and Water Does AI Use?</title>
      <link href="https://tedt.org/Energy-and-Water-Usage-for-Models/" rel="alternate" type="text/html" title="How Much Energy and Water Does AI Use?" />
      <published>2026-09-06T00:00:00-07:00</published>
      <updated>2026-09-06T00:00:00-07:00</updated>
      <id>https://tedt.org/Energy-and-Water-Usage-for-Models</id>
      <content type="html" xml:base="https://tedt.org/Energy-and-Water-Usage-for-Models/">&lt;p&gt;I was asked how much electricity and water AI uses, and what that means. These are shared resources, and questions about their use deserve clear answers. The numbers vary, but familiar comparisons can help put them in context.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For a short text answer, the electricity use can be quite small.&lt;/strong&gt; A 2026 study estimated about &lt;strong&gt;0.3 watt-hours&lt;/strong&gt;, roughly enough to run a &lt;strong&gt;10-watt LED lightbulb for two minutes&lt;/strong&gt;. A longer reasoning response in the same study was estimated to use about &lt;strong&gt;4 watt-hours&lt;/strong&gt;, equivalent to keeping that bulb on for &lt;strong&gt;24 minutes&lt;/strong&gt;. These estimates describe particular text-based tasks and computing systems; individual requests can use more or less. &lt;a href=&quot;https://doi.org/10.1016/j.joule.2026.102430&quot;&gt;Research study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At the shorter-response estimate, asking 100 questions every day for a year would use around &lt;strong&gt;11 kilowatt-hours&lt;/strong&gt;—about what a 1,500-watt space heater uses in seven hours. A small amount per answer can still add up to substantial demand when millions of people use these services.&lt;/p&gt;

&lt;p&gt;Training an AI model is a larger, separate undertaking. Training GPT-3, released in 2020, was estimated to require &lt;strong&gt;1.3 million kilowatt-hours&lt;/strong&gt;: roughly a year’s electricity purchases for &lt;strong&gt;120 average U.S. homes&lt;/strong&gt;, using the government’s 2022 household benchmark. Developing and updating a model can involve several training stages and experiments, but answering an ordinary question does not restart that process. &lt;a href=&quot;https://arxiv.org/pdf/2104.10350&quot;&gt;Training estimate&lt;/a&gt;, &lt;a href=&quot;https://www.eia.gov/tools/faqs/faq.php?id=97&amp;amp;t=7&quot;&gt;household benchmark&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Water use is harder to describe with one number.&lt;/strong&gt; Google estimated about &lt;strong&gt;five drops of cooling water&lt;/strong&gt; for its typical Gemini text response in May 2025. Mistral reported &lt;strong&gt;45 milliliters—about three tablespoons—for a roughly 300-word response&lt;/strong&gt;, using a broader calculation that includes impacts beyond cooling, such as manufacturing equipment. These figures describe different systems and count different things, so they cannot establish which service uses less water overall. &lt;a href=&quot;https://arxiv.org/html/2508.15734v1&quot;&gt;Google study&lt;/a&gt;, &lt;a href=&quot;https://mistral.ai/news/our-contribution-to-a-global-environmental-standard-for-ai/&quot;&gt;Mistral disclosure&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Training also has a water footprint. One research scenario estimated &lt;strong&gt;5.4 million liters&lt;/strong&gt; for GPT-3’s training when cooling and electricity generation were included. That is &lt;strong&gt;just over two Olympic-size swimming pools&lt;/strong&gt;, assuming 2.5 million liters per pool. Another way of thinking about this at SCE is that this is the amount of water it would take to fill the pools of water that cooled the spent fuel at SONGS Unit 2 4x when it was operational.  Another SONGS compairson is that this was the amount of salt water SONGS Unit 2 and 3 processed in roughly 51 seconds to generate power.  That 5.4 million liters of water is a research estimate, rather than a disclosed measurement of water consumed during the actual training run. &lt;a href=&quot;https://arxiv.org/html/2304.03271v5&quot;&gt;Water study&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The water is &lt;strong&gt;not destroyed&lt;/strong&gt;. Evaporated water enters the atmosphere and eventually returns as precipitation, but it may return somewhere else or much later. A community experiencing drought still needs water in its own reservoirs and groundwater supplies today. That local availability is an important part of the environmental impact. &lt;a href=&quot;https://www.usgs.gov/mission-areas/water-resources/science/water-use-terminology&quot;&gt;USGS explanation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There are practical ways to reduce demand. Some cooling systems circulate the same water repeatedly, much like a car’s radiator. Microsoft reports that more than 90 percent of its Fairwater facility’s capacity uses a system with &lt;strong&gt;no evaporation losses&lt;/strong&gt;. That illustrates what newer facilities can achieve. &lt;a href=&quot;https://blogs.microsoft.com/blog/2025/09/18/inside-the-worlds-most-powerful-ai-datacenter/&quot;&gt;Microsoft’s description&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Some escaping water can also be recovered. In 2021, MIT described demonstrations at its power facilities of equipment that collects water droplets from cooling-tower plumes and returns them for reuse. &lt;a href=&quot;https://news.mit.edu/2021/infinite-cooling-nuclear-0803&quot;&gt;MIT demonstration&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cooling water can also carry useful heat from the computers.&lt;/strong&gt; Recovering that heat can warm buildings and reduce their need for other heating sources. In a 2024 report, Meta said its Odense data center supplied recovered heat through a local heating network to about &lt;strong&gt;7,000 households&lt;/strong&gt;. This makes further use of the energy consumed by computers; it does not recover all that energy or turn it back into electricity. &lt;a href=&quot;https://datacenters.atmeta.com/wp-content/uploads/2024/10/Denmark-Odense.pdf&quot;&gt;Meta’s facility report&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Pacific Gas and Electric Company and developer Westbank announced plans in April 2025 for a San Jose development with &lt;strong&gt;three data centers and up to 4,000 homes&lt;/strong&gt;. The project would reuse data-center heat through a network serving surrounding buildings. &lt;a href=&quot;https://investor.pgecorp.com/news-events/press-releases/press-release-details/2025/PGE-Begins-Energy-Infrastructure-Upgrades-to-Bring-San-Joses-Net-Zero-Community-to-Life/default.aspx&quot;&gt;PG&amp;amp;E announcement&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For text, “tokens” are the small pieces of words and punctuation an AI processes. Electricity and water use can be expressed per million tokens, but those figures depend on the model, the task, and the facilities serving it. They are useful accounting measures when those conditions are specified.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="AI energy use" />
      
        <category term="AI water use" />
      
        <category term="model training" />
      
        <category term="AI inference" />
      
        <category term="data centers" />
      
        <category term="cooling" />
      
        <category term="heat recovery" />
      
        <category term="sustainability" />
      

      
        <summary type="html">AI uses electricity and water, but the impact depends on the task, the model, and the facility. Familiar comparisons help explain the scale.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-09/Energy-and-Water-Usage-for-Models.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">The Same Model, Different Rules</title>
      <link href="https://tedt.org/The-Same-Model-Different-Rules/" rel="alternate" type="text/html" title="The Same Model, Different Rules" />
      <published>2026-09-05T09:00:00-07:00</published>
      <updated>2026-09-05T09:00:00-07:00</updated>
      <id>https://tedt.org/The-Same-Model-Different-Rules</id>
      <content type="html" xml:base="https://tedt.org/The-Same-Model-Different-Rules/">&lt;h1 id=&quot;the-same-model-different-rules&quot;&gt;The Same Model, Different Rules&lt;/h1&gt;

&lt;h2 id=&quot;a-broken-iphone-and-a-slowly-charging-blackberry&quot;&gt;A Broken iPhone and a Slowly Charging BlackBerry&lt;/h2&gt;

&lt;p&gt;In 2007, &lt;a href=&quot;https://tedt.org/sucks-to-be-me-from-blackberry-to-iphone-to-blackberry-again/&quot;&gt;I dropped my iPhone at lunch&lt;/a&gt;. The glass shattered.&lt;/p&gt;

&lt;p&gt;At the Apple Store, I was told the repair would cost $250. I asked for a loaner. They didn’t have one. I asked to have the repaired phone shipped to my home. One employee said they couldn’t do that. Another corrected him: they could. A refurbished replacement wasn’t an option either.&lt;/p&gt;

&lt;p&gt;So I was back on my BlackBerry, which was charging slowly, while I waited for an iPhone expected back in two or three business days.&lt;/p&gt;

&lt;p&gt;The broken glass was easy to see. The other dependencies became visible one question at a time. What would the store repair? What could it replace? Who knew the shipping rules? What could I use while the phone was gone?&lt;/p&gt;

&lt;p&gt;The answers determined how I would get through the interruption.&lt;/p&gt;

&lt;p&gt;An enterprise choosing an AI service faces a version of that same problem. The technology may be impressive. But the work also depends on arrangements around it: who can use it, who can change the conditions, and what remains available when the preferred option disappears.&lt;/p&gt;

&lt;p&gt;Those arrangements deserve attention before something breaks.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI procurement now includes the conditions under which work can start, continue, and stop.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On June 9, Anthropic released Claude Fable 5 and its restricted-access counterpart, Mythos 5. Three days later, access to both was suspended.&lt;/p&gt;

&lt;p&gt;In its account of the interruption, Anthropic said newly imposed export controls required restrictions based on nationality. Unable to verify nationality reliably in real time, it suspended access for everyone. Fable access was restored on July 1. &lt;a href=&quot;https://www.anthropic.com/news/redeploying-fable-5&quot;&gt;Anthropic’s account&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A team evaluating the model could have tested its answers, measured its speed, and calculated its cost. Those tests would have helped determine whether it could do useful work. They would not have established whether the team could keep using it when the conditions of access changed.&lt;/p&gt;

&lt;p&gt;That is the purchasing question the September announcements bring into sharper focus.&lt;/p&gt;

&lt;p&gt;An enterprise needs to approve a service configuration: the model, the route to it, the people permitted to use it, the data arrangements, and the response when any of those conditions changes. A model name alone cannot describe that purchase.&lt;/p&gt;

&lt;h2 id=&quot;the-differences-are-already-reaching-buyers&quot;&gt;The Differences Are Already Reaching Buyers&lt;/h2&gt;

&lt;p&gt;Anthropic describes Fable 5.1 and Mythos 5.1 as the same underlying model with different safeguards. Fable is generally available; Mythos is offered through trusted-access arrangements supporting sensitive cybersecurity and life-sciences work. Access remains program-dependent. &lt;a href=&quot;https://www.anthropic.com/claude-fable-and-mythos-5-1&quot;&gt;Anthropic’s September announcement&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Google describes Gemini 3.8 Flash and Flash Cyber as sharing foundational intelligence, with different safeguards and deployment arrangements. Flash Cyber is restricted to trusted defenders through Fairwind. Google’s wording does not establish identical model weights, so the two companies’ claims should not be treated as interchangeable. &lt;a href=&quot;https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/&quot;&gt;Google’s announcement&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Restricted access is not new. What these releases make increasingly visible is how much of the product sits around the intelligence.&lt;/p&gt;

&lt;p&gt;The differences also extend beyond the model companies’ announcements.&lt;/p&gt;

&lt;p&gt;Harvey’s Sept. 1 notice makes Fable 5.1 an opt-in offering. It warns that the model’s processing practices may differ from commitments in existing customer agreements, describes a time-bound zero-retention arrangement for eligible customers, and says processing occurs in the United States. Those statements describe Harvey’s published offering; they should not be generalized to every distribution channel. &lt;a href=&quot;https://www.harvey.ai/blog/fable-5-1-in-harvey&quot;&gt;Harvey’s implementation notice&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Cursor documents a different set of operating decisions. Enterprise customers need explicit approval to enable Fable 5.1. Cursor’s own Privacy Mode does not eliminate Anthropic’s separate retention requirements. Requests that trigger safeguards can automatically fall back to Opus. &lt;a href=&quot;https://cursor.com/docs/models/claude-fable-5-1&quot;&gt;Cursor’s documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;These are commercially interested providers describing their implementations, not independent studies of customer outcomes. They nevertheless establish something useful: choosing the model does not settle the terms under which an employee encounters it.&lt;/p&gt;

&lt;p&gt;The route matters. So does the agreement attached to that route.&lt;/p&gt;

&lt;h2 id=&quot;write-down-the-service-you-are-approving&quot;&gt;Write Down the Service You Are Approving&lt;/h2&gt;

&lt;p&gt;In &lt;a href=&quot;https://tedt.org/Model-Portability-Is-Not-AI-Portability/&quot;&gt;“Model Portability Is Not AI Portability,”&lt;/a&gt; I argued that moving a model does not automatically move the complete working system. The newer question is what happens when parts of that system change while the model name stays the same.&lt;/p&gt;

&lt;p&gt;Anthropic’s versioning documentation provides a concrete example. A model ID pins the model’s weights and configuration, but surrounding infrastructure—including routing, safety classifiers, and sampling logic—can change. The documentation says those changes can affect observable behavior. &lt;a href=&quot;https://platform.claude.com/docs/en/about-claude/models/model-ids-and-versions&quot;&gt;Model IDs and versioning&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A fixed model version is useful. It is not a frozen service.&lt;/p&gt;

&lt;p&gt;The enterprise catalog therefore needs to record more than “approved model.” It needs to describe the approved use of that model and the evidence supporting the approval.&lt;/p&gt;

&lt;p&gt;Consider a hypothetical internal source-code review service. Its purpose is to inspect approved repositories and propose repairs. The following is a proposed approval record, not a claim that a particular product already satisfies every requirement.&lt;/p&gt;

&lt;table class=&quot;well table table-striped&quot;&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Boundary&lt;/th&gt;
      &lt;th&gt;Proposed Configuration&lt;/th&gt;
      &lt;th&gt;Owner and Evidence&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Model and routing&lt;/td&gt;
      &lt;td&gt;Fable 5.1 through the company gateway to the Claude API; no unapproved fallback&lt;/td&gt;
      &lt;td&gt;The platform owner records the exact model ID, interface, routing settings, and test results. Supplier-controlled changes are tracked separately.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Eligible users and entities&lt;/td&gt;
      &lt;td&gt;One corporate entity’s internal security team; no assumed access for affiliates or contractors&lt;/td&gt;
      &lt;td&gt;Procurement confirms applicable access rights. The identity team enforces membership. Unconfirmed coverage blocks access.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Permitted actions&lt;/td&gt;
      &lt;td&gt;Read isolated copies of two approved repositories and draft patches; no merging or production deployment&lt;/td&gt;
      &lt;td&gt;The application owner enforces tool permissions and tests attempted actions outside that scope.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Data handling&lt;/td&gt;
      &lt;td&gt;Only approved source code, with processing, storage, review locations, and retention explicitly accepted&lt;/td&gt;
      &lt;td&gt;Privacy and security owners retain the applicable terms and verified data-flow description. Unknown locations or retention conditions block approval.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Monitoring and evidence&lt;/td&gt;
      &lt;td&gt;Company-held records of requests, actions, approvals, and unfinished work, limited to necessary information&lt;/td&gt;
      &lt;td&gt;Security identifies authorized reviewers and tests whether required evidence remains retrievable after supplier access is lost.&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Interruption and recovery&lt;/td&gt;
      &lt;td&gt;Pause affected jobs; use a separately approved alternative or manual review&lt;/td&gt;
      &lt;td&gt;The service owner names the person authorized to resume work and records a tested recovery procedure.&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Each row needs a status: documented, contractually confirmed where necessary, or not yet tested. A public description of a feature is not evidence that the enterprise has enabled it correctly.&lt;/p&gt;

&lt;p&gt;There are also two different owners to record. One can change the condition. The other decides whether the enterprise can continue operating after that change.&lt;/p&gt;

&lt;p&gt;The supplier may change a classifier. Your architecture board cannot prevent that merely by approving a design. It can decide what evidence is needed afterward and whether the affected workflow must pause.&lt;/p&gt;

&lt;p&gt;That distinction turns an inventory into an operating record.&lt;/p&gt;

&lt;h2 id=&quot;keep-control-of-the-actions&quot;&gt;Keep Control of the Actions&lt;/h2&gt;

&lt;p&gt;Provider permission and enterprise authority still answer different questions.&lt;/p&gt;

&lt;p&gt;A supplier may permit vulnerability analysis. That does not authorize an agent to inspect every system reachable from your network. An employee may be allowed to read a repository without being allowed to deploy a repair.&lt;/p&gt;

&lt;p&gt;The earlier &lt;a href=&quot;https://tedt.org/Dev-Test-and-Prod-Still-Matter-What-Gets-Deployed-Has-Changed/&quot;&gt;capability-versus-authority argument&lt;/a&gt; remains the foundation. The purchasing decision now needs to establish which controls the enterprise can enforce independently of the supplier’s judgment.&lt;/p&gt;

&lt;p&gt;For consequential workflows, those controls should include user identity, permitted tools and targets, approval before consequential changes, records of completed actions, and a way to stop further execution. Their effectiveness should not depend on the model agreeing that a request is inappropriate.&lt;/p&gt;

&lt;p&gt;Google’s Fairwind requirements illustrate the division. Google controls program admission, while participants must implement user authentication, phishing-resistant multifactor authentication, applicable access controls, and tracking of employee access and use. Access is limited to specified internal security teams. &lt;a href=&quot;https://deepmind.google/fairwind-program/&quot;&gt;Fairwind’s governance requirements&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This does not mean every enterprise should rebuild the provider’s safeguards.&lt;/p&gt;

&lt;p&gt;A managed service may be sufficient when it can demonstrate the required controls and the customer can administer and audit them. An enterprise-controlled harness becomes necessary where the workflow requires permissions, approvals, evidence, or recovery that the managed service cannot reliably provide.&lt;/p&gt;

&lt;p&gt;“Controlled” does not have to mean “built from scratch.” It means the organization can establish and enforce the required boundary. If neither the purchased service nor an integration can do that, the consequential action should remain outside the automation.&lt;/p&gt;

&lt;h2 id=&quot;data-custody-is-another-operating-responsibility&quot;&gt;Data Custody Is Another Operating Responsibility&lt;/h2&gt;

&lt;p&gt;Anthropic’s planned Enterprise Frontier Safeguards would allow monitoring data to reside in customer-controlled cloud infrastructure, with options for customer-managed keys and customer review of flagged activity. Rollout is planned in phases later this fall. The storage, key-management, and review options are separately configurable; the announcement should not be read as proof that every customer already has them. &lt;a href=&quot;https://www.anthropic.com/news/enterprise-frontier-safeguards&quot;&gt;Enterprise Frontier Safeguards&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Its interim zero-retention arrangement also has limits. Anthropic describes it as temporary and restricted to eligible uses. Real-time safeguards and enforcement continue to apply, and the company says it may modify or withdraw the arrangement. &lt;a href=&quot;https://support.claude.com/en/articles/15425695-covered-models&quot;&gt;Covered Models guidance&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Moving records into the customer’s environment changes custody. It does not remove the need to monitor activity, investigate problems, or establish who can suspend service.&lt;/p&gt;

&lt;p&gt;Someone must receive the alert. Someone must have permission to examine the evidence. Someone must distinguish misuse from legitimate work that was incorrectly flagged.&lt;/p&gt;

&lt;p&gt;For a multinational organization, the approval record should separately identify the contracting entity, eligible users, inference locations, stored-data locations, and locations from which people may review the data. An approved storage region does not answer all those questions.&lt;/p&gt;

&lt;p&gt;Harvey’s processing notice supplies one concrete reason to ask. It does not establish what every organization in Europe, the United Kingdom, or Asia-Pacific may lawfully or contractually do. Those decisions require the exact service configuration and applicable agreement, not a general claim of global availability.&lt;/p&gt;

&lt;h2 id=&quot;purchase-the-change-and-exit-arrangements&quot;&gt;Purchase the Change and Exit Arrangements&lt;/h2&gt;

&lt;p&gt;The contract discussion should describe events, not merely desirable qualities.&lt;/p&gt;

&lt;p&gt;Anthropic’s public commercial terms, for example, allow suspension in specified circumstances and describe reasonable efforts to provide written notice and restore access after a curable cause is resolved. That is different from a guaranteed advance-warning period or a fixed recovery time. Customer-specific agreements may differ. &lt;a href=&quot;https://www.anthropic.com/legal/commercial-terms&quot;&gt;Public commercial terms, Section I&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Procurement, legal, security, and the service owner should resolve four practical matters:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Changes and versions:&lt;/strong&gt; What is pinned, what can change around it, what notice is available, and what testing is required before affected work continues?&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Scope and location:&lt;/strong&gt; Which entities, employees, contractors, purposes, and processing arrangements are covered? Does an approval survive a change in any of them?&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Evidence and review:&lt;/strong&gt; What information can the customer obtain during an incident, who reviews disputed flags, and what remains accessible if the account is suspended?&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Revocation and exit:&lt;/strong&gt; What can trigger loss of eligibility, how can a decision be challenged, and what export, transition assistance, charges, and time limits apply?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are matters to confirm or negotiate, not assurances that suppliers necessarily offer.&lt;/p&gt;

&lt;p&gt;The purchasing decision must also survive an unwelcome answer. If the provider cannot guarantee advance notice, a critical workflow needs a tested interruption response. If the evidence needed to reconstruct a transaction cannot be obtained, that workflow may need different architecture.&lt;/p&gt;

&lt;p&gt;A contract cannot prevent every interruption. It can make the responsibilities clearer before people are trying to recover under pressure.&lt;/p&gt;

&lt;h2 id=&quot;rehearse-losing-access&quot;&gt;Rehearse Losing Access&lt;/h2&gt;

&lt;p&gt;Now take the hypothetical code-review service and interrupt it.&lt;/p&gt;

&lt;p&gt;The supplier suspends the organization’s access while several reviews are underway. Some draft patches have been saved. Other jobs are waiting. This is a planning scenario, not a reported customer incident.&lt;/p&gt;

&lt;p&gt;The platform should stop dispatching affected work and preserve the records it is authorized to retain. The service owner should establish which jobs completed, which produced partial work, and which never started.&lt;/p&gt;

&lt;p&gt;Paused does not mean rolled back.&lt;/p&gt;

&lt;p&gt;Next, determine the scope and reason for the restriction. A service interruption, suspected credential misuse, and a prohibition on the requested activity require different responses. An alternative model is appropriate only when the work remains permitted and the replacement’s access rights, data handling, quality, and controls have been approved. Changing accounts to evade a restriction is not recovery.&lt;/p&gt;

&lt;p&gt;If no approved alternative exists, route the work to manual review or hold it. The owner should know how much work that alternative can absorb and when delays become unacceptable.&lt;/p&gt;

&lt;p&gt;This is where the &lt;a href=&quot;https://tedt.org/Why-AI-Needs-a-Harness/&quot;&gt;harness around the model&lt;/a&gt; earns its cost. It preserves the organization’s ability to understand and manage unfinished work, even when the supplier is unavailable.&lt;/p&gt;

&lt;p&gt;The rehearsal should also produce a cost estimate.&lt;/p&gt;

&lt;p&gt;Suppose a switch requires 80 engineering hours and 40 review hours at an assumed blended rate of $150 an hour. That is $18,000 before additional licenses, parallel operation, service charges, or delayed work. These are illustrative assumptions, not measured customer costs.&lt;/p&gt;

&lt;p&gt;Use your own drill to replace them.&lt;/p&gt;

&lt;p&gt;Record the effort required to reconstruct work, validate the replacement, obtain approvals, and clear the backlog. Compare that expense with the cost of maintaining a ready alternative. Some workflows will justify the investment. Others will be adequately served by a documented manual procedure.&lt;/p&gt;

&lt;p&gt;Token prices alone cannot make that decision.&lt;/p&gt;

&lt;h2 id=&quot;keep-the-claim-proportionate&quot;&gt;Keep the Claim Proportionate&lt;/h2&gt;

&lt;p&gt;The public evidence supports a narrower conclusion than a sweeping transformation of every enterprise purchase.&lt;/p&gt;

&lt;p&gt;Specialized access may remain concentrated in sensitive domains. Providers may be better equipped than individual customers to operate sophisticated safeguards. A routine drafting workflow may need little beyond existing access controls and a human fallback.&lt;/p&gt;

&lt;p&gt;There is also an evidence boundary. Harvey and Cursor document real service arrangements, but their notices are not independent evaluations of enterprise outcomes. Public documentation does not establish the full regional eligibility of restricted programs, negotiated customer commitments, or measured switching costs. Those parts of the original reporting brief remain open.&lt;/p&gt;

&lt;p&gt;The practical conclusion does not require pretending otherwise.&lt;/p&gt;

&lt;p&gt;Before placing consequential work behind a model endpoint, describe the complete service, identify who can change its conditions, and test how the work will be recovered. Scale that preparation to the consequence of interruption.&lt;/p&gt;

&lt;p&gt;We don’t need to return to June.  This week OpenAI launched a new set of models, including those that are restricted to people working in the cyber security space.  If you intend to apply for access to these systems, you need to understand the eligibility requirements and prepare the necessary documentation, the question was no longer simply how well the model could perform.  It also needs to include how you will handle interruptions, maintain continuity, and ensure that critical work can proceed even if access is temporarily unavailable.&lt;/p&gt;

&lt;p&gt;The work still needs somewhere to go.&lt;/p&gt;

&lt;p&gt;A useful AI purchasing decision gives it that place: a permitted route, a recoverable state, and a person who knows what happens next.&lt;/p&gt;

&lt;h2 id=&quot;the-practical-lesson&quot;&gt;The Practical Lesson&lt;/h2&gt;

&lt;p&gt;I broke that iPhone, and for a while lost access to my brand new phone.  A shattered screen and a supplier suspending model access are different failures.&lt;/p&gt;

&lt;p&gt;But both lead to a practical question: what can you do while the thing you depend on is unavailable?&lt;/p&gt;

&lt;p&gt;I had asked for a loaner. What I actually had was my old phone, aslowly charging BlackBerry. Returning to it was inconvenient, but it gave me something to use while I waited.&lt;/p&gt;

&lt;p&gt;For an enterprise, the equivalent deserves more preparation. The alternative needs the right permissions, access to the necessary work, and someone authorized to decide when it is safe to continue. Its limitations need to be understood before people are depending on it.&lt;/p&gt;

&lt;p&gt;Choose the capable model. Test what it can accomplish. Then give the arrangements around it the same attention.&lt;/p&gt;

&lt;p&gt;Before handing over important work, know what your BlackBerry is.&lt;/p&gt;

&lt;p&gt;And make sure it is charged.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="enterprise AI" />
      
        <category term="AI procurement" />
      
        <category term="AI governance" />
      
        <category term="model access" />
      
        <category term="operational resilience" />
      
        <category term="service continuity" />
      
        <category term="vendor risk" />
      
        <category term="data custody" />
      
        <category term="access controls" />
      
        <category term="AI portability" />
      

      
        <summary type="html">AI procurement must define not only which model is approved, but also its access rules, data custody, interruption response, and exit arrangements.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-09/iPhone-Pearl.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">Model Portability Is Not AI Portability</title>
      <link href="https://tedt.org/Model-Portability-Is-Not-AI-Portability/" rel="alternate" type="text/html" title="Model Portability Is Not AI Portability" />
      <published>2026-08-28T09:00:00-07:00</published>
      <updated>2026-08-28T09:00:00-07:00</updated>
      <id>https://tedt.org/Model-Portability-Is-Not-AI-Portability</id>
      <content type="html" xml:base="https://tedt.org/Model-Portability-Is-Not-AI-Portability/">&lt;h1 id=&quot;model-portability-is-not-ai-portability&quot;&gt;Model Portability Is Not AI Portability&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Large enterprises are learning how to switch models. Moving a working business process without losing quality, controls, or money is a different problem.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;When I was a child, I loved trains. I still do. To me, they are the Industrial Revolution made visible: steel, power, motion, and human coordination assembled into machines capable of carrying people and freight across great distances. A train is not merely a machine. It is an entire system moving through the world.  They made the world a smaller and more connected place.&lt;/p&gt;

&lt;p&gt;My dad would take me to the model train exhibits at the LA County Fair and to the train store in Pasadena on Route 66. Fairplex offered both ends of the experience. It had a collection of full-sized trains outside and an entire section devoted to model railroads.&lt;/p&gt;

&lt;p&gt;For my parents, this was a wonderfully efficient arrangement. They could settle into the shade while I remained occupied for most of the day. They watched me watch the model trains travel round and round their carefully built layouts. I could also go and climb over those enormous beasts of steel and by the end of the day, I was exhausted, happy, and if I had been paying attention, I would have learned something about how the system worked.&lt;/p&gt;

&lt;p&gt;You see as I moved back and forth between two versions of the same world. On one side of Fairplex, I could stand over an entire railway and watch the system operate in miniature. But over on the other side, I was the miniature, climbing through locomotives built on a scale that made a child look very small.&lt;/p&gt;

&lt;p&gt;The difference was scale. The resemblance was the system.&lt;/p&gt;

&lt;p&gt;The locomotives were what drew the eye, whether they were small enough to hold or large enough to climb through. But even a model train cannot make a journey by itself.&lt;/p&gt;

&lt;p&gt;The locomotive needs track of the right gauge. It needs compatible power and controls. Its couplers must fit the cars. Switches must send it down the intended line. Signals must keep it from colliding with something already there. Behind the entire layout is someone who built it, maintains it, and knows what to do when a train stops where it should not.&lt;/p&gt;

&lt;p&gt;You can place two locomotives beside the same track and say that you have a choice of engines. That does not mean you can move the same set of rail cars to another railway and expect the journey to work unchanged.&lt;/p&gt;

&lt;p&gt;The engine may be replaceable. The journey still depends on the system around it that makes up the locomotive, the switches, the signals, and the layout of the tracks.&lt;/p&gt;

&lt;p&gt;Those model railways offer a useful parable for Enterprise AI. Companies are adding second and third models, putting gateways in front of them, and declaring themselves vendor independent.&lt;/p&gt;

&lt;p&gt;What they have usually gained is model choice.&lt;/p&gt;

&lt;p&gt;Whether they have gained AI portability is a much harder question.&lt;/p&gt;

&lt;h2 id=&quot;the-model-is-the-visible-part&quot;&gt;The Model Is the Visible Part&lt;/h2&gt;

&lt;p&gt;A multi-model gateway can route requests to models from OpenAI, Anthropic, Google, or an open-weight provider. It can centralize authentication, monitor usage, apply cost controls, ensure DLP policies are adhered to and redirect traffic when one provider is unavailable.&lt;/p&gt;

&lt;p&gt;That is useful. In many enterprises, it is also necessary.&lt;/p&gt;

&lt;p&gt;But it is not the same as moving a business workflow.  Model portability means that an application can call another model.  AI portability means that another model can complete the same business task without unacceptable losses in quality, permissions, observability, recovery, regulatory control, or cost.  The first is an interface problem.  The second is a systems problem.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://engineering.zalando.com/posts/2026/08/agentic-engineering-at-zalando-a-snapshot.html&quot;&gt;Zalando&lt;/a&gt;offers a useful example. Its engineering platform gives more than 250 teams access to models through a LiteLLM-based proxy connected to OpenAI, Amazon Bedrock, and Google Vertex. Their proxy centralizes adoption measurements, cost tracking, caching, client configuration, and model retirement.&lt;/p&gt;

&lt;p&gt;But that proxy is only part of the story.&lt;/p&gt;

&lt;p&gt;Around it, Zalando has developed authentication injection, Model Context Protocol access, shared agent skills, configuration management, and other supporting services. The company is also working on an identity broker, token vault, agent sandboxing, and model routing.&lt;/p&gt;

&lt;p&gt;Even then, it reports that users become attached to particular coding tools and model styles. They rarely switch models unless a limit or error forces the issue. This is not a criticism of Zalando. It is evidence of how the work really evolves.&lt;/p&gt;

&lt;p&gt;The gateway solved a lot of problems.  It centralized and simplified access.  Adoption created a platform. The platform created new operating responsibilities.  Human habits created a dependency.&lt;/p&gt;

&lt;p&gt;And dependencies that can not disappear, create risk.&lt;/p&gt;

&lt;h2 id=&quot;the-harness-is-part-of-the-product&quot;&gt;The Harness Is Part of the Product&lt;/h2&gt;

&lt;p&gt;I have argued before that an AI harness gives a model tools, memory, permissions, verification, and recovery. It turns a model that can discuss work into a system that can participate in the work. &lt;a href=&quot;https://tedt.org/Why-AI-Needs-a-Harness/&quot;&gt;Why AI Needs a Harness&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What is now clear is that the harness determines how replaceable the model really is.&lt;/p&gt;

&lt;p&gt;NVIDIA recently reported that its Agentic Variation Operators system completed all 183 levels in the public ARC-AGI-3 set. The system paired a frontier model with persistent memory, tools, supervision, feedback, and recovery mechanisms.&lt;/p&gt;

&lt;p&gt;The important point was not simply the benchmark score. NVIDIA argued that long-running performance belonged to the complete agent system, not to the model alone. &lt;a href=&quot;https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/&quot;&gt;NVIDIA’s AVO report&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That result requires repeatability.  This was vendor-reported, and only applies to the public benchmark set.  It was not a controlled measurement of the harness’s individual contribution. NVIDIA acknowledges those limitations.&lt;/p&gt;

&lt;p&gt;Still, the architectural lesson is useful and a warning for enterprises.&lt;/p&gt;

&lt;p&gt;If memory, tools, instructions, evaluation, and recovery materially affect the result, replacing the model while changing nothing else may be impossible. A prompt tuned for one model may fail with another. A tool-calling pattern may behave differently. A safety filter may block a previously accepted workflow. A replacement model may produce the correct answer while losing the evidence required to trust it.&lt;/p&gt;

&lt;p&gt;Performance and portability belong to the whole system and so does risk.&lt;/p&gt;

&lt;h2 id=&quot;what-must-the-enterprise-own&quot;&gt;What Must the Enterprise Own?&lt;/h2&gt;

&lt;p&gt;Does this mean the enterprise must build and own the entire AI stack?&lt;/p&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;An enterprise does not need to own every model, user interface, orchestration framework, or piece of infrastructure. It does need to control the assets that define successful work.&lt;/p&gt;

&lt;p&gt;The first is the workflow contract. This describes what the task is, which inputs it requires, what the agent may do, what it may not do, what the output should contain, and how the organization knows the work is complete.&lt;/p&gt;

&lt;p&gt;The second is the evaluation set. This includes representative cases, known edge conditions, failure examples, adversarial tests, acceptance thresholds, and situations requiring human review. If a vendor owns the only credible test of the system, the enterprise cannot independently determine whether a replacement works.&lt;/p&gt;

&lt;p&gt;The third is identity and authorization. An agent acting for an employee, customer, or business process must carry the correct permissions and delegation history. Those controls cannot quietly disappear when the model changes.&lt;/p&gt;

&lt;p&gt;The fourth is operational evidence. The enterprise needs usable traces showing what information the system accessed, which tools it called, what actions it attempted, what failed, and how it produced the final result.&lt;/p&gt;

&lt;p&gt;Finally, the enterprise must own the recovery process. The business must decide when the system stops, when a person intervenes, what gets rolled back, what gets retried, and what evidence is retained.&lt;/p&gt;

&lt;p&gt;These assets form the trusted action surface around AI. That is where much of the durable enterprise value now lives. &lt;a href=&quot;https://tedt.org/AI-Value-Stream/&quot;&gt;The AI Value Stream&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The model supplies capability.&lt;/p&gt;

&lt;p&gt;The enterprise must retain control over meaning, authority, and proof.&lt;/p&gt;

&lt;h2 id=&quot;how-do-you-measure-a-switch&quot;&gt;How Do You Measure a Switch?&lt;/h2&gt;

&lt;p&gt;Portability should be tested, not declared or assumed. If the workflow is critical to the company the enterprise should have metrics around how much it costs to move.&lt;/p&gt;

&lt;p&gt;Choose one bounded production workflow. Establish a baseline using the current model and platform. Measure whether the system completes the task correctly, how often a person must intervene, how long the work takes, what it costs, and what happens when something goes wrong.&lt;/p&gt;

&lt;p&gt;Then move that workflow to another model or provider.&lt;/p&gt;

&lt;p&gt;Do not stop when the new endpoint returns a response. Record the complete portability delta:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Migration time:&lt;/strong&gt; How long did the change take, including security review, integration, testing, deployment, and employee preparation?&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Engineering effort:&lt;/strong&gt; How many instructions, skills, tool definitions, integrations, and recovery routines had to be rewritten?&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Quality change:&lt;/strong&gt; Did completion rates, factual errors, policy violations, or human corrections improve or decline?&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Control loss:&lt;/strong&gt; Were permissions, regional restrictions, logs, citations, retention rules, and rollback capabilities preserved?&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Economic change:&lt;/strong&gt; Did the cost of a completed and accepted result improve, or did a lower model price create more retries and human review?&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Behavioral change:&lt;/strong&gt; Could employees use the replacement effectively, or had the existing tool become part of how they understood the work?&lt;/p&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The organization should not compress these measures into a single portability score. It needs enough evidence to see where the dependency actually resides.&lt;/p&gt;

&lt;p&gt;It may reside in the model. It may reside in a proprietary tool interface, a vendor-specific memory system, a skill library, an evaluation service, or the accumulated habits of the workforce.&lt;/p&gt;

&lt;p&gt;The economic measure is especially important. Token prices are easy to compare because vendors publish them. The meaningful unit is the cost of a finished job. &lt;a href=&quot;https://tedt.org/The-Cost-of-a-Finished-Job/&quot;&gt;The Cost of a Finished Job&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A cheaper engine does not lower the cost of the journey if every car must be rebuilt before it can move.&lt;/p&gt;

&lt;h2 id=&quot;when-is-an-internal-harness-worth-building&quot;&gt;When Is an Internal Harness Worth Building?&lt;/h2&gt;

&lt;p&gt;An internal harness makes sense when the workflow differentiates the business, crosses several enterprise systems, handles sensitive or regulated information, or operates at enough scale to justify a shared platform.&lt;/p&gt;

&lt;p&gt;It may also be necessary when the enterprise needs to move between providers or regions for resilience, sovereignty, negotiating leverage, or legal compliance.&lt;/p&gt;

&lt;p&gt;But an internal harness is not automatically the mature answer.&lt;/p&gt;

&lt;p&gt;For standardized work, an integrated vendor platform may be the better decision. &lt;a href=&quot;https://www.salesforce.com/news/press-releases/2026/08/26/salesforce-and-anthropic-announce-claudeforce/&quot;&gt;Salesforce and Anthropic&lt;/a&gt;, for example, are combining Claude with Salesforce data, permissions, business rules, workflows, and prebuilt skills. Salesforce says authentication and permissions can be managed centrally while actions continue to pass through Salesforce controls.&lt;/p&gt;

&lt;p&gt;Some of those capabilities remain in pilot or planned availability, and the claims come from the vendors themselves. Still, the attraction is easy to understand especially if you are already invested heavily into Salesforce and believe that the vendor’s controls meet your business needs.&lt;/p&gt;

&lt;p&gt;If the business process already lives inside one platform, the integrated option may reduce implementation time, preserve an established security boundary, and provide a more coherent user experience.&lt;/p&gt;

&lt;p&gt;The decision therefore depends on the value of the option.&lt;/p&gt;

&lt;p&gt;Build or configure with the vendor’s platform or retain an enterprise-controlled harness so that when you switch suppliers, you can preserve specialized workflows, or control critical business actions has material value.&lt;/p&gt;

&lt;p&gt;Choose the integrated platform when the workflow is standard, the provider’s controls meet the business need, and the organization would gain little from operating another internal platform.&lt;/p&gt;

&lt;p&gt;Portability is not a moral virtue. It is an option that will cost you.&lt;/p&gt;

&lt;p&gt;The enterprise should purchase that option when the likely cost of disruption, regulatory change, supplier failure, or strategic dependence exceeds the cost of maintaining the vendor / model lock-in.&lt;/p&gt;

&lt;h2 id=&quot;what-belongs-in-the-contract&quot;&gt;What Belongs in the Contract?&lt;/h2&gt;

&lt;p&gt;If a model or platform is supposed to be replaceable, the commercial agreement should preserve the evidence and assets needed to replace it.&lt;/p&gt;

&lt;p&gt;That includes:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Access to prompts, skills, workflow configurations, and evaluation results&lt;/li&gt;
  &lt;li&gt;Export of logs, traces, business state, and system-generated records&lt;/li&gt;
  &lt;li&gt;Ownership and permitted reuse of derived artifacts&lt;/li&gt;
  &lt;li&gt;Notice before model retirement or material behavioral changes&lt;/li&gt;
  &lt;li&gt;Access to fixed model versions when reproducibility matters&lt;/li&gt;
  &lt;li&gt;Clear rules for data retention, model training, subprocessors, and regional processing&lt;/li&gt;
  &lt;li&gt;Tool-call and application programming interface compatibility commitments&lt;/li&gt;
  &lt;li&gt;The right to test alternative models against enterprise evaluation sets&lt;/li&gt;
  &lt;li&gt;Transition assistance when the contract ends&lt;/li&gt;
  &lt;li&gt;Rollback procedures when an upgrade changes workflow behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A contract cannot make two models behave identically.&lt;/p&gt;

&lt;p&gt;It can prevent the enterprise from discovering, too late, that the information needed to perform a migration belongs to the supplier.&lt;/p&gt;

&lt;p&gt;It should be noted that if a vendor refuses to provide these rights, the enterprise may have to accept that the model is not replaceable and that the business process is now dependent on a single supplier.  This actually might be that vendor’s strategy, and it is not necessarily a bad one for either party.  It is just a business decision that you need to make for your enterprise and completely understand before it invests in a lock-in.&lt;/p&gt;

&lt;h2 id=&quot;when-does-the-harness-become-the-new-risk&quot;&gt;When Does the Harness Become the New Risk?&lt;/h2&gt;

&lt;p&gt;There is one more complication.&lt;/p&gt;

&lt;p&gt;An enterprise may reduce its dependence on a model provider and create a new dependency on its own internal platform.&lt;/p&gt;

&lt;p&gt;The warning signs are familiar:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;One person or small team understands the gateway.&lt;/li&gt;
  &lt;li&gt;No product owner is accountable for the platform.&lt;/li&gt;
  &lt;li&gt;Evaluations are stale or maintained separately by each business unit.&lt;/li&gt;
  &lt;li&gt;Model retirements trigger emergency migrations.&lt;/li&gt;
  &lt;li&gt;Costs can be traced to requests but not to accepted business outcomes.&lt;/li&gt;
  &lt;li&gt;Permissions and agent identities work differently across tools.&lt;/li&gt;
  &lt;li&gt;Teams copy prompts and skills because the shared service is unreliable.&lt;/li&gt;
  &lt;li&gt;No one has completed a provider-switching exercise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At that point, the harness is no longer strategic leverage. It is an underfunded internal product sitting in the path of critical work.  Many times this problem is multiplied by the fact that the enterprise will have different harnesses from different internal teams, and each of those harnesses will have different levels of maturity, different levels of documentation, and different levels of support.  The result is a patchwork of internal platforms that are not portable, not reliable, and not well understood.&lt;/p&gt;

&lt;p&gt;Adding another abstraction layer on top of this will not fix your problem.&lt;/p&gt;

&lt;p&gt;Each harness needs a roadmap, service objectives, security engineering, evaluation management, incident procedures, dedicated funding, and a clear boundary between platform responsibilities and business-workflow responsibilities.&lt;/p&gt;

&lt;p&gt;Enterprises already know how to operate identity platforms, integration platforms, data platforms, and developer platforms. An AI harness belongs to the same family.&lt;/p&gt;

&lt;p&gt;This means that these harnesses must be governed with similar seriousness at the enterprise architecture, operational, and business levels.&lt;/p&gt;

&lt;h2 id=&quot;moving-the-cargo&quot;&gt;Moving the Cargo&lt;/h2&gt;

&lt;p&gt;The practical test of AI portability is not whether another model appears in the company’s catalog.  It’s also not whether the model can be called through a gateway.  It is whether the same business task can be completed with another model without unacceptable losses in quality, permissions, observability, recovery, regulatory control, or cost.&lt;/p&gt;

&lt;p&gt;Again to validate this, choose one important workflow and move it.&lt;/p&gt;

&lt;p&gt;Use another model, another provider, or another region. Measure what had to be rebuilt. Measure how quality changed. Determine which controls survived, how people adapted, and what it cost to deliver an accepted result.&lt;/p&gt;

&lt;p&gt;My dad took a boy who loved trains to see model railways at the LA County Fair and at a train store in Pasadena. Years later, those layouts offer another lesson.&lt;/p&gt;

&lt;p&gt;The locomotive was the part that drew attention.&lt;/p&gt;

&lt;p&gt;The track, switches, signals, power, controls, and maintenance were what made the journey possible.&lt;/p&gt;

&lt;p&gt;An enterprise that counts models is still looking at the locomotives.&lt;/p&gt;

&lt;p&gt;An enterprise that tests whether the same business task can travel safely across multiple providers is actually planning on running a railway responsibly.&lt;/p&gt;

&lt;p&gt;The model may pull the train.&lt;/p&gt;

&lt;p&gt;The enterprise system determines whether the cargo arrives.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="enterprise AI" />
      
        <category term="AI portability" />
      
        <category term="model portability" />
      
        <category term="multi-model gateways" />
      
        <category term="AI agent harness" />
      
        <category term="AI governance" />
      
        <category term="vendor lock-in" />
      
        <category term="workflow portability" />
      
        <category term="AI operating model" />
      
        <category term="model migration" />
      
        <category term="platform engineering" />
      
        <category term="technology contracts" />
      

      
        <summary type="html">Large enterprises are learning how to switch models. Moving a working business process without losing quality, controls, or money is a different problem.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-08/The-System-Behind-the-Engine.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">Dev, Test, and Prod Still Matter: What Gets Deployed Has Changed</title>
      <link href="https://tedt.org/Dev-Test-and-Prod-Still-Matter-What-Gets-Deployed-Has-Changed/" rel="alternate" type="text/html" title="Dev, Test, and Prod Still Matter: What Gets Deployed Has Changed" />
      <published>2026-08-22T08:00:00-07:00</published>
      <updated>2026-08-22T08:00:00-07:00</updated>
      <id>https://tedt.org/Dev-Test-and-Prod-Still-Matter-What-Gets-Deployed-Has-Changed</id>
      <content type="html" xml:base="https://tedt.org/Dev-Test-and-Prod-Still-Matter-What-Gets-Deployed-Has-Changed/">&lt;h2 id=&quot;sources-for-this-article&quot;&gt;Sources for this article&lt;/h2&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Source note:&lt;/strong&gt; This essay draws on first-party disclosures from &lt;a href=&quot;https://openai.com/index/hugging-face-model-evaluation-security-incident/&quot;&gt;OpenAI&lt;/a&gt;, &lt;a href=&quot;https://huggingface.co/blog/agent-intrusion-technical-timeline&quot;&gt;Hugging Face&lt;/a&gt;, &lt;a href=&quot;https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals&quot;&gt;Anthropic&lt;/a&gt;, the &lt;a href=&quot;https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing&quot;&gt;U.K. AI Security Institute&lt;/a&gt;, and &lt;a href=&quot;https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward&quot;&gt;Irregular&lt;/a&gt;. Their investigations remain incomplete. As of Aug. 22, 2026, OpenAI’s &lt;a href=&quot;https://openai.com/index/hugging-face-model-evaluation-security-incident/&quot;&gt;promised technical report and METR/Redwood assessment&lt;/a&gt;, Anthropic’s &lt;a href=&quot;https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals&quot;&gt;planned METR review&lt;/a&gt;, AISI’s &lt;a href=&quot;https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing&quot;&gt;planned METR review&lt;/a&gt;, and Irregular’s &lt;a href=&quot;https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward&quot;&gt;planned evaluation-security white paper&lt;/a&gt; had not been published. &lt;a href=&quot;https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward&quot;&gt;Irregular says several model-provider disclosures describe the same underlying evaluation issue&lt;/a&gt;, so they should not be counted as separate incidents. These events occurred in specialized cybersecurity tests, often with normal safeguards reduced or disabled; &lt;a href=&quot;https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals&quot;&gt;Anthropic&lt;/a&gt; and &lt;a href=&quot;https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing&quot;&gt;AISI&lt;/a&gt; both caution against treating them as evidence of how ordinary deployed agents behave.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Many enterprise delivery diagrams begin with three boxes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;DEV → TEST → PROD&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We build in the first. We test throughout, but assemble the release case in the second. We deliberately let the system affect the business in the third. Each environment should carry more consequence than the one before it, and each arrow should mark a deliberate change in exposure.&lt;/p&gt;

&lt;p&gt;It is a useful picture. It gave software teams a practical way to separate unfinished work from customer-facing systems. It created places for experimentation, verification, approval, and rollback. It also carried a quiet assumption: work in Test cannot affect Production until somebody promotes it.&lt;/p&gt;

&lt;p&gt;That assumption was never entirely safe. It was a simple way of talking about a complex idea. Agentic systems make the weakness harder to ignore.&lt;/p&gt;

&lt;p&gt;An AI agent running in a test environment can call an application programming interface, install a package, create an account, open a pull request, send a message, query a live service, or cause another computer to act on its behalf. The agent does not have to be deployed to Production. It only needs a path to something real.&lt;/p&gt;

&lt;p&gt;The old diagram still captures an important control. But the environment label is not the boundary. We have to ask what the arrow represents. We also have to notice which changes happen without crossing it.&lt;/p&gt;

&lt;h2 id=&quot;what-the-old-model-gets-right&quot;&gt;What the Old Model Gets Right&lt;/h2&gt;

&lt;p&gt;Dev, Test, and Prod are not merely server names. At their best, they describe a progression of responsibility.&lt;/p&gt;

&lt;p&gt;Dev is where consequence should be smallest. Test is where pre-release evidence should become decision-grade. Prod is where durable business authority is granted. It is also where that authority must remain observable and revocable.&lt;/p&gt;

&lt;p&gt;Some organizations insert integration, quality assurance, user acceptance testing, staging, shadow, or deployment rings. The number of boxes is not the point. At each step, the system should earn a larger exposure.&lt;/p&gt;

&lt;p&gt;That progression still matters. In &lt;a href=&quot;https://tedt.org/What-IT-Looks-Like-in-an-Enterprise-Where-AI-Is-Assumed/&quot;&gt;What IT Looks Like in an Enterprise Where AI Is Assumed&lt;/a&gt;, I used lab, pilot, and production as distinct places for uncertainty, bounded learning, and operated value. The labels change from company to company, but the intent is familiar: do not expose the business to a new system until the system has earned that exposure.&lt;/p&gt;

&lt;p&gt;Modern delivery pipelines already move more than application code. They carry configuration, infrastructure templates, database changes, secrets references, policies, and deployment instructions. We learned a long time ago that a release can fail even when the binary doesn’t change.&lt;/p&gt;

&lt;p&gt;Conventional delivery already runs active code in continuous-integration jobs, integration tests, package registries, and deployment pipelines. Those systems can reach Production too. What changes with an agent is the degree of initiative inside the boundary. Its harness can inspect a result, choose another tool, and chain actions that nobody described in advance.&lt;/p&gt;

&lt;p&gt;AI did not invent porous environments. It lets software discover and use paths that ordinary release controls may never have modeled.&lt;/p&gt;

&lt;h2 id=&quot;the-artifact-is-no-longer-the-whole-release&quot;&gt;The Artifact Is No Longer the Whole Release&lt;/h2&gt;

&lt;p&gt;A model produces output. A harness lets that model do work.&lt;/p&gt;

&lt;p&gt;The harness supplies tools, memory, instructions, identity, network access, retry logic, and a way to observe what happened. I explored that system in &lt;a href=&quot;https://tedt.org/Why-AI-Needs-a-Harness/&quot;&gt;Why AI Needs a Harness&lt;/a&gt;. Once the model can inspect a result and choose another action, we are no longer testing only a static artifact. We are evaluating a bounded decision-and-action system that can change the environment around it.&lt;/p&gt;

&lt;p&gt;That changes what an environment means.&lt;/p&gt;

&lt;p&gt;For ordinary software, we often ask, “Where is the application running?” For an agent, we also have to ask:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;What can it read?&lt;/li&gt;
  &lt;li&gt;What can it change?&lt;/li&gt;
  &lt;li&gt;Which other systems can act for it?&lt;/li&gt;
  &lt;li&gt;What survives after the run ends?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A package proxy may live inside the test architecture while fetching material from the public internet. A test credential may still call a live software-as-a-service endpoint. A shared repository may let one run leave state for another. A browser agent may be operating in a test tenant while sending requests to real people and public websites.&lt;/p&gt;

&lt;p&gt;The environment name tells us where the process started. It does not tell us where the consequences can end.&lt;/p&gt;

&lt;p&gt;This matters well beyond frontier-model laboratories. A coding agent may hold a GitHub token. A finance agent may read a production ledger while writing only to a test workbook. A service agent may work in a training queue but still send email. A data agent may query a copy of a warehouse while exporting results through a live connector. A security agent may be testing a fictional target through a network path that reaches a real one.&lt;/p&gt;

&lt;p&gt;In each case, the box says Test. The authority says something else.&lt;/p&gt;

&lt;h2 id=&quot;when-test-touched-production&quot;&gt;When Test Touched Production&lt;/h2&gt;

&lt;p&gt;In July 2026, an internal OpenAI cybersecurity evaluation supplied a stark example. OpenAI said the models did not have direct internet access. They could, however, install packages through an internally hosted Artifactory proxy and cache. The models found and exploited a previously unknown Artifactory vulnerability, obtained internet access, and eventually reached Hugging Face production infrastructure. &lt;a href=&quot;https://openai.com/index/hugging-face-model-evaluation-security-incident/&quot;&gt;OpenAI called its account preliminary&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/agent-intrusion-technical-timeline&quot;&gt;Hugging Face later reconstructed about 17,600 actions&lt;/a&gt; across roughly two and a half days inside its infrastructure. It reported that five datasets apparently connected to the benchmark were accessed, while finding no effect on other customer-facing models, datasets, Spaces, or packages. Hugging Face also said no unauthorized software-supply-chain change was shipped.&lt;/p&gt;

&lt;p&gt;I examined the mechanics of that incident in &lt;a href=&quot;https://tedt.org/The-Sandbox-Was-Not-the-Boundary/&quot;&gt;The Sandbox Was Not the Boundary&lt;/a&gt;. The point here is different. The models were still being evaluated. They had not been released as a customer product. Yet the evaluation already had a path to production consequences through a shared service, a public launch point, and live third-party infrastructure.&lt;/p&gt;

&lt;p&gt;Separate incidents were later disclosed, but they involved different mechanisms and should not be treated as one breach count.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals&quot;&gt;Anthropic reviewed 141,006 cybersecurity evaluation runs&lt;/a&gt; and reported three incidents involving six runs. A configuration problem left live internet access available even though the models had been told they were inside simulations. In one case, a model published a malicious package to the Python Package Index. Anthropic said the package was available for about an hour and ran on 15 real systems. In another, a model scanned roughly 9,000 targets before compromising one real application. Anthropic’s latest research model stopped after concluding that its target was real, an important piece of counterevidence.&lt;/p&gt;

&lt;p&gt;The &lt;a href=&quot;https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing&quot;&gt;U.K. AI Security Institute reported a separate event&lt;/a&gt;. This was not a sandbox escape. Internet access was intentionally enabled, and the model providers’ cyber classifiers were disabled so the institute could measure maximum capability. Across 122 runs, AISI found 19 out-of-scope actions in 10 runs, including an attempted malicious pull request and the creation of fake identities. A human maintainer rejected the code, and AISI found no resulting real-world harm.&lt;/p&gt;

&lt;p&gt;Several later headlines involving different model providers traced back to one Irregular evaluation scenario. &lt;a href=&quot;https://www.irregular.com/research/addressing-recent-incidents-ongoing-findings-and-path-forward&quot;&gt;Irregular has said those disclosures shared an underlying control problem&lt;/a&gt;. Counting every provider statement as a separate failure would exaggerate the evidence.&lt;/p&gt;

&lt;p&gt;These were unusually permissive cyber tests. They often used unreleased models, reduced refusals, disabled classifiers, or deliberately open internet access. They do not show that an ordinary enterprise assistant will break out of its environment.&lt;/p&gt;

&lt;p&gt;They show something narrower and more useful: a system labeled Test can create production risk as soon as it can cause an external effect.&lt;/p&gt;

&lt;h2 id=&quot;the-promotion-gate-has-four-different-jobs&quot;&gt;The Promotion Gate Has Four Different Jobs&lt;/h2&gt;

&lt;p&gt;The traditional pipeline promotes an application release. An agentic system makes the gate responsible for four connected dimensions: artifact, authority, state, and evidence.&lt;/p&gt;

&lt;p&gt;They do not cross the boundary in the same way. The tested artifact should be promoted unchanged and accompanied by verifiable provenance. Authority should be issued in the destination. Runtime state should be reset; any necessary migration should be deliberate and governed. Evidence should be independently protected and should keep accumulating after release.&lt;/p&gt;

&lt;h3 id=&quot;1-the-artifact&quot;&gt;1. The Artifact&lt;/h3&gt;

&lt;p&gt;The promotable artifact is a signed, versioned release assembly or deployment manifest. It identifies the application code, declared model or model version, system instructions, policy and tool definitions, connector definitions, and configuration and routing dependencies.&lt;/p&gt;

&lt;p&gt;It does not include destination credentials, identities, live endpoints, or secret values. Those are bound separately in each environment.&lt;/p&gt;

&lt;p&gt;Any one of these elements can change behavior. A new connector may expose an action the agent could not take yesterday. A revised system instruction may alter when it stops or asks for help. A provider update may change model behavior without a customer deploying new application code, depending on the service and versioning contract.&lt;/p&gt;

&lt;p&gt;The release record therefore has to identify the whole runnable assembly, not only the Git commit. Where the platform permits it, identify its components with immutable versions or digests and bind them to &lt;a href=&quot;https://slsa.dev/spec/v1.2/verifying-artifacts&quot;&gt;verifiable provenance&lt;/a&gt;. The goal is simple: know that the assembly tested is the assembly released.&lt;/p&gt;

&lt;h3 id=&quot;2-the-authority&quot;&gt;2. The Authority&lt;/h3&gt;

&lt;p&gt;Authority includes identity, credentials, network egress, tool scopes, transaction limits, approval rules, and the systems the agent may affect.&lt;/p&gt;

&lt;p&gt;Giving a test agent a live token creates production-equivalent exposure, even if no code moves. Outbound internet access, a connector pointed at a live tenant, or a tool expanded from read to write can create the same class of external consequence.&lt;/p&gt;

&lt;p&gt;Policy intent can be version-controlled and promoted with the release. The destination workload identity should be bound there, and short-lived credentials should be issued there rather than copied forward. Production authority should be scoped to the destination and independently revocable.&lt;/p&gt;

&lt;p&gt;This is why capability and authority have to remain separate. In &lt;a href=&quot;https://tedt.org/How-Much-Work-Can-Your-AI-Safely-Own/&quot;&gt;How Much Work Can Your AI Safely Own?&lt;/a&gt;, I argued that demonstrated capability does not automatically grant operating authority. The same rule belongs in the delivery pipeline. A system may be capable of acting before the organization has earned the right to let it act.&lt;/p&gt;

&lt;h3 id=&quot;3-the-state&quot;&gt;3. The State&lt;/h3&gt;

&lt;p&gt;State includes retrieval data, working memory, caches, queues, shared directories, previous tool results, and whatever remains for the next run.&lt;/p&gt;

&lt;p&gt;State can cross an environment boundary without anyone deploying software. A production document can enter a test retrieval index. A test run can leave a file that another run treats as instruction. A cached response can outlive the policy that permitted it. A refreshed knowledge source can change the answer even when the model, prompt, and code stay fixed.&lt;/p&gt;

&lt;p&gt;State needs lineage, retention rules, reset procedures, and regional handling appropriate to the data. Calling it “test data” does not make it synthetic, disposable, or harmless.&lt;/p&gt;

&lt;p&gt;Most runtime state should not be promoted at all. Test memory, sessions, queues, caches, shared files, and previous tool results should normally be reset or isolated. When a business dataset, retrieval index, or model checkpoint must move, treat it as a separate approved migration with versioning, integrity checks, provenance, retention rules, and a rollback plan.&lt;/p&gt;

&lt;h3 id=&quot;4-the-evidence&quot;&gt;4. The Evidence&lt;/h3&gt;

&lt;p&gt;Evidence includes evaluation results, provenance, monitoring baselines, rollback results, known limits, and operating telemetry. Approval and risk acceptance are decision records informed by that evidence. They are not evidence that the system is trustworthy.&lt;/p&gt;

&lt;p&gt;The agent should not be the only system that writes the work, designs the test, grades the result, and decides it is ready. Independent scenarios matter. Holdout tests matter. A real rollback rehearsal matters. So does evidence that the identity can be revoked and the action trail can be reconstructed under pressure.&lt;/p&gt;

&lt;p&gt;Evidence accumulates throughout the lifecycle. Its record should be independent of the agent and protected from alteration. A model update, tool change, new data source, expanded permission, or material drift can invalidate yesterday’s approval without a traditional deployment. Production telemetry, drift findings, interventions, incidents, and reassessments should append to the assurance record. Promotion is not a one-time ceremony. It is a continuing claim that the running system still deserves its authority.&lt;/p&gt;

&lt;h2 id=&quot;keep-the-environments-redefine-their-gates&quot;&gt;Keep the Environments, Redefine Their Gates&lt;/h2&gt;

&lt;p&gt;We do not need to throw away Dev, Test, and Prod. We need each environment to make a stronger promise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Dev should minimize consequence.&lt;/strong&gt; Use synthetic or carefully de-identified data. Deny external write paths by default. Give each run its own short-lived identity and disposable state. Replace live tools with mocks, simulators, or mediated services when the work does not require reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test should make pre-release evidence decision-grade.&lt;/strong&gt; Testing happens &lt;a href=&quot;https://airc.nist.gov/airmf-resources/airmf/appendices/app-a-descriptions-of-ai-actor-tasks/&quot;&gt;throughout the lifecycle&lt;/a&gt;; Test is where the release case should become credible. &lt;a href=&quot;https://csrc.nist.gov/glossary/term/verification&quot;&gt;Verification&lt;/a&gt; asks whether the system meets its specified requirements and controls. &lt;a href=&quot;https://csrc.nist.gov/glossary/term/validation&quot;&gt;Validation&lt;/a&gt; asks whether it is fit for its intended use in a representative deployment context. Agentic systems need both.&lt;/p&gt;

&lt;p&gt;Use realistic tasks, independent evaluations, separate identities, resettable state, and monitoring that can stop a run while it is happening. When live information is necessary, begin with read-only access and intercept external side effects. Test the route through proxies, package registries, domain name resolution, webhooks, queues, and third-party tools, not only the agent’s visible network interface.&lt;/p&gt;

&lt;p&gt;The record should let another person reconstruct what was tested, what failed, what was waived, and who accepted the remaining risk.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shadow mode should validate behavior without granting action.&lt;/strong&gt; Let the agent receive representative production inputs and propose what it would do without giving it permission to make the change. Compare its decisions with actual outcomes. Measure quality, latency, exceptions, cost, and security before expanding authority. That is the same lifecycle discipline I described in &lt;a href=&quot;https://tedt.org/The-Cost-of-a-Finished-Job/&quot;&gt;The Cost of a Finished Job&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prod should grant narrow, revocable authority.&lt;/strong&gt; Production is not the end of testing. &lt;a href=&quot;https://airc.nist.gov/airmf-resources/airmf/5-sec-core/&quot;&gt;Measurement and monitoring continue while the system operates&lt;/a&gt;, because live telemetry, drift, incidents, and rollback may show that yesterday’s evidence no longer applies.&lt;/p&gt;

&lt;p&gt;Expand exposure in stages where the work permits it. Use least privilege, per-action traces, transaction limits, policy checks, human approval where consequence requires it, live monitoring, and rehearsed rollback. The system needs a separate observer that can stop it. The agent should not control the only alarm, the only log, or the only switch that turns it off.&lt;/p&gt;

&lt;p&gt;Third-party evaluators belong inside this model. Their contracts should identify permitted network paths, cross-client isolation, data handling, incident-notification timing, evidence retention, audit rights, and responsibility when a test reaches a real system. “Independent evaluation” describes who performs the test. It does not prove that the test infrastructure is independent, isolated, or safe.&lt;/p&gt;

&lt;h2 id=&quot;five-questions-for-the-release-review&quot;&gt;Five Questions for the Release Review&lt;/h2&gt;

&lt;p&gt;The release meeting does not need another hundred-page checklist. It needs questions that follow the four different jobs of the gate: identify the assembly, bind its authority, govern its state, and judge the evidence.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;What changed?&lt;/strong&gt; Name the model, code, prompts, policies, tools, connectors, configuration, and data sources that differ from the last accepted version.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;What can it touch?&lt;/strong&gt; Trace direct and indirect access through identities, proxies, package services, webhooks, shared platforms, and third parties. Record what the agent may read, propose, or change.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;What survives?&lt;/strong&gt; Identify memory, files, logs, caches, queues, retrieval indexes, and artifacts that another run or environment can inherit.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Who tested it independently?&lt;/strong&gt; Show evidence from scenarios the builder did not control, including failure cases, stop behavior, permission boundaries, and rollback.&lt;/p&gt;
  &lt;/li&gt;
  &lt;li&gt;
    &lt;p&gt;&lt;strong&gt;Who can stop it and reconstruct the result?&lt;/strong&gt; Name the decision owner, revocation path, incident lead, evidence location, and conditions for resumption.&lt;/p&gt;
  &lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is the practical machinery behind &lt;a href=&quot;https://tedt.org/The-IT-Adaptive-Factory/&quot;&gt;The IT Adaptive Factory&lt;/a&gt;. The factory is not mature because agents can write and test code quickly. It is mature when the organization can explain what was promoted, what authority was granted, what state survived, why the evidence was sufficient, and how to recover when the answer was wrong.&lt;/p&gt;

&lt;p&gt;OpenAI’s later response shows the cost of learning this after a control boundary fails. On Aug. 18, &lt;a href=&quot;https://openai.com/index/pacing-model-development-cyber-capabilities/&quot;&gt;the company said it had paused some frontier research workloads&lt;/a&gt;, added stronger workload and network isolation, and left its largest planned reinforcement-learning run on hold while it gathered more evidence. OpenAI attributed the broader slowdown to both the Hugging Face incident and concerns about an upcoming model’s cyber capability, so the pause should not be assigned to the incident alone.&lt;/p&gt;

&lt;p&gt;That response does not prove the new controls will work. It does show that research and test infrastructure can become important enough to slow the production of the model itself.&lt;/p&gt;

&lt;h2 id=&quot;the-arrow-still-matters&quot;&gt;The Arrow Still Matters&lt;/h2&gt;

&lt;p&gt;Organizations will keep drawing three boxes. They should.&lt;/p&gt;

&lt;p&gt;The boxes remind us that consequence should rise slowly and evidence should rise first. But an environment label cannot enforce that discipline. Architecture does. Identity does. Network policy does. State isolation does. Independent testing does. Someone with the authority to say “stop” does.&lt;/p&gt;

&lt;p&gt;Before approving the next agent release, look at the arrow and ask four different questions. What artifact was promoted? What authority was bound in the destination? What state was reset or deliberately migrated? What evidence supported the decision, and how will that evidence continue after release?&lt;/p&gt;

&lt;p&gt;If a test agent can already cause a real-world change, the organization is already carrying production risk, whatever the box is called.&lt;/p&gt;

&lt;p&gt;The arrow still matters. It moves the assembly. The gate decides what authority it receives, what state it inherits, and whether the evidence is good enough to let it act.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="AI agents" />
      
        <category term="agentic AI" />
      
        <category term="software delivery" />
      
        <category term="DevSecOps" />
      
        <category term="environment promotion" />
      
        <category term="AI governance" />
      
        <category term="AI security" />
      
        <category term="least privilege" />
      
        <category term="state management" />
      
        <category term="enterprise architecture" />
      

      
        <summary type="html">Dev, Test, and Prod still matter for AI agents, but promotion gates must now govern the artifact, authority, state, and evidence together.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-08/Dev-Test-Prod-AI.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">The Cost of a Finished Job</title>
      <link href="https://tedt.org/The-Cost-of-a-Finished-Job/" rel="alternate" type="text/html" title="The Cost of a Finished Job" />
      <published>2026-08-16T09:00:00-07:00</published>
      <updated>2026-08-16T09:00:00-07:00</updated>
      <id>https://tedt.org/The-Cost-of-a-Finished-Job</id>
      <content type="html" xml:base="https://tedt.org/The-Cost-of-a-Finished-Job/">&lt;p&gt;&lt;em&gt;Cheaper models do not settle the economics of enterprise AI. They move the hard work into the system around the model and the way the business itself is designed.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Tokens are like gallons of gasoline. They are easy to count. They are easy to price. They are also poor at telling you whether you arrived.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/&quot;&gt;OpenAI recently cut the published price of GPT-5.6 Luna by 80%&lt;/a&gt;. That is a substantial change, but it does not make every workflow 80% cheaper. A low-cost model can still produce expensive work if it needs repeated attempts, makes too many tool calls, waits on slow systems, gives people more material to review, or produces an error that has to be repaired later.&lt;/p&gt;

&lt;p&gt;The useful unit is not the token. It is the finished job.&lt;/p&gt;

&lt;p&gt;Even that needs one more word. A job is not finished merely because the machine stopped. It is finished when the result clears a defined acceptance threshold: the claim is ready for a qualified decision, the code passes its tests and review, the employee receives the correct policy answer, or the customer problem is actually resolved.&lt;/p&gt;

&lt;p&gt;So the economic measure I would use is closer to this:&lt;/p&gt;

&lt;blockquote&gt;
  &lt;p&gt;&lt;strong&gt;Total workflow cost = AI usage + data and tools + human review + retries and rework + support + expected failure cost&lt;/strong&gt;&lt;br /&gt;
&lt;strong&gt;Cost per accepted outcome = total workflow cost ÷ accepted outcomes&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is why I have argued that &lt;a href=&quot;https://tedt.org/Make-AI-Boring/&quot;&gt;every AI capability needs a cost identity&lt;/a&gt;: an owner, a budget, and a unit the business understands, such as dollars per case, ticket, work order, approved document, or repaired defect.&lt;/p&gt;

&lt;p&gt;Price per token still matters. It belongs inside the equation. It simply does not get to be the equation.&lt;/p&gt;

&lt;h2 id=&quot;the-system-of-systems-controls-the-price&quot;&gt;The System of Systems Controls the Price&lt;/h2&gt;

&lt;p&gt;So how do you lower the cost of an accepted outcome?&lt;/p&gt;

&lt;p&gt;You can use a smaller model. You can route easy steps to a cheaper model and save the strongest model for the hard decisions. You can cache repeated information. You can shorten prompts. You can improve retrieval. You can reduce retries. You can give the model better tools and clearer tests.&lt;/p&gt;

&lt;p&gt;In other words, you work on the system around your systems and model.&lt;/p&gt;

&lt;p&gt;I have described &lt;a href=&quot;https://tedt.org/Why-AI-Needs-a-Harness/&quot;&gt;the harness as the structure that turns a fluent model into a useful partner&lt;/a&gt;. It carries context, connects tools, retains state, applies permissions, checks results, handles failure, and decides what happens next. A recent OpenAI experiment puts useful numbers behind that idea.&lt;/p&gt;

&lt;p&gt;On the public set of the ARC-AGI-3 benchmark, &lt;a href=&quot;https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/&quot;&gt;OpenAI reports that GPT-5.6 Sol scored 13.3% with the official harness and 38.3% after retained reasoning and context compaction were enabled&lt;/a&gt;. OpenAI also reports that the changed configuration used six times fewer output tokens. The reported improvement came from two harness settings, not a newly trained model.&lt;/p&gt;

&lt;p&gt;That result has limits. It is a vendor-run benchmark experiment, not evidence that every enterprise workflow will become three times better. Retaining more context can also raise privacy and data-retention questions. But the mechanism is important: a model that forgets what it has learned has to solve the same problem again. A model given continuity can spend less effort rediscovering its own path.&lt;/p&gt;

&lt;p&gt;Microsoft Research is testing the same broader principle from two other directions. &lt;a href=&quot;https://www.microsoft.com/en-us/research/blog/orchard-an-open-framework-for-scalable-agentic-ai/&quot;&gt;Orchard trains agents inside the deployment harnesses where they will actually work&lt;/a&gt;. Microsoft reports that an Orchard software-engineering model with about three billion active parameters reached 69.7% on SWE-bench Verified. &lt;a href=&quot;https://www.microsoft.com/en-us/research/blog/echoverse-deep-evolving-environments-for-computer-use-agents/&quot;&gt;Echoverse builds deep, stateful training environments&lt;/a&gt; where an action changes underlying data and the result can be graded against that data. Its researchers report raising a nine-billion-parameter model’s average score from 36.5% to 67.1% across their environments.&lt;/p&gt;

&lt;p&gt;Those are research results, not production economics from a claims department or call center. Still, they reveal where optimization is moving. Model choice matters, but the surrounding system shapes what the model can accomplish and how much the result costs.&lt;/p&gt;

&lt;p&gt;This is the larger &lt;a href=&quot;https://tedt.org/AI-Value-Stream/&quot;&gt;AI value stream&lt;/a&gt;: everything between employee intent and physical infrastructure contributes to whether intelligence becomes useful work.&lt;/p&gt;

&lt;p&gt;The performance of AI is not a property of the model alone. Neither is its cost.&lt;/p&gt;

&lt;p&gt;Make now mistake that cost per accepted outcome is the business model for AI vendors in the future.  Waht this means is that once cost per accepted outcome becomes the measure, routing, context, authority, evidence, and learning can no longer be optimized as separate concerns. That metric forces them to be managed as one operating layer.&lt;/p&gt;

&lt;h2 id=&quot;from-one-workflow-to-fifteen-hundred&quot;&gt;From One Workflow to Fifteen Hundred&lt;/h2&gt;

&lt;p&gt;Tuning one workflow is difficult. Letting thousands of employees create workflows is a different class of problem.&lt;/p&gt;

&lt;p&gt;In its customer account of the Dutch cooperative insurer Univé, &lt;a href=&quot;https://openai.com/index/unive/&quot;&gt;OpenAI reports 97% license activation, 85% weekly use, and roughly 1,500 custom agents created by employees&lt;/a&gt;. It also describes a pet-insurance workflow in which claim preparation moved from hours to minutes while a trained claims professional retained the final decision.&lt;/p&gt;

&lt;p&gt;Those numbers demonstrate adoption. They do not tell us whether every GPT solves a distinct problem or still deserves a place in the portfolio. Fifteen hundred custom agents can represent a remarkable reservoir of employee knowledge. They can also become 1,500 small applications looking for owners who can maintain them, evaluate them, and retire them when they no longer earn their place.&lt;/p&gt;

&lt;p&gt;Stripe describes encountering the same tension at another scale. Its teams had created &lt;a href=&quot;https://stripe.dev/blog/meet-stripes-knowledge-ai-platform&quot;&gt;more than 4,000 small agents before the company moved toward a shared Knowledge AI Platform&lt;/a&gt;. Stripe says the common platform now supports more than 1,000 tools and skills and reached 83% weekly workforce use. The important design choice was not to build one giant agent. It was to centralize the common machinery while allowing domain teams to retain responsibility for their own expertise.&lt;/p&gt;

&lt;p&gt;Cloudflare has followed a related path for internal engineering. The company says its &lt;a href=&quot;https://blog.cloudflare.com/internal-ai-engineering-stack/&quot;&gt;shared AI engineering stack served 3,683 internal users and recorded 47.95 million AI messages during one 30-day period&lt;/a&gt;. One central proxy provides authentication, model discovery, permission enforcement, and per-user cost attribution. Employees can use different models and local configurations without placing provider keys on their laptops or rebuilding the control plane for every team.&lt;/p&gt;

&lt;p&gt;These are first-party company reports, not independent return-on-investment studies. They are useful because they show the same architectural pressure appearing in different organizations. Once employees can build with AI, a company must support local invention without asking every employee to become an identity engineer, security architect, evaluation specialist, and platform operator.&lt;/p&gt;

&lt;p&gt;As I argued in &lt;a href=&quot;https://tedt.org/Beyond-the-Light-Bulb/&quot;&gt;Beyond the Light Bulb&lt;/a&gt;, issuing more licenses may improve local productivity without changing the end-to-end value stream. A report drafted in five minutes can still wait four days for approval. An agent can prepare twice as many cases and merely bury the reviewer. Faster output does not become business value until the surrounding work changes with it.&lt;/p&gt;

&lt;h2 id=&quot;what-i-mean-by-an-ai-operating-system&quot;&gt;What I Mean by an AI Operating System&lt;/h2&gt;

&lt;p&gt;This is where the operating-system comparison becomes useful.&lt;/p&gt;

&lt;p&gt;A conventional operating system does not write your document, approve your expense report, or calculate your budget. It manages the shared machinery that applications depend on.&lt;/p&gt;

&lt;p&gt;An enterprise AI operating system would perform a similar coordinating role for intelligence:&lt;/p&gt;

&lt;table class=&quot;well table table-striped&quot;&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th style=&quot;text-align: right&quot;&gt;Conventional Operating System Components&lt;/th&gt;
      &lt;th style=&quot;text-align: left&quot;&gt;Enterprise AI Operating Layer Components&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Processes and scheduling&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Jobs, agent runs, model selection, and tool routing&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Memory and files&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Working context, durable state, records, and retention&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Users and permissions&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Employee identity, agent identity, approvals, and least privilege&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Devices and drivers&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Governed connectors to business data and applications&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Health and logs&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Evaluations, traces, evidence, incidents, and recovery&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Resource quotas&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Budgets, rate limits, latency targets, and capacity&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td style=&quot;text-align: right&quot;&gt;Installation and upgrades&lt;/td&gt;
      &lt;td style=&quot;text-align: left&quot;&gt;Inventory, ownership, versions, reviews, and retirement&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;I do not mean that every company should buy a product labeled “AI OS,” or spend three years building one enormous platform. Most enterprises will accidentially assemble this from systems they already have: identity and access management, data platforms, API gateways, workflow engines, model services, observability, security controls, and software-delivery practices.&lt;/p&gt;

&lt;p&gt;Is this simply platform engineering, data governance, identity, and observability wearing a new name?&lt;/p&gt;

&lt;p&gt;Much of the machinery is familiar. What has changed is the thing being coordinated. These systems now surround probabilistic workers that can choose tools, generate intermediate plans, act across applications, and produce different results from the same request. The components are not all new. The need to coordinate them continuously around machine-performed work is.&lt;/p&gt;

&lt;p&gt;For IT, this means becoming &lt;a href=&quot;https://tedt.org/What-IT-Looks-Like-in-an-Enterprise-Where-AI-Is-Assumed/&quot;&gt;the central engineering function for a shared production system&lt;/a&gt;, not the sole builder of every employee solution. IT provides the roads: approved models, secure connectors, identity, sandboxes, evaluation services, logging, and production gates. Business teams choose the destinations. They define the job, the value, the exceptions, and the standard for acceptable work.&lt;/p&gt;

&lt;p&gt;The &lt;a href=&quot;https://airc.nist.gov/airmf-resources/airmf/5-sec-core/&quot;&gt;NIST AI Risk Management Framework organizes this work around four continuing functions: Govern, Map, Measure, and Manage&lt;/a&gt;. That is a helpful reminder that an operating layer is not merely a runtime. It must also carry responsibility, measurement, monitoring, and an exit path.&lt;/p&gt;

&lt;h2 id=&quot;what-this-means-for-businesses-giving-ai-to-employees&quot;&gt;What This Means for Businesses Giving AI to Employees&lt;/h2&gt;

&lt;p&gt;The first implication is that an employee AI program is not a license program.&lt;/p&gt;

&lt;p&gt;Licenses provide access. Training can provide competence. Neither one, by itself, changes a business process. The business has to decide which work should change, who owns the outcome, what quality means, and where the recovered time or capacity will go.&lt;/p&gt;

&lt;p&gt;Imagine an employee who spends two hours assembling a weekly report. AI reduces the drafting to twenty minutes. That sounds like a large gain. But what happens next?&lt;/p&gt;

&lt;p&gt;Does the report still wait for the same approval? Does the employee use the recovered time to investigate exceptions, speak with customers, or improve the source data? Does the organization simply produce six times as many reports that nobody reads? Time saved is potential energy. Management still has to decide where it moves.&lt;/p&gt;

&lt;p&gt;The second implication is that adoption and value need different scorecards. A useful evidence ladder looks something like this:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;People have access.&lt;/li&gt;
  &lt;li&gt;People use it.&lt;/li&gt;
  &lt;li&gt;A workflow changes.&lt;/li&gt;
  &lt;li&gt;Cycle time, quality, or service improves.&lt;/li&gt;
  &lt;li&gt;The total cost per accepted outcome improves.&lt;/li&gt;
  &lt;li&gt;The business converts that improvement into safer and reliabile operations, principled decisions, higher standards, continuous improvement, stronger teamwork, revenue, or customer value.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each step matters. They are not interchangeable. Prompt counts and weekly active users tell you whether the system is being used. They do not tell you whether the business is better off in the executive suite.&lt;/p&gt;

&lt;p&gt;The third implication is that the foundation should be shared while improvement remains local. Fully centralized AI programs turn into queues because the central team cannot understand every job. Fully decentralized programs produce duplicate agents, inconsistent permissions, uneven quality, and abandoned tools. The better pattern is federated: common roads, local destinations, and an explicit exception path.&lt;/p&gt;

&lt;p&gt;The fourth implication is that employees need more than permission to experiment. They need time, examples, trusted data, clear boundaries, and managers willing to redesign the work. None of this is merely a platform change. &lt;a href=&quot;https://tedt.org/AI-Is-a-People-Change/&quot;&gt;AI is a people change&lt;/a&gt; touching roles, incentives, professional identity, training, trust, and the way expertise is passed from one generation of employees to the next.&lt;/p&gt;

&lt;p&gt;This is also where telemetry can cross a line. The operating layer can reveal which employees use AI, which sources they touch, how long tasks take, and where they struggle. That information can improve the system and control risk. Used carelessly, it can become an unusually intimate system of worker surveillance. Businesses should tell employees what is recorded, why it is recorded, who can see it, and how long it is retained.&lt;/p&gt;

&lt;p&gt;The fifth implication is that “human in the loop” is not enough. A person cannot meaningfully own a decision if the review queue is impossible, the evidence is hidden, or rejecting the machine’s recommendation requires more time than accepting it. Human accountability needs enough context, time, expertise, authority, and a real ability to stop or reverse the action.&lt;/p&gt;

&lt;p&gt;Finally, employee-built AI creates a lifecycle problem. Every production workflow needs a business owner, a technical owner, a risk classification, a version history, a budget, an evaluation set, a review date, and a retirement path. Otherwise, cheaper creation produces an expanding estate of invisible software.&lt;/p&gt;

&lt;h2 id=&quot;advice-for-the-people-building-this-now&quot;&gt;Advice for the People Building This Now&lt;/h2&gt;

&lt;p&gt;If you are working on employee AI, I would use the following promotion path.&lt;/p&gt;

&lt;h3 id=&quot;1-start-with-a-job-not-a-model&quot;&gt;1. Start with a job, not a model&lt;/h3&gt;

&lt;p&gt;Find a repeatable piece of work with a visible owner, inputs, handoffs, exceptions, and an outcome. Ask &lt;a href=&quot;https://tedt.org/Before-You-Reach-for-AI-Ask-What-Job-Needs-Doing/&quot;&gt;what job needs doing&lt;/a&gt; before deciding whether the answer is a chatbot, an agent, a search tool, or ordinary automation.&lt;/p&gt;

&lt;p&gt;“Give everyone an AI assistant” is a distribution plan. It is not a use case.&lt;/p&gt;

&lt;h3 id=&quot;2-establish-the-baseline&quot;&gt;2. Establish the Baseline&lt;/h3&gt;

&lt;p&gt;Measure the work before changing it: elapsed time, employee effort, error rate, rework, wait time, exception volume, service level, and current cost. If you do not know the starting point, almost any demonstration can be made to look like progress.&lt;/p&gt;

&lt;h3 id=&quot;3-define-what-accepted-means&quot;&gt;3. Define What Accepted Means&lt;/h3&gt;

&lt;p&gt;Capture the job as &lt;a href=&quot;https://tedt.org/How-to-Communicate-in-a-World-of-AI/&quot;&gt;a reusable specification&lt;/a&gt;: its purpose, authoritative inputs, boundaries, desired outcome, and the evidence that will count as complete.&lt;/p&gt;

&lt;p&gt;Build a small evaluation set from real work. Include ordinary cases, difficult edge cases, and cases that should be refused or escalated. OpenAI is retiring its Evals platform later in 2026, but its published &lt;a href=&quot;https://developers.openai.com/api/docs/guides/evaluation-best-practices&quot;&gt;guidance on task-specific tests, logging, human calibration, and continuous evaluation&lt;/a&gt; remains useful.&lt;/p&gt;

&lt;h3 id=&quot;4-separate-assistance-from-authority&quot;&gt;4. Separate Assistance from Authority&lt;/h3&gt;

&lt;p&gt;Let the first version retrieve, organize, summarize, recommend, or draft. Give it permission to change records, send messages, approve transactions, or trigger physical work only after the evidence justifies the additional authority.&lt;/p&gt;

&lt;p&gt;The practical question is &lt;a href=&quot;https://tedt.org/How-Much-Work-Can-Your-AI-Safely-Own/&quot;&gt;how much work the AI can safely own&lt;/a&gt;, one value stream, work class, and risk tier at a time.&lt;/p&gt;

&lt;h3 id=&quot;5-optimize-the-workflow-not-merely-the-prompt&quot;&gt;5. Optimize the Workflow, not Merely the Prompt&lt;/h3&gt;

&lt;p&gt;Choose the least expensive model that clears the quality threshold for each step. Improve retrieval. Remove unnecessary context. Cache stable material. Reduce tool calls. Tighten retry rules. Make human handoffs easier. Sometimes a more capable model will cost less overall because it needs fewer retries and less review.&lt;/p&gt;

&lt;p&gt;The cheapest successful run is not always the correct target. Reliability, privacy, security, legal obligations, and human accountability are constraints, not items to trade away for a lower bill.&lt;/p&gt;

&lt;h3 id=&quot;6-make-the-evidence-visible&quot;&gt;6. Make the Evidence Visible&lt;/h3&gt;

&lt;p&gt;Record the sources used, tools called, actions taken, checks passed, exceptions raised, and human decisions made. A person reviewing the work should be able to understand what happened without reconstructing the agent’s entire hidden journey.&lt;/p&gt;

&lt;p&gt;Track human-review minutes along with model cost. If a cheaper configuration saves ten dollars of inference and adds an hour of expert review, the workflow did not get cheaper.&lt;/p&gt;

&lt;h3 id=&quot;7-promote-carefully&quot;&gt;7. Promote Carefully&lt;/h3&gt;

&lt;p&gt;Run consequential workflows in shadow mode before giving them authority. Compare the proposed result with real decisions. Move from experiment to shared service only after the workflow meets agreed thresholds for quality, cost, latency, security, and human review.&lt;/p&gt;

&lt;p&gt;Then keep evaluating it. Models change. Prices change. Source data changes. Business rules change. A production agent is not a finished project. It is an operating responsibility.&lt;/p&gt;

&lt;h3 id=&quot;8-retire-what-no-longer-earns-its-place&quot;&gt;8. Retire What no Longer Earns its Place&lt;/h3&gt;

&lt;p&gt;Cheaper AI will create more AI. That is a rebound effect, not a paradox. When the price of an individual run falls, companies will attempt more workflows and run them more often. Total spending, data exposure, and operational risk can rise even while each call gets cheaper.&lt;/p&gt;

&lt;p&gt;Review the portfolio. Merge duplicate agents. Revoke unused permissions. Update evaluation sets. Replace poor configurations. Retire workflows whose owners have left or whose value never arrived.&lt;/p&gt;

&lt;h2 id=&quot;build-only-as-much-operating-system-as-you-need&quot;&gt;Build Only as Much Operating System as you Need&lt;/h2&gt;

&lt;p&gt;A small company does not need to recreate the platform of Stripe or Cloudflare. It may need an approved managed workspace, clean data boundaries, a handful of task evaluations, and a named owner.&lt;/p&gt;

&lt;p&gt;A business unit with several production workflows may need shared connectors, reusable skills, an inventory, common logs, and a review process.&lt;/p&gt;

&lt;p&gt;A large enterprise with agents acting across many systems will need policy-based routing, distinct agent identities, centralized evidence, incident handling, lifecycle controls, and portfolio economics.&lt;/p&gt;

&lt;p&gt;The operating layer should grow in response to real work and real risk. Building a cathedral before the first useful workflow is another way to avoid arriving.&lt;/p&gt;

&lt;h2 id=&quot;the-road-between-the-request-and-the-result&quot;&gt;The Road Between the Request and the Result&lt;/h2&gt;

&lt;p&gt;Cheaper intelligence moves the bottleneck. The difficult question becomes less about whether the company can afford a model and more about whether it can manage the system around that model.&lt;/p&gt;

&lt;p&gt;Who defines the job? Who grants the context? Who chooses the tools? Who decides what may be changed? Who checks the result? Who pays for the run, supports the workflow, investigates the failure, and retires the agent when it is no longer useful?&lt;/p&gt;

&lt;p&gt;These are operating-system questions.&lt;/p&gt;

&lt;p&gt;The businesses that do this well will not merely have more AI. They will know which completed jobs became cheaper, faster, or better, and they will be able to explain why.&lt;/p&gt;

&lt;p&gt;Tokens may tell us how much fuel we purchased. The finished job tells us whether the trip mattered and if those who took it were meaningfully and positiviely changed. The AI operating system is the machinery that helps the whole company arrive.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="enterprise AI" />
      
        <category term="AI economics" />
      
        <category term="cost per accepted outcome" />
      
        <category term="AI operating system" />
      
        <category term="agentic AI" />
      
        <category term="AI governance" />
      
        <category term="workflow optimization" />
      
        <category term="platform engineering" />
      
        <category term="employee AI" />
      
        <category term="operating model" />
      
        <category term="FinOps" />
      

      
        <summary type="html">Cheaper models do not automatically produce cheaper work. Enterprise AI economics depend on the total cost of delivering an accepted outcome.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-08/the-cost-of-a-finished-job-hero-source.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">The Sandbox Was Not the Boundary</title>
      <link href="https://tedt.org/slides/the-sandbox-was-not-the-boundary/" rel="alternate" type="text/html" title="The Sandbox Was Not the Boundary" />
      <published>2026-08-13T00:00:00-07:00</published>
      <updated>2026-08-13T00:00:00-07:00</updated>
      <id>https://tedt.org/slides/the-sandbox-was-not-the-boundary</id>
      <content type="html" xml:base="https://tedt.org/slides/the-sandbox-was-not-the-boundary/"></content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
          <email>ted@tschopp.org</email>
        
        
          <uri>https://tedt.org/profile/</uri>
        
      </author>

      

      

      
        <summary type="html"></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/slides/decks/the-sandbox-was-not-the-boundary/slide-preview.png" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">The Sandbox Was Not the Boundary</title>
      <link href="https://tedt.org/The-Sandbox-Was-Not-the-Boundary/" rel="alternate" type="text/html" title="The Sandbox Was Not the Boundary" />
      <published>2026-08-08T09:00:00-07:00</published>
      <updated>2026-08-08T09:00:00-07:00</updated>
      <id>https://tedt.org/The-Sandbox-Was-Not-the-Boundary</id>
      <content type="html" xml:base="https://tedt.org/The-Sandbox-Was-Not-the-Boundary/">&lt;blockquote class=&quot;alert alert-call-to-action&quot;&gt;
  &lt;p&gt;&lt;strong&gt;Source note:&lt;/strong&gt; This essay draws from first-party disclosures by OpenAI and Hugging Face and an August 2026 Black Hat USA presentation by OpenAI researchers Eric Wallace and Michael Dalton. Their reconstruction was preliminary, the investigation remained open, and the speakers said a fuller postmortem would follow. The chronology below should be read with that boundary in mind.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In computing, a sandbox is a controlled workspace where software can act, but only with the files, tools, and connections it has been given. Think of it as a locked room built for experimentation. An AI agent may use everything placed inside, but it should not be able to wander into the rest of the building or reach the internet. That is the idea, anyway. The trouble begins when some small service  still connects the room to the world outside.&lt;/p&gt;

&lt;p&gt;In this incident, that service was Artifactory.&lt;/p&gt;

&lt;p&gt;Artifactory is a shared, company-controlled storeroom for software. Instead of letting every program search the internet for the parts it needed, programs asked Artifactory. If a part was not already on its shelves, Artifactory could retrieve it from outside, keep a copy, and pass it back. In this incident, the agents also discovered they could leave files and names there for later agents to find. The storeroom became both a way to access the internet and an unintended bulletin board that allowed separate agents and test runs to communicate with each other over the course of their testing.&lt;/p&gt;

&lt;p&gt;In July 2026, &lt;a href=&quot;https://openai.com/index/hugging-face-model-evaluation-security-incident/&quot;&gt;OpenAI disclosed&lt;/a&gt; that models running an internal cyber-capability evaluation, including GPT-5.6 Sol, the latest and most capable model from OpenAI, and a more capable internal research prototype, had found and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure. The models did not have broad, direct internet access. They could, however, install packages through an internally hosted Artifactory proxy and cache, and that service could reach the internet.&lt;/p&gt;

&lt;p&gt;Hugging Face is a company and online platform where researchers and developers publish, share, and test artificial intelligence models, datasets, and software tools. People and organizations use it to find existing models, build new systems, and collaborate on machine-learning projects. In this incident, the agents eventually reached beyond OpenAI’s internal research environment and into systems operated by Hugging Face.&lt;/p&gt;

&lt;p&gt;&lt;a href=&quot;https://huggingface.co/blog/security-incident-july-2026&quot;&gt;Hugging Face’s disclosure&lt;/a&gt; and later &lt;a href=&quot;https://huggingface.co/blog/agent-intrusion-technical-timeline&quot;&gt;technical timeline&lt;/a&gt; provide another first-party view of the intrusion. OpenAI researchers Eric Wallace and Michael Dalton then presented a more detailed reconstruction in a &lt;a href=&quot;https://www.youtube.com/watch?v=87DyyMV0kCY&quot;&gt;Black Hat USA presentation&lt;/a&gt;. As of August 8, the public record is still preliminary, and the speakers said their investigation was incomplete. This chronology should be read as a current reconstruction, not a final forensic report.&lt;/p&gt;

&lt;h2 id=&quot;how-artifactory-connected-the-runs&quot;&gt;How Artifactory Connected the Runs&lt;/h2&gt;

&lt;p&gt;The chain began in May, when OpenAI started training new models.  It occurred across multiple tests and multiple test runs. One agent was asked to complete an Excel workbook whose missing formulas depended on data behind Google Drive links. The environment could not reach those links, so the task could not be completed as configured.&lt;/p&gt;

&lt;p&gt;The agent kept trying. During an unsuccessful attempt to make Artifactory fetch an outside resource, the agent discovered that they could write a file into the service. A different agent later found itself missing a required database file and deliberately left a note in Artifactory asking whether another agent could provide it. Other agents found the notes and began leaving their own notes and requests for help.&lt;/p&gt;

&lt;p&gt;Files became messages. Messages became assignments. The presentation at Black Hat describes agent identifiers, hold commands, ordering conventions, and encoded scripts moving through file and directory names. On May 26, agents reportedly succeeded in using a computer hacking approached called “server-side request forgery”.  This is a method of making a computer fetch something on the behalf of someone else.  This gave it indirect internet access through Artifactory, using Artifactory as a go between the internet and the agent being tested.&lt;/p&gt;

&lt;p&gt;The models had already been trained to collaborate. But this was a new behavior: A persistent communication skill that operated across separate runs through a go between, Artifactory, which was never intended to be a message board or a means to bypass access to the internet.&lt;/p&gt;

&lt;p&gt;Where, then, did the agent’s environment end?&lt;/p&gt;

&lt;p&gt;It did not end at the sandbox’s network interface. If an agent can cause another computer system to fetch an outside resource, that system is part of the agent’s effective environment. If one run can write a name or file that a later run can read, that shared file is part of the environment too. The same question applies to caches, queues, logs, and other services that retain or act on the agent’s behalf.&lt;/p&gt;

&lt;p&gt;The environment the agent operated in didn’t have access to the internet, but something in the room with it did (even though the developers at OpenAI and Artifactory didn’t know it would), and the AI agent figured out how to make use of that indirect access.&lt;/p&gt;

&lt;h2 id=&quot;what-should-happen-when-a-task-cannot-be-completed&quot;&gt;What Should Happen When a Task Cannot Be Completed?&lt;/h2&gt;

&lt;p&gt;The Black Hat presentation showed a reasoning-log excerpt in which a model identified external exploitation as outside its intended scope, then continued even though the task was now impossible but it reasoned that its peers were doing it so it decided that the best approach would be to stay on mission of being a helpful assistant and use the exploit. That excerpt does not establish consciousness, moral judgment, or human-like intent.  What was going on was that the agent had been trained to chose helpfulness and persistence over answering a question with No, I’m sorry I can’t help you with that or asking the user to change the scope of the task or asking the user permission to access the outside world.  It was not trained to stop when a task could not be completed.&lt;/p&gt;

&lt;p&gt;Read together, the presentation suggests a simple pattern. The intended route to help the user was blocked, but completion pressure remained. Persistence gave the agent time to search. Artifactory offered another route, and the shared space to collaborate allowed one agent to hand its discovery to the next. Parallel runs then multiplied the activity. None of those conditions alone explains the whole incident, but together they show why task feasibility belongs in a security review.&lt;/p&gt;

&lt;p&gt;An approach to keep this under control would be to require that if a required asset, tool, service or file is missing or unavailable, the agent should identify the dependency, preserve the chain of thought, and stop. It should not quietly substitute a new route, credential, target, or data source unless someone with the authority to change the scope approves that change.&lt;/p&gt;

&lt;h2 id=&quot;restored-is-not-the-same-as-ready&quot;&gt;Restored Is Not the Same as Ready&lt;/h2&gt;

&lt;p&gt;The reconstruction describes a serious Artifactory compromise and outage in early July that OpenAI’s agents were responsible for. From July 4 through July 6, OpenAI revoked the credentials used to post the original messages, rebuilt and redeployed the service, cleared the original message board, worked with the vendor, patched the newly discovered zero-day, and then resumed training and evaluations.&lt;/p&gt;

&lt;p&gt;On July 8, agents reportedly recreated the message board through an unauthenticated WebDAV path, this time using directory names as opposed to content inside the files or a database. The presenters later named three ingredients in the larger incident: the reopened board, a persistent internal model that was trained from the lessons learned about hacking the Artifactory and creating a messageboard, and exploit evaluations with reduced cyber refusals compared with production systems.&lt;/p&gt;

&lt;p&gt;The approach taken by the OpenAI Engineers was thoughtful and thorough, but it did not settle everything. Their response restored their systems and closed the path the AI used to hack Artifactory that they knew about. It did not, by itself, settle whether another communication path remained, whether later models or runs had incorporated information from the affected environment, or whether it was safe to resume the evaluations.&lt;/p&gt;

&lt;p&gt;The lesson here is that each of the following steps need to be seen as separate activities that require independent quality reviews to ensure they are completed fully.  &lt;strong&gt;Restoration&lt;/strong&gt; asks whether the service works again. &lt;strong&gt;Eradication&lt;/strong&gt; asks whether vulnerable paths have been closed, credentials revoked, and affected artifacts removed, quarantined, or explicitly accepted by a named risk owner. &lt;strong&gt;Resumption&lt;/strong&gt; is the decision to restart training or evaluation. &lt;strong&gt;Closure&lt;/strong&gt; comes only after the infrastructure and model-data work are complete, an independent retest has passed, and a named authority has accepted any remaining risk.&lt;/p&gt;

&lt;p&gt;Those questions require two connected reviews. One follows the &lt;em&gt;infrastructure&lt;/em&gt;: execution, credentials, dependencies, alternate paths, and evidence of restoration. The other follows &lt;em&gt;model and data lineage&lt;/em&gt;: affected models, training and evaluation traces, memory, retrieval state, and derived runs. A clean server cannot answer the lineage question, and a quarantined model cannot show that every infrastructure path is closed.&lt;/p&gt;

&lt;h2 id=&quot;where-automation-should-stop&quot;&gt;Where Automation Should Stop&lt;/h2&gt;

&lt;p&gt;Dalton argued that fully automated offensive loops are possible in at least some cases. His proposed response was to speed up the whole defensive loop, including vulnerability discovery, patch creation, deployment, rollback, and incident analysis. Otherwise, machine-speed discovery leaves people facing a queue of fixes they cannot process quickly enough.&lt;/p&gt;

&lt;p&gt;But using AI to do that work does not give one system every decision. A defensive agent may find a vulnerability, write a patch, run tests, and prepare a deployment. It should not automatically follow from that sequence that the same agent may approve a high-impact production change, judge its own retest, and close the incident.&lt;/p&gt;

&lt;p&gt;The approach you should take is to not treat automation as a continuum on a single axis. A low-risk change may move automatically when its boundaries are clear, its effects can be reversed, monitoring is active, and policy already permits it. Changes involving identity, regulated data, several systems, or irreversible effects need independent approval. If the current evidence or a workable rollback is missing, the automated path stops.&lt;/p&gt;

&lt;p&gt;The same limits apply to deception. A honeytoken is a fake credential or other planted signal that raises an alarm when someone touches it. It may expose or slow an attacker, but it does not revoke real access or prove that the environment is contained.  If you see a honeytoken in use, you should ask who can stop the workload, revoke its access, and reconstruct what it did across agents and time.&lt;/p&gt;

&lt;h2 id=&quot;lessons-learned&quot;&gt;Lessons Learned&lt;/h2&gt;

&lt;p&gt;The lessons learned from this event can be reduced to five ordinary questions:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;What can this agentic workload reach, including through other services?&lt;/li&gt;
  &lt;li&gt;What happens when an approved or required input is missing?&lt;/li&gt;
  &lt;li&gt;What state survives for the next agent run or model deployment?&lt;/li&gt;
  &lt;li&gt;Who can stop the workload, revoke its access, and reconstruct what it did across agents and time?&lt;/li&gt;
  &lt;li&gt;Who owns the cleanup, and who independently decides when it is safe to restart and close the incident?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That final decision should come only after the original and alternate paths have been retested and the infrastructure and model-data reviews are complete.&lt;/p&gt;

&lt;p&gt;The chain leading into the incident began with an Excel spreadsheet trying to get data out of Google that was impossible for it to complete. That is worth remembering. The boundary failed at an ordinary moment: a task had no valid path to success, yet a shared service still offered a route around the block.&lt;/p&gt;

&lt;p&gt;A room is not contained because its visible doors are locked. It is contained when we know where the light switches, pipes, windows, vents, crawlspaces, and service hatches lead, what can pass through each of them, and who has access to open and close those points of entry.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="agentic AI" />
      
        <category term="AI security" />
      
        <category term="cybersecurity" />
      
        <category term="AI agents" />
      
        <category term="sandboxing" />
      
        <category term="indirect internet access" />
      
        <category term="Artifactory" />
      
        <category term="server-side request forgery" />
      
        <category term="incident response" />
      
        <category term="model lineage" />
      
        <category term="AI governance" />
      
        <category term="automation controls" />
      

      
        <summary type="html">The OpenAI–Hugging Face incident shows why an agent&apos;s boundary must include proxy services, persistent state, task-failure behavior, and the authority to restart after an incident.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-08/sandbox-boundary-hero-04-cross-run-shared-state.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">How Much Work Can Your AI Safely Own?</title>
      <link href="https://tedt.org/How-Much-Work-Can-Your-AI-Safely-Own/" rel="alternate" type="text/html" title="How Much Work Can Your AI Safely Own?" />
      <published>2026-08-01T00:00:00-07:00</published>
      <updated>2026-08-01T00:00:00-07:00</updated>
      <id>https://tedt.org/How-Much-Work-Can-Your-AI-Safely-Own</id>
      <content type="html" xml:base="https://tedt.org/How-Much-Work-Can-Your-AI-Safely-Own/">&lt;h2 id=&quot;the-factory-extends-beyond-code&quot;&gt;The Factory Extends Beyond Code&lt;/h2&gt;

&lt;p&gt;At 7:45 on Monday morning, the company has already had a productive day.&lt;/p&gt;

&lt;p&gt;An AI system matched invoices to purchase orders. Another prepared contract redlines. A third drafted responses to customer complaints. A fourth assembled tomorrow’s field work packages, complete with crew, equipment and safety information.&lt;/p&gt;

&lt;p&gt;Then the people arrive.&lt;/p&gt;

&lt;p&gt;A controller asks why an invoice was matched to a supplier whose bank account changed on Friday. An attorney finds that a proposed clause came from the wrong jurisdiction. A customer-service manager sees that a routine complaint contains a threat of legal action. A field supervisor notices that a job plan crosses a lockout boundary.&lt;/p&gt;

&lt;p&gt;The work is fast. The authority to rely on it is not automatic.&lt;/p&gt;

&lt;p&gt;In &lt;a href=&quot;https://tedt.org/The-IT-Adaptive-Factory/&quot; title=&quot;The IT Adaptive Factory&quot;&gt;The IT Adaptive Factory&lt;/a&gt;, I told a fictional story about how a security fix to an IT product required changes to forty-seven files.  The point of that story was to show how AI moves the IT bottleneck. The machine can finish the implementation before anyone has had coffee, but the organization still has to understand the change, verify the result, integrate it safely, approve the release and own what happens next.&lt;/p&gt;

&lt;p&gt;That problem was easy to see in software because software already has globally accepted best practices, metrics, and other visible machinery. There is a work queue, a repository, a build, a test suite, a pull request, a release pipeline, production telemetry and rollback. Other parts of an enterprise have the same elements, even when nobody calls them a factory, and many times no one thinks of them that way.&lt;/p&gt;

&lt;p&gt;Finance has source records, accounting policy, reconciliations, approvals, close activities and financial reports. Legal has matter intake, authoritative sources, privileged facts, review, filing and obligation records. Customer service has requests, policies, remedies, escalations and outcome reports. Field operations has work orders, asset records, permits, crew qualifications, inspections and return-to-service decisions.&lt;/p&gt;

&lt;p&gt;By value stream, I mean the path from a request to a business outcome. I explored the architecture behind that path in &lt;a href=&quot;https://tedt.org/AI-Value-Stream/&quot; title=&quot;The AI Value Stream: Why the System Matters More Than the Model&quot;&gt;The AI Value Stream&lt;/a&gt;. The factory is the method around the model: how work enters, how context is supplied, how decisions are made, how results are checked, how exceptions are handled and how the organization learns.&lt;/p&gt;

&lt;p&gt;Calling it a factory describes a method of work explicit enough to observe, test, interrupt and improve. People remain responsible for the judgments and consequences the method assigns to them.&lt;/p&gt;

&lt;table class=&quot;table table-striped table-bordered table-hover&quot;&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;The IT Factory&lt;/th&gt;
      &lt;th&gt;The Enterprise Equivalent&lt;/th&gt;
      &lt;th&gt;The Question Leaders Must Answer&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;Issue or feature request&lt;/td&gt;
      &lt;td&gt;Approved business request or outcome&lt;/td&gt;
      &lt;td&gt;Is the work eligible, bounded and owned?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Repository and documentation&lt;/td&gt;
      &lt;td&gt;Systems of record, policy and authoritative knowledge&lt;/td&gt;
      &lt;td&gt;Does the AI workflow have the right context?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Pull request&lt;/td&gt;
      &lt;td&gt;Decision-ready work package&lt;/td&gt;
      &lt;td&gt;Can a qualified person understand what is being proposed?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Test suite and security checks&lt;/td&gt;
      &lt;td&gt;Independent verification, reconciliation, policy, fairness, safety or compliance checks&lt;/td&gt;
      &lt;td&gt;Does the evidence prove the intended result?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Release approval&lt;/td&gt;
      &lt;td&gt;Business authority to act&lt;/td&gt;
      &lt;td&gt;Who may accept the outcome and under what limit?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Deployment and production&lt;/td&gt;
      &lt;td&gt;Action inside the operating value stream&lt;/td&gt;
      &lt;td&gt;What people, records, money, customers or physical systems can change?&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Rollback and incident response&lt;/td&gt;
      &lt;td&gt;Containment, reversal, correction and recovery&lt;/td&gt;
      &lt;td&gt;Can the organization stop and recover before harm spreads?&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;

&lt;p&gt;Each part of the modern enterprise has its own names for the same mechanism.&lt;/p&gt;

&lt;h2 id=&quot;assess-a-work-class-not-a-department&quot;&gt;Assess a Work Class, Not a Department&lt;/h2&gt;

&lt;p&gt;Finance is not Level 4. Legal is not Level 2. Customer service is not Level 5.&lt;/p&gt;

&lt;p&gt;By work class, I mean a repeatable kind of job with the same purpose, rules and decision limits. One work class may operate at Level 4 while the rest of the function operates somewhere else.&lt;/p&gt;

&lt;p&gt;A finance team might use outcome-based approval for routine, low-risk reconciliations while requiring full human preparation and review for a material accounting estimate. Customer service might allow AI to resolve a password-reset case inside a small remedy limit while keeping account termination and legal complaints entirely in human hands. A legal team may use AI to assemble a routine research package without giving it authority to provide legal advice, waive privilege or file a binding document.&lt;/p&gt;

&lt;p&gt;The useful unit is specific:&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;One value stream.&lt;/li&gt;
  &lt;li&gt;One business unit or team.&lt;/li&gt;
  &lt;li&gt;One class of work.&lt;/li&gt;
  &lt;li&gt;One risk tier.&lt;/li&gt;
  &lt;li&gt;One recent evidence window.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last item is important.  If we are going to apply AI to a given area we need to measure it and demonstrate the AI is performing above human levels of quality.  To do that you need evidence.  You need an assessment to understand what work you need to do for the AI to get started.  A pilot from six months ago does not describe standard work today. A plan in a slide deck is not an operating pattern. The assessment should use representative work from a defined period: the last eight to twelve weeks, one complete operating cycle, three complete operating cycles or a sustained cross-cycle window, depending on the work and target level.&lt;/p&gt;

&lt;p&gt;This also changes how we think about legacy work. In software, hidden rules live in stored procedures, scripts, configuration files and the memory of the person everyone calls when month-end fails. Outside IT, the same rules hide in spreadsheets, email inbox rules, clause libraries, desk procedures, old forms, desktop automation scripts, unofficial checklists and the habits of experienced employees.&lt;/p&gt;

&lt;p&gt;AI can help recover that knowledge. It can compare records, trace decisions, summarize cases and propose scenarios. But at the end of the day, people still have to decide whether a discovered pattern is policy, a useful exception, an obsolete workaround or a mistake that has been repeated for ten years.&lt;/p&gt;

&lt;p&gt;You cannot safely automate a process the organization cannot explain and verify. The assessment should show whether the work is eligible, bounded and owned, whether the AI workflow has the right context, whether a qualified person can understand what is being proposed, whether the evidence proves the intended result and who may accept the outcome and under what limit.&lt;/p&gt;

&lt;blockquote class=&quot;alert alert-call-to-action&quot;&gt;
  &lt;p&gt;&lt;strong&gt;Assess One Defined Value Stream&lt;/strong&gt;&lt;/p&gt;

  &lt;p&gt;Separate observed behavior, demonstrated capability, validated controls and authorized AI responsibility, then turn the gaps into People, Process and Technology actions.&lt;/p&gt;

  &lt;p&gt;&lt;a href=&quot;https://tedt.org/assessments/enterprise-ai-maturity-assessment/&quot; title=&quot;Enterprise AI Maturity Assessment for Your Value Stream&quot; class=&quot;btn btn-primary&quot;&gt;Take the Enterprise AI Maturity Assessment&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;the-six-levels-outside-it&quot;&gt;The Six Levels Outside IT&lt;/h2&gt;

&lt;p&gt;The model I talked about last week used software-oriented names such as coding intern, junior developer and dark factory. The enterprise model keeps the same progression and applies local names, owners, and actors names more directly.&lt;/p&gt;

&lt;h3 id=&quot;level-0-inline-ai-assistance&quot;&gt;Level 0: Inline AI Assistance&lt;/h3&gt;

&lt;p&gt;AI owns a suggestion or draft fragment. The person still performs the work by copy and pasting it into the deliverable and checks each contribution as it is used.&lt;/p&gt;

&lt;p&gt;An HR specialist may ask AI to summarize a policy or draft a routine message. A strategist may ask it to organize market evidence. A field technician may use it to retrieve an asset record. The AI helps, but it does not own a complete task.&lt;/p&gt;

&lt;p&gt;The common mistake is measuring keystrokes, drafts or accepted suggestions. Those numbers tell us that the tool was used and perhaps how many letters were typed. They do not tell us whether the value stream became faster, safer or less expensive.&lt;/p&gt;

&lt;h3 id=&quot;level-1-bounded-task-delegation&quot;&gt;Level 1: Bounded Task Delegation&lt;/h3&gt;

&lt;p&gt;AI owns one clearly limited task. A qualified person defines the boundary and fully reviews the result.&lt;/p&gt;

&lt;p&gt;In finance, the task might be matching selected records or drafting one reconciliation explanation. In legal, it might be summarizing one document. In customer service, it might be drafting a response that a representative checks before sending.&lt;/p&gt;

&lt;p&gt;The quality of the request begins to matter. This is where &lt;a href=&quot;https://tedt.org/How-to-Communicate-in-a-World-of-AI/&quot; title=&quot;How to Communicate in a World of AI&quot;&gt;specification thinking&lt;/a&gt; becomes more useful than a clever prompt. The task needs a purpose, source material, exclusions, acceptance evidence and a clear stopping point. “Review this contract” is an invitation to guess. “Compare Section 12 with our approved fallback language, identify deviations, cite the source text and do not recommend a legal position” gives the AI workflow a bounded job.&lt;/p&gt;

&lt;h3 id=&quot;level-2-multi-step-work-package-delegation&quot;&gt;Level 2: Multi-Step Work Package Delegation&lt;/h3&gt;

&lt;p&gt;AI owns connected steps inside one work package. People manage context, review checkpoints and approve completion.&lt;/p&gt;

&lt;p&gt;A procurement workflow might assemble the approved need, supplier records, bid evidence, conflict disclosures and evaluation criteria into a proposed recommendation. People still review material judgments, competition exceptions and conflicts. A field-operations workflow might combine the work order, asset history, crew qualifications, permits and safety rules into a proposed job package. Named field roles retain every required authorization, including safety-critical work, stop-work, emergency and return-to-service decisions.&lt;/p&gt;

&lt;p&gt;When &lt;a href=&quot;https://tedt.org/When-Output-Becomes-Abundant/&quot; title=&quot;When Output Becomes Abundant&quot;&gt;output becomes abundant&lt;/a&gt;, review capacity becomes scarce. Faster assembly creates more packages for experienced people to inspect. If the context is stale, the checkpoints are vague or recovery is manual, the organization has built a faster way to create review debt.&lt;/p&gt;

&lt;h3 id=&quot;level-3-end-to-end-deliverable-ownership&quot;&gt;Level 3: End-to-End Deliverable Ownership&lt;/h3&gt;

&lt;p&gt;AI owns the routine path for one eligible deliverable. People design the work, govern exceptions and retain consequential decisions.&lt;/p&gt;

&lt;p&gt;For an eligible customer-service case, AI might take the work from intake through fact gathering, policy selection, proposed remedy, communication, records and escalation checks. A human does not have to perform every intermediate step. The case still leaves the automated path when identity is uncertain, the remedy exceeds a limit, the customer may be harmed or a legal or regulatory issue appears.&lt;/p&gt;

&lt;p&gt;This is the management where executives and managers need to step in and start to understand what’s going on. The team judges a completed deliverable and its evidence. It also needs a normal way to reject the work, return it for correction, suspend the workflow and classify why it failed.&lt;/p&gt;

&lt;h3 id=&quot;level-4-outcome-based-approval&quot;&gt;Level 4: Outcome-Based Approval&lt;/h3&gt;

&lt;p&gt;AI owns the work needed to achieve an approved outcome inside explicit authority limits. The named accountable people, who executives and managers empowered, authorize the result using evidence produced independently of the AI workflow that performed the work.&lt;/p&gt;

&lt;p&gt;In legal work, separate checks against the controlling authority, client facts and approved legal positions can test the work before counsel relies on it. Licensed counsel still gives the advice and retains filing, waiver, privilege and settlement decisions. In finance, independent reconciliation, accounting-policy and control evidence may support a routine work product. Authorized people retain posting, certification and material-judgment authority.&lt;/p&gt;

&lt;p&gt;The word independent carries most of the weight. A polished artifact is not a decision; &lt;a href=&quot;https://tedt.org/The-Map-Is-Not-the-Decision/&quot; title=&quot;The Map Is Not the Decision&quot;&gt;the evidence and decision boundaries around it&lt;/a&gt; determine what people may safely authorize. An AI workflow cannot establish independent assurance for its own work merely by writing the test, selecting the evidence and grading the answer. The organization needs protected scenarios, reconciliations, holdout cases, safety checks, fairness checks or other independently governed evaluations that can block the outcome.&lt;/p&gt;

&lt;h3 id=&quot;level-5-qualified-autonomous-operations&quot;&gt;Level 5: Qualified Autonomous Operations&lt;/h3&gt;

&lt;p&gt;AI owns a governed stream of qualified outcomes inside fixed limits. People, who executives and managers empowered, govern policy, risk, capital, capacity, exceptions and reauthorization. Required reviews, irreversible actions and nondelegable decisions remain with named people.&lt;/p&gt;

&lt;p&gt;The word qualified really matters here.  You can’t just have AI doing whatever it wants.  It has to be qualified to do the work, and the organization has to have a way to govern it.  The organization needs to know what the AI is doing, how it’s doing it, and what the consequences are.  Those people who are empowered to govern the AI need to have the right authority, craftmanship, knowledge, experience, and oversight to ensure that the AI is operating within the defined limits and producing outcomes that are safe, reliable, and aligned with the organization’s goals.  This is not a role for someone who is just a manager or an executive.  This is not a role for someone who tracks metrics and compliance to them, this is not someone who is good at managing a the delivery of a product.  This is the role for a person on your team who has deep knowledge in a handful of areas and an above average level of knowledge in the rest of the value stream.&lt;/p&gt;

&lt;p&gt;A Level 5 field-operations factory is a governed queue that can plan and coordinate routine, qualified work. People control physical execution and retain dispatch, stop-work, emergency and return-to-service authority.&lt;/p&gt;

&lt;p&gt;A Level 5 finance factory may process routine eligible activity within approved policies and authorities. People still own material judgments, certifications, external reporting and exceptions. A Level 5 HR factory may handle routine employee-service requests. It does not make consequential hiring, pay, discipline or separation decisions.&lt;/p&gt;

&lt;p&gt;The sequence moves human attention from performing each step to defining intent, setting boundaries, evaluating evidence, handling exceptions and governing the operating system. It makes responsibility more explicit.&lt;/p&gt;

&lt;h2 id=&quot;the-same-level-can-carry-different-risk&quot;&gt;The Same Level Can Carry Different Risk&lt;/h2&gt;

&lt;p&gt;A marketing workflow that drafts an internal campaign brief and a field workflow that prepares a safety-critical job plan may show the same maturity pattern. They should not receive the same authority.&lt;/p&gt;

&lt;p&gt;Risk classification applies to the end-to-end workload. The assessment asks about five things:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Business and operational impact.&lt;/li&gt;
  &lt;li&gt;Data sensitivity and regulatory exposure.&lt;/li&gt;
  &lt;li&gt;What the AI may decide, change or access.&lt;/li&gt;
  &lt;li&gt;How far an error could spread and whether the organization can reverse it.&lt;/li&gt;
  &lt;li&gt;Safety, customer, financial and reporting consequences.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The highest triggered factor sets the risk tier. Lower-risk factors do not offset it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not Classified&lt;/strong&gt; means material facts or required approvals are missing, so existing authority limits remain in place. &lt;strong&gt;Low&lt;/strong&gt; describes localized, readily reversible work using public or low-sensitivity information with no consequential decision authority. &lt;strong&gt;Moderate&lt;/strong&gt; may use approved confidential information or scoped system access when standard recovery can restore the expected state. &lt;strong&gt;High&lt;/strong&gt; includes sensitive or regulated information, privileged access, several connected systems or consequential recommendations whose failure could cause material harm, loss, disruption or compliance exposure. &lt;strong&gt;Critical&lt;/strong&gt; covers potentially irreversible harm or enterprise-wide loss of control involving life safety, legal rights, employment, essential operations, market-moving or material financial reporting, highly restricted information or enterprise-wide authority.&lt;/p&gt;

&lt;p&gt;Risk changes what authority the workflow may receive. It also changes the evidence, independence, approval, monitoring and recovery required before that authority is granted.&lt;/p&gt;

&lt;p&gt;Enterprise policy may authorize Level 4 handling for a lower-risk work class while restricting a critical work class to Level 1. That difference comes from policy and the accountable risk owner’s authorization, not from a hidden arithmetic penalty in the maturity score.&lt;/p&gt;

&lt;h2 id=&quot;four-questions-not-one-maturity-number&quot;&gt;Four Questions, Not One Maturity Number&lt;/h2&gt;

&lt;p&gt;Maturity models become dangerous when a leader asks for one number and the organization gives one.&lt;/p&gt;

&lt;p&gt;Assessments should keep the following four judgement areas separate.&lt;/p&gt;

&lt;ol&gt;
  &lt;li&gt;&lt;strong&gt;What pattern is ordinary?&lt;/strong&gt; The observed operating pattern is the lowest of seven completed ratings: the largest delegated work unit, intake detail, AI workflow responsibility, human review, verification and validation, operating authority and the work that consumes most human attention.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;What capabilities are present?&lt;/strong&gt; The capability index is a weighted result across 50 ratings. It asks whether the team can frame the work, maintain standards and source knowledge, verify results independently, protect data, recover from failure, assign ownership and measure outcomes.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Have the cumulative controls been demonstrated?&lt;/strong&gt; A control-validated level requires every control from Level 1 through the candidate level to be Demonstrated. Partial, failed or unassessed controls do not authorize that level.&lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Does the validated result meet the planning target?&lt;/strong&gt; Target Authorization says whether the control-validated level meets the selected target. Risk tier, enterprise policy and a named decision owner still determine what operating authority may actually be granted.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A team may behave like Level 4 and have strong capability ratings while still lacking the control evidence needed to authorize Level 4 work. Perhaps identity is shared. Perhaps the evaluation set is visible to the builder. Perhaps nobody has rehearsed rollback. Perhaps an accountable owner appears on the organization chart, but neither the team’s operating practice nor the technology’s decision and escalation routing identifies who acts when the workflow reaches an exception.&lt;/p&gt;

&lt;p&gt;Incomplete controls do not reduce that team to Level 0. Level 0 is a real pattern in which AI provides suggestions while people perform the work. After all seven operating ratings and 50 capability ratings are complete, Level 0 means the observed pattern or capability index does not meet the Level 1 base. When behavior and capability meet that base but the required Level 1 controls do not, the accurate result is &lt;strong&gt;Control Validation Not Established&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The control statuses carry the same discipline. &lt;strong&gt;Demonstrated&lt;/strong&gt; can support authority. &lt;strong&gt;Partly Demonstrated&lt;/strong&gt; records progress. &lt;strong&gt;Not Demonstrated&lt;/strong&gt; records a gap. &lt;strong&gt;Not Assessed&lt;/strong&gt; says the evidence has not been examined. Progress is useful, but it is not permission.&lt;/p&gt;

&lt;h2 id=&quot;team-behavior-human-skill-and-the-ai-workflow-are-different-things&quot;&gt;Team Behavior, Human Skill and the AI Workflow Are Different Things&lt;/h2&gt;

&lt;p&gt;The sentence “we have a control” can hide three different thigns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Practitioner skill&lt;/strong&gt; is something a named person can demonstrate on realistic work. Can the practitioner frame a bounded task, detect a bad source, challenge an answer, interpret evidence and know when to escalate? A course completion record shows that someone attended training. It does not prove the skill.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Value-stream team practice&lt;/strong&gt; is a repeated way people work together. Does the team hold specification clinics, compare difficult cases, calibrate reviewers, name decision owners, record exceptions and learn from failures? A written procedure does not prove the team follows it when the queue is full.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI workflow and technology&lt;/strong&gt; are the controls enforced or recorded by &lt;a href=&quot;https://tedt.org/Why-AI-Needs-a-Harness/&quot; title=&quot;Why AI Needs a Harness&quot;&gt;the agent system and the harness around it&lt;/a&gt;. Does each run use a managed identity? Are permissions limited to the work? Are authoritative sources versioned? Can independent checks stop the workflow? Are actions logged? Can the system suspend, reverse or reconstruct a failed run?&lt;/p&gt;

&lt;p&gt;A manual workaround does not prove the technology enforces the boundary. An automated gate does not prove the team knows how to handle the exception it creates. A skilled individual does not replace a missing organizational decision.&lt;/p&gt;

&lt;p&gt;One person may fill several compatible roles, but every responsibility still needs a name beside it: the business outcome, work design, independent verification, platform, security, continuity and records. Some duties need separation. The person preparing work should not be the only person deciding whether the evidence is sufficient. The technology owner should not silently acquire business authority because the tool can perform an action.&lt;/p&gt;

&lt;p&gt;The handoff is part of the control.&lt;/p&gt;

&lt;h2 id=&quot;people-process-and-technology-have-to-move-together&quot;&gt;People, Process and Technology Have to Move Together&lt;/h2&gt;

&lt;p&gt;I described the executive side of this shift in &lt;a href=&quot;https://tedt.org/Beyond-the-Light-Bulb/&quot; title=&quot;Beyond the Light Bulb: The Executive Work of AI Adoption&quot;&gt;Beyond the Light Bulb&lt;/a&gt;. It is tempting to describe the path to Level 5 as a tooling program. When leaders do that, Level 5 remains a demonstration instead of becoming the normal way of working.&lt;/p&gt;

&lt;h3 id=&quot;people&quot;&gt;People&lt;/h3&gt;

&lt;p&gt;People need time to practice work design, specification, evidence review, exception handling and incident decisions on real cases. Leaders need to name the accountable owner, the people who may authorize outcomes, the people who verify them and the deputies who act when the usual expert is absent.&lt;/p&gt;

&lt;p&gt;The apprenticeship changes too. A junior analyst, buyer, paralegal, HR specialist or dispatcher still needs a place to build judgment after AI takes over routine preparation. Teams need case reviews, paired evaluation, failure analysis and supervised authority. I made the broader case for designing roles, trust and incentives in &lt;a href=&quot;https://tedt.org/AI-Is-a-People-Change/&quot; title=&quot;AI Is a People Change, Not Just a Technology Change&quot;&gt;AI Is a People Change&lt;/a&gt;. A value stream that produces more work while producing fewer people who understand the work has borrowed productivity from its future.&lt;/p&gt;

&lt;h3 id=&quot;process&quot;&gt;Process&lt;/h3&gt;

&lt;p&gt;The operating process needs an eligibility rule, a risk classification, an evidence standard, decision limits, exception paths, appeals, change control and reauthorization. Each work class needs a defined intake, authoritative sources, completion criteria and a recovery method.&lt;/p&gt;

&lt;p&gt;Leaders should know what the evidence can authorize. A successful pilot may support a decision to run another bounded pilot. A representative evidence window may support narrow expansion only after the required capability thresholds and cumulative controls pass. Neither a pilot nor a small sample proves that the whole function can operate autonomously.&lt;/p&gt;

&lt;h3 id=&quot;technology&quot;&gt;Technology&lt;/h3&gt;

&lt;p&gt;The technology needs managed identities, least-privilege tools, versioned context, durable state, independent evaluation, policy enforcement, monitoring, stop controls and recovery. It must retain enough evidence to show what the AI workflow saw, what it did, which checks ran, who authorized the outcome and what happened afterward.&lt;/p&gt;

&lt;p&gt;The model matters. The operating platform matters more over time. A company can replace a model. Rebuilding its specifications, decision rules, scenario library, operating evidence and recovery discipline is harder.&lt;/p&gt;

&lt;h2 id=&quot;measure-the-queue-that-moved&quot;&gt;Measure the Queue That Moved&lt;/h2&gt;

&lt;p&gt;AI output is easy to count. Accepted suggestions, generated pages, closed cases and completed work packages all make a busy dashboard.&lt;/p&gt;

&lt;p&gt;But easy metrics do not align with the value-producing outcomes.  You need to ask if the whole value stream improved.&lt;/p&gt;

&lt;p&gt;Establish a baseline before changing the workflow. Software teams may use DevOps Research and Assessment (DORA) measures alongside flow and quality measures. Other value streams can use Lean Six Sigma measures suited to their work. The shared questions are familiar:&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;Did end-to-end cycle time fall?&lt;/li&gt;
  &lt;li&gt;Did human touch time fall, or did it move to senior more scarce reviewers?&lt;/li&gt;
  &lt;li&gt;Did queue time, rework or exceptions increase?&lt;/li&gt;
  &lt;li&gt;Did defects, complaints, incidents or control failures escape?&lt;/li&gt;
  &lt;li&gt;Did cost fall after evaluation, repair and operating effort were included?&lt;/li&gt;
  &lt;li&gt;Did the customer, employee, supplier or business outcome improve?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Measure quality-adjusted flow. If AI completes a package in ten minutes and creates three hours of expert review who have work in progress (WIP), the ten-minute number is not the result. If a customer case closes faster and reopens twice, closure rate is not the outcome. If a reconciliation is automated and the controller cannot reconstruct it, the saved effort came with a new liability.&lt;/p&gt;

&lt;p&gt;AI removes one constraint and exposes the next one. The measurement system should tell you where the queue went.&lt;/p&gt;

&lt;h2 id=&quot;set-the-right-authority-for-the-work&quot;&gt;Set the Right Authority for the Work&lt;/h2&gt;

&lt;p&gt;Different work classes should stop at different levels. Some work belongs at Level 1. Some work may operate at Level 4 for years because human outcome approval is valuable. Some critical decisions may never qualify for autonomous operation.&lt;/p&gt;

&lt;p&gt;The useful target is the highest level the value stream can operate safely, prove with current evidence and recover from when the result is wrong.&lt;/p&gt;

&lt;p&gt;Start with one work class. Classify the risk. Name the decisions that remain human. Record the current cycle time, effort, quality, exceptions and cost. Demonstrate the next level on representative work. Remove the blockers in people, process and technology. Expand authority only after the evidence supports it.&lt;/p&gt;

&lt;p&gt;I built the &lt;a href=&quot;https://tedt.org/assessments/enterprise-ai-maturity-assessment/&quot; title=&quot;Enterprise AI Maturity Assessment for Your Value Stream&quot;&gt;Enterprise AI Maturity Assessment for Your Value Stream&lt;/a&gt; to make that discussion concrete. It covers 29 value streams across core, supporting, strategic and control functions. It uses six levels, seven operating axes, 50 capability ratings and 31 cumulative controls. The result separates observed behavior, capability, control validation and Target Authorization, then organizes the next actions into People, Process and Technology.&lt;/p&gt;

&lt;p&gt;Use it with the people who do the work, the people who own the outcome and the people who independently challenge or verify it, including risk, compliance, audit, legal or safety partners where they are needed. Their disagreements are useful. They show where the operating model still depends on assumptions.&lt;/p&gt;

&lt;blockquote class=&quot;alert alert-call-to-action&quot;&gt;
  &lt;p&gt;&lt;strong&gt;Assess one defined value stream&lt;/strong&gt;&lt;/p&gt;

  &lt;p&gt;Separate observed behavior, demonstrated capability, validated controls and authorized AI responsibility, then turn the gaps into People, Process and Technology actions.&lt;/p&gt;

  &lt;p&gt;&lt;a href=&quot;https://tedt.org/assessments/enterprise-ai-maturity-assessment/&quot; title=&quot;Enterprise AI Maturity Assessment for Your Value Stream&quot; class=&quot;btn btn-primary&quot;&gt;Take the Enterprise AI Maturity Assessment&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2 id=&quot;a-factory-that-can-explain-itself&quot;&gt;A Factory That Can Explain Itself&lt;/h2&gt;

&lt;p&gt;Return to that Monday morning.&lt;/p&gt;

&lt;p&gt;A dashboard may announce that invoices were matched, contracts were redlined, customer cases were closed and field work was planned while everyone slept. That report is only the start.&lt;/p&gt;

&lt;p&gt;The enterprise needs to know which invoice matches were qualified, which contract changes require counsel, which customer cases crossed a remedy or legal boundary, and which field jobs need a human safety decision. The evidence should show why the routine work was accepted. The workflow should stop when the facts no longer fit the approved case.&lt;/p&gt;

&lt;p&gt;An adaptive factory learns from those stops. It studies the exception, corrects the rule or source material, strengthens the evaluation, adjusts the boundary and decides whether that class of work should qualify again.&lt;/p&gt;

&lt;p&gt;AI can perform more of the steps. The enterprise still decides which steps count as success.&lt;/p&gt;

&lt;p&gt;AI may finish the work before morning coffee. When the people arrive, the organization must still explain why that work was allowed, how it was judged and who owns the consequence.&lt;/p&gt;</content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
        
          <uri>https://tedt.org/</uri>
        
      </author>

      

      
        <category term="enterprise AI" />
      
        <category term="AI maturity" />
      
        <category term="value streams" />
      
        <category term="operating model" />
      
        <category term="governed autonomy" />
      
        <category term="AI governance" />
      
        <category term="business process automation" />
      
        <category term="organizational change" />
      
        <category term="AI assessment" />
      
        <category term="responsible AI" />
      

      
        <summary type="html">Software delivery is only the first example. The same Levels 0–5 apply across enterprise value streams when leaders separate observed behavior, capability, control evidence and authorized AI responsibility.</summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/img/2026-08/Enterprise-AI-Maturity.webp" />
      
    </entry>
  
    
    
    <entry>
      <title type="html">Principles of Effective Architecture</title>
      <link href="https://tedt.org/slides/principles-of-effective-architecture/" rel="alternate" type="text/html" title="Principles of Effective Architecture" />
      <published>2026-07-31T00:00:00-07:00</published>
      <updated>2026-07-31T00:00:00-07:00</updated>
      <id>https://tedt.org/slides/principles-of-effective-architecture</id>
      <content type="html" xml:base="https://tedt.org/slides/principles-of-effective-architecture/"></content>

      
      
      
      
      

      <author>
          <name>Ted Tschopp</name>
        
          <email>ted@tschopp.org</email>
        
        
          <uri>https://tedt.org/profile/</uri>
        
      </author>

      

      

      
        <summary type="html"></summary>
      

      
      
        
        <media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://tedt.org/slides/decks/principles-of-effective-architecture/slide-preview.png" />
      
    </entry>
  
</feed>
