{
    "version": "https://jsonfeed.org/version/1",
    "title": "Natalia Cummings",
    "description": "",
    "home_page_url": "https://nataliacummings.com",
    "feed_url": "https://nataliacummings.com/feed.json",
    "user_comment": "",
    "icon": "https://nataliacummings.com/media/website/Logo_NC_tagline_caps_for-light-bg.png",
    "author": {
        "name": "Natalia Cummings"
    },
    "items": [
        {
            "id": "https://nataliacummings.com/the-model-you-chose-is-already-behind.html",
            "url": "https://nataliacummings.com/the-model-you-chose-is-already-behind.html",
            "title": "The Model You Chose Is Already Behind",
            "summary": "It is Friday evening in July, and your best engineer is still at her desk. She's reworking one of your most valuable production applications, since she just found out the foundation model that it runs on will be shut off in October. Well before then&hellip;",
            "content_html": "<p><span style=\"font-weight: 400;\">It is Friday evening in July, and your best engineer is still at her desk. She's reworking one of your most valuable production applications, since she just found out the foundation model that it runs on will be shut off in October. Well before then she has to find, test, and evaluate other models, and migrate the load. She's figuring this out on her own because this is the first time she's encountered this, and central IT never put up a system to manage this. That is her weekend, and probably her next few.</span></p>\n<p><span style=\"font-weight: 400;\">She only found out yesterday, by chance. On a routine call, one of your hyperscalers' account managers asked whether you were still using the old model, and mentioned in passing that it was being retired. The formal notice had gone out months earlier to a technical inbox nobody monitors. Nobody whose application depended on that model had seen it.</span></p>\n<p><span style=\"font-weight: 400;\">Here is the part that should worry you more than the lost weekend: Models age on a curve. The oldest get retired outright. The next tier still runs but the provider stops giving you room to grow on it. The rest simply fall behind newer models that cost less and do more. She caught this one because it was being retired and someone happened to mention it. You are running many models across many applications. She got lucky on one. Is anybody watching the rest?</span></p>\n<h2><span style=\"font-weight: 400;\">Falling behind when nothing's broken</span></h2>\n<p><span style=\"font-weight: 400;\">The knock-on effect of this issue at scale goes beyond overworked engineers. For eighteen months the application ran on a model nobody revisited while newer releases undercut it on price and performance, and nothing on any dashboard flagged it, because nothing was broken. The scramble at least announces itself. This just leaves money on the table, and no one is going to flag it for you.</span></p>\n<h2><span style=\"font-weight: 400;\">The playbook that used to work</span></h2>\n<p><span style=\"font-weight: 400;\">Most executives I work with formed their instincts about platform decisions around data. Data behaves differently from foundation models, and that difference is why continuous re-evaluation sounds like churn rather than discipline.</span></p>\n<p><span style=\"font-weight: 400;\">Think about how you chose a data foundation. You selected a warehouse or a database, you built on it, and you ran it for the better part of a decade. Upgrades were events you scheduled around the business, trading a planned downtime for years of stability and the assumption of better data products. You may have standardized on one. More likely you run several. Either way, each was chosen (or inherited), on the expectation it would still be there in five years, and that expectation was usually right.</span></p>\n<p><span style=\"font-weight: 400;\">AI applications don't work that way, because foundation model release cycles are far shorter than data platform cycles. Generations turn over in months. This hits your business on both ends: you are forced off an old model sooner than you planned for, and better, cheaper options arrive faster than your planning cycle can absorb them. A decision you used to make once has become a decision you have to keep making, and most organizations have no mechanism for a decision that repeats.</span></p>\n<figure class=\"post__image post__image--center\"><img loading=\"lazy\"  src=\"https://nataliacummings.com/media/posts/4/nc-releases-curve-v2-2.png\" alt=\"\" width=\"1080\" height=\"660\" sizes=\"(max-width: 1920px) 100vw, 1920px\" srcset=\"https://nataliacummings.com/media/posts/4/responsive/nc-releases-curve-v2-2-xs.png 640w ,https://nataliacummings.com/media/posts/4/responsive/nc-releases-curve-v2-2-sm.png 768w ,https://nataliacummings.com/media/posts/4/responsive/nc-releases-curve-v2-2-md.png 1024w ,https://nataliacummings.com/media/posts/4/responsive/nc-releases-curve-v2-2-lg.png 1366w ,https://nataliacummings.com/media/posts/4/responsive/nc-releases-curve-v2-2-xl.png 1600w ,https://nataliacummings.com/media/posts/4/responsive/nc-releases-curve-v2-2-2xl.png 1920w\"></figure>\n<h2><span style=\"font-weight: 400;\">By the time the email arrives, you are already late</span></h2>\n<p><span style=\"font-weight: 400;\">The providers are not hiding the timetable. OpenAI commits to at least six months of notice for generally available models and as little as two weeks for anything with \"preview\" in the name to stop responding. Microsoft gives at least sixty days before an Azure model is retired, and Anthropic sixty days for publicly released models. The notice that reached the wrong inbox in the opening is a real failure worth fixing, but fixing it only means you hear the bad news sooner.</span></p>\n<p><span style=\"font-weight: 400;\">The deeper problem is that in most companies an end-of-life notice is the only thing that ever triggers a fresh look at the model an application runs on. There is no standing process to track what the providers have released, evaluate it against the workloads you actually run, promote it where it wins, and decommission what it replaces. A company that has that process treats the October email as administrative: it evaluated the successor in May because the successor was cheaper, and the notice confirms a decision already made. A company without one has no strategy other than the email, and the email jumps the queue ahead of whatever the team had planned.</span></p>\n<p><span style=\"font-weight: 400;\">Capacity behaves the same way, and it is not confined to the cloud providers. Serving capacity is allocated per model, and it follows demand toward the newest ones. Through 2026 the frontier labs have rationed throughput rather than raise prices, tightening limits at peak hours as demand outran the hardware. Standardize on an older model and you may find there is no more of it to be had in the week you need it most.</span></p>\n<h2><span style=\"font-weight: 400;\">Not choosing was a choice</span></h2>\n<p><span style=\"font-weight: 400;\">Ask who owns keeping your AI current and the answer in most companies is that nobody does, not obviously.</span></p>\n<p><span style=\"font-weight: 400;\">Sometimes central IT treats it as an annual procurement question, if it’s asked at all. More often leadership never set a protocol, so each line of business decides for itself when to move and what to move to, with no standard and no trigger. Some teams stay sharp. Most let their applications age, because there is always something more urgent than upgrading something that already works.</span></p>\n<p><span style=\"font-weight: 400;\">What I see across engagements is that the architecture is rarely what fails. Teams build reasonable systems. What is missing is that nobody has been made accountable for noticing when the ground has shifted, and nobody has been told what should oblige them to act.</span></p>\n<p><span style=\"font-weight: 400;\">You also cannot fix that with a committee. Past a handful of use cases, a central group choosing models can't keep up. The work has to be divided on purpose. The center owns the what and the why: which models are approved, what guardrails apply, and what events trigger a fresh look at a model choice. It also owns the platform everything runs through, the gateway whose metadata ties each model call back to a line of business, which is what makes usage and ownership visible without anyone maintaining a list. The builders own the how: evaluate what is new, adopt it where it makes sense, migrate, decommission the old. Centralize the decision and you get a bottleneck. Leave it to chance and every team does something different, and most do nothing.</span></p>\n<h2><span style=\"font-weight: 400;\">What Intuit built, and what to borrow</span></h2>\n<p><span style=\"font-weight: 400;\">The clearest example of doing this well is one I watched up close. For several years I worked alongside Intuit as a partner, meeting monthly with their VPs of Data and of AI to work through new model testing and plan migrations together.</span></p>\n<p><span style=\"font-weight: 400;\">What struck me was not just the quality of their engineers, but that they had built a system to enable the whole organization to keep pace without anyone losing a weekend to it. The system is called GenOS, and everything I am describing is public: it's a model-agnostic platform giving developers an extensible catalog of frontier, open-source, and proprietary models- including Intuit's own financial ones, with a runtime layer that selects for the job and an internal leaderboard ranking models by task. Their CTO, Alex Balazs, notes that this model-agnostic approach is exactly how they future-proof the platform.</span></p>\n<p><span style=\"font-weight: 400;\">The leaderboard is the part worth replicating, because what it really did was democratize knowledge. Teams could see which models were winning which tasks, why they won, and the prompts that got them there. Nobody was reinventing the wheel. When Intuit shipped an agent starter kit on that foundation, nine hundred internal developers built hundreds of agents in just five weeks.</span></p>\n<p><span style=\"font-weight: 400;\">You don't need to build GenOS, but you do need what it does, which is make switching models a routine engineering task instead of an improvised one. Pointing an application at a different model takes minutes. What follows is the real work, because a prompt tuned for one model routinely underperforms on another, and while a newer model can now do much of the re-tuning, the research still shows it cannot be trusted to do it unsupervised. A good system buys bounded work rather than no work. What bounds it is a set of real cases with known good answers, versioned alongside the prompts and run against every candidate, so a migration becomes a defined debugging exercise rather than an open-ended experiment.</span></p>\n<p><span style=\"font-weight: 400;\">That expectation also has to reach your partners. If system integrators build your AI, ask them before you sign how they intend to keep the application portable across model generations, and what they will hand over so your own team can run the next migration without them. The answer tells you a great deal about what you're buying.</span></p>\n<h2><span style=\"font-weight: 400;\">What two years did to the price of the same result</span></h2>\n<p><span style=\"font-weight: 400;\">Stanford's Human-Centered AI Institute tracks the price of capability rather than the price of any particular model. Between November 2022 and October 2024, the cost of running a system at GPT-3.5 level of performance fell from twenty dollars per million tokens to seven cents, a decline of more than 280 times in under two years, measured at fixed benchmark parity rather than by substituting a weaker model.</span></p>\n<p><span style=\"font-weight: 400;\">Put a real workload against that. An application consuming five hundred million tokens a month cost ten thousand dollars a month at the opening price and thirty-five dollars a month at the closing one, for the same task, at the same quality, at the same volume of queries. Nobody had to optimize anything to earn that reduction. The market produced it, but it only reaches the companies that can act on it.</span></p>\n<p><span style=\"font-weight: 400;\">Capability is converging at the same time. Stanford's 2026 index found the leading models from six different labs, Anthropic, OpenAI, Google, xAI, Alibaba and DeepSeek, scoring within a hair of one another in public head-to-head rankings. When every leading model is good enough, the advantage moves to whoever runs the cheapest, most reliable one. Your competitor may not be winning on model quality at all. They may simply be delivering a comparable result more cheaply and reliably, and reinvesting the difference everywhere else.</span></p>\n<h2><span style=\"font-weight: 400;\">Ninety-one days is not the problem</span></h2>\n<p><span style=\"font-weight: 400;\">Go back to Friday evening, and your engineer at seven o'clock, judging claims one at a time. She is doing by hand what a mature team would have automated: pulling real examples, judging the new model's answers against the old, adjusting prompts, checking again. Whatever she works out stays with her, in her head and a spreadsheet, not in anything the company owns. When the next team hits the same wall in February, they start over.</span></p>\n<p><span style=\"font-weight: 400;\">The weekend and the ninety-one days are the least of it. What your organization actually lost is the exercise itself. It went through all of that and came out the other side with no more capability than it had going in, and it will pay again the next time the frontier moves.</span></p>\n<p><span style=\"font-weight: 400;\">The decision in front of you concerns the system rather than the technology. Whether the next one of these produces a repeatable process or another Friday night for your best engineer is a governance question, it belongs to you rather than to your architects, and you can set the mandate in an afternoon.</span></p>\n<h2><span style=\"font-weight: 400;\">Three things to do Monday morning</span></h2>\n<p><strong>Fund the test set.</strong><span style=\"font-weight: 400;\"> Require every AI application in production to have a set of real cases with known good answers, versioned with its prompts and run against any candidate model. This is the asset that makes every future migration routine, and running it costs a typical team a few hundred dollars a month in evaluation API calls.</span></p>\n<p><strong>Assume nobody owns this, and assign it.</strong><span style=\"font-weight: 400;\"> You won't find an existing owner, because in most companies there isn't one. So name an owner for each AI application, and give them the short list of events that should trigger a fresh look at the model: a new release, a price drop, a capacity limit, an end-of-life notice.</span></p>\n<p><strong>Ask one question in your next business review.</strong><span style=\"font-weight: 400;\"> What models are we running in production, who owns each one, when was each last evaluated, and what deprecation notices apply. If nobody can answer inside a week, you have found the gap, and the answer costs you nothing to demand.</span></p>\n<p><i><span style=\"font-weight: 400;\">All thoughts, ideas, and opinions expressed here are my own.</span></i></p>\n<h3><span style=\"font-weight: 400;\">Sources</span></h3>\n<ol>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">OpenAI, </span><i><span style=\"font-weight: 400;\">Deprecations</span></i><span style=\"font-weight: 400;\">, developer documentation, accessed July 2026 (minimum notice periods: at least six months for generally available models, three months for specialized variants, roughly two weeks for preview models; notification by email to accounts actively using the model).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">OpenAI, batch deprecation notice, 22 April 2026 (more than twenty-five model IDs scheduled for shutdown across 23 July and 23 October 2026, including gpt-3.5-turbo, gpt-4, o1, o3-mini and o4-mini).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Microsoft, </span><i><span style=\"font-weight: 400;\">Azure OpenAI in Microsoft Foundry model deprecations and retirements</span></i><span style=\"font-weight: 400;\">, accessed July 2026 (not-sooner-than retirement date set at launch, 365 days for generally available models and 90–120 days for preview; at least 60 days' notice before GA retirement).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Anthropic, </span><i><span style=\"font-weight: 400;\">Model deprecations</span></i><span style=\"font-weight: 400;\">, accessed July 2026 (at least 60 days' notice for publicly released models).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Microsoft, </span><i><span style=\"font-weight: 400;\">Azure OpenAI in Microsoft Foundry Models quotas and limits</span></i><span style=\"font-weight: 400;\">, accessed July 2026 (quota increase request process; priority given to customers actively using existing allocation; TPM and RPM limits defined per region, per subscription, per model).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Scientific American, </span><i><span style=\"font-weight: 400;\">What is the AI compute crunch, and why are AI tools hitting usage limits?</span></i><span style=\"font-weight: 400;\">, May 2026 (providers currently prefer rate limiting to price increases under capacity constraint). Contemporary reporting also documents Anthropic tightening peak-hour usage limits in March 2026 and GitHub pausing new Copilot signups in April 2026 on compute grounds.</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Stanford HAI, </span><i><span style=\"font-weight: 400;\">The 2025 AI Index Report</span></i><span style=\"font-weight: 400;\">, April 2025 (inference cost for a system performing at GPT-3.5 level fell more than 280-fold between November 2022 and October 2024, from $20.00 to $0.07 per million tokens). </span><i><span style=\"font-weight: 400;\">Note: this is a benchmark-parity measure, standardized at an MMLU score equivalent to GPT-3.5, not a like-for-like enterprise invoice. The $10,000 and $35 monthly figures in the text apply those published unit prices to an illustrative 500-million-token workload and are arithmetic, not a measured case.</span></i></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Stanford HAI, </span><i><span style=\"font-weight: 400;\">The 2026 AI Index Report</span></i><span style=\"font-weight: 400;\">, 2026 (as of March 2026, leading models from Anthropic, xAI, Google, OpenAI, Alibaba and DeepSeek clustered within 25 Elo points on the Arena leaderboard, shifting competition toward cost and reliability).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Constellation Research, </span><i><span style=\"font-weight: 400;\">Intuit embraces LLM choice for multiple use cases</span></i><span style=\"font-weight: 400;\">, September 2024 (GenOS AI Workbench includes an LLM leaderboard, prompt management and automated evaluation; CTO Alex Balazs on model-agnostic architecture as future-proofing, Intuit Investor Day).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Intuit Engineering, </span><i><span style=\"font-weight: 400;\">Intuit's Custom LLM Leaderboard: Optimizing Model Selection for Financial Use Cases</span></i><span style=\"font-weight: 400;\">, October 2024 (leaderboard architecture, evaluation framework, model registry, benchmark management).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Intuit, </span><i><span style=\"font-weight: 400;\">Technology at Intuit</span></i><span style=\"font-weight: 400;\">, accessed July 2026 (GenStudio provides an extensible catalog of commercial, open source, and proprietary LLMs, including Anthropic Claude via AWS Bedrock, Gemini, LLaMa, and Mistral, alongside Intuit's own financial models).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">VentureBeat, </span><i><span style=\"font-weight: 400;\">Inside Intuit's GenOS update</span></i><span style=\"font-weight: 400;\">, December 2025 (Agent Starter Kit enabled 900 internal developers to build hundreds of AI agents within five weeks).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Jahani et al., </span><i><span style=\"font-weight: 400;\">Prompt adaptation as a dynamic complement in generative AI systems</span></i><span style=\"font-weight: 400;\">, 2026, as discussed in arXiv:2604.27082, </span><i><span style=\"font-weight: 400;\">When Your LLM Reaches End-of-Life: A Framework for Confident Model Migration in Production Systems</span></i><span style=\"font-weight: 400;\">, April 2026 (prompt adaptation is a primary driver of results when switching models; automatic prompt rewriting is not yet reliably effective).</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Practitioner literature on production LLM migration, 2026 (semantic rather than syntactic failure modes on model swap; versioned prompt repositories tied to validated model versions; canary routing of a small share of live traffic scored against the offline rubric). </span><i><span style=\"font-weight: 400;\">Note: this is established practitioner consensus rather than a single primary study; a primary source would be preferable but I haven't yet found one.</span></i></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Evaluation API cost for running a modest test set (on the order of a few hundred dollars a month for a team with two to three prompt paths), from 2026 LLM testing tooling surveys. </span><i><span style=\"font-weight: 400;\">Note: figure is illustrative and order-of-magnitude, drawn from vendor-adjacent tooling estimates rather than an audited benchmark; actual cost varies with test-set size, model pricing, and evaluation frequency.</span></i></li>\n</ol>",
            "image": "https://nataliacummings.com/media/posts/4/Screenshot-2026-07-25-at-11.16.07-PM.png",
            "author": {
                "name": "Natalia Cummings"
            },
            "tags": [
            ],
            "date_published": "2026-07-25T22:22:41+02:00",
            "date_modified": "2026-07-25T23:17:44+02:00"
        },
        {
            "id": "https://nataliacummings.com/youre-funding-ai-wrong.html",
            "url": "https://nataliacummings.com/youre-funding-ai-wrong.html",
            "title": "You&#x27;re Funding AI Wrong",
            "summary": "It's a Tuesday afternoon in the quarterly business review and your CTO is eight slides deep into the AI agent program update. The room is expensive: twelve senior leaders, two hours blocked, catering nobody's touching. Slide after slide shows pilot results, accuracy metrics, demo videos,&hellip;",
            "content_html": "<p>It's a Tuesday afternoon in the quarterly business review and your CTO is eight slides deep into the AI agent program update. The room is expensive: twelve senior leaders, two hours blocked, catering nobody's touching. Slide after slide shows pilot results, accuracy metrics, demo videos, architecture diagrams. Everything looks impressive.</p>\n<p>Then your CFO leans forward and asks the question that's been building for two quarters: \"We've put $4M into this program. What's it giving us back?\" The CTO pulls up a slide about \"expected efficiency gains\" and \"projected time savings,\" but the numbers are modeled, not measured. Nobody in the room can point to a dollar saved, a process shortened, or a customer outcome changed. Five pilots have been funded, with promising results in controlled environments, but there's still nothing the CFO can take to the board and defend.</p>\n<p>What happens next is the part I've watched play out in company after company. The CTO talks about model maturity and architecture decisions that are \"de-risking production readiness.\" The CISO reframes the delay as responsible stewardship: three open risk assessments before any agent touches production data. The CDO produces a data-readiness roadmap that conveniently completes next fiscal year. Everyone is protecting their program, their budget, their headcount. Nobody is lying. But nobody is solving the actual problem either, because the problem isn't in any single leader's domain. It sits in the gaps between them.</p>\n<p>Here's what I tell every leadership team living this moment. Most enterprises that stall with AI are funding it as a technology program when what they need is an operating model transformation. The models were never the hard part.</p>\n<p>If you've been in enterprise technology long enough, the symptoms are recognizable. Funded programs, talented teams, promising technology, no production outcomes. We watched the same movie play out with data platforms about a decade ago. Companies spent millions standing up data lakes, hired the teams to run them, and the first use cases proved the value was real. Then it stalled, because nobody built the layer between a lake full of raw data and a business that could safely use it, and every new team that wanted in had to solve access, quality, and compliance from scratch. Gartner had a name for where most of them ended up: the data swamp. The one difference is speed- what took five years to become a crisis with data is taking about eighteen months with agents.</p>\n<p>But failing to ship is only one of three ways this goes wrong. In my experience it isn't even the most expensive one.</p>\n<h3>Three ways to fail at AI agents</h3>\n<p><strong><span class=\"fm\">Failure mode 1: Never ship.</span></strong> Roughly 83% of US enterprises have funded agentic AI projects; about 41% have gotten one to production (Codiste, 2026). The pilots work. The demos impress leadership. Then the project hits security review, data-access negotiations, and compliance questions it can't answer, and it sits there until the next reorg quietly kills it.</p>\n<p><strong><span class=\"fm\">Failure mode 2: Ship but can't prove value.</span> </strong>This one is more insidious because from the outside it looks like success. The agent is in production. It's doing something. But nobody defined what success looked like before launch, so there's no baseline, no way to know if it's improving or degrading, no feedback loop tying it to the outcome that justified the spend. Six months later someone asks \"what's the ROI?\" and the honest answer is \"we don't know.\" The agent keeps running because nobody wants to be the person who turns it off.</p>\n<p><strong><span class=\"fm\">Failure mode 3: Ship too fast and lose control.</span></strong> This is where the organizations that moved aggressively without governance end up. They pressed GO across the enterprise, hoping that by enabling their employees with AI that would somehow turn into efficiencies. Now costs are climbing. Licenses are piling up across multiple vendors. Agents are running that nobody owns. Shadow AI is proliferating because lines of business deployed their own tools outside IT. When leadership asks \"what do we decommission?\" the answer is \"we don't know what's running, let alone which ones to keep.\" The speed they gained by skipping governance is now costing more in cleanup than governance would have cost to build.</p>\n<p>The pattern I see repeatedly: these aren't alternatives, they're sequential. Organizations hit #1, relax controls to break the bottlenecks, then hit #2 and #3 within a year.</p>\n<h3>The number that should end the debate</h3>\n<p>Databricks analyzed 20,000 organizations for its 2026 State of AI Agents report and found one result that speaks to all three failure modes at once: companies that implemented AI governance pushed 12x more projects to production than those that didn't.</p>\n<p>That stat usually gets filed under \"shipping.\" But that undersells it. Governance forces you to define success criteria before deployment, which is what fixes failure mode 2. And it creates the registry, identity, and evaluation infrastructure that prevents sprawl, which is what fixes failure mode 3. The same report found organizations using structured evaluation tools moved 6x more systems to production, not because evaluation slows things down, but because it gives the CISO evidence to approve a deployment in days instead of blocking it for months. Those same evaluation frameworks then double as the measurement layer that tells you whether an agent is delivering value after launch.</p>\n<p>One investment addresses all three failure modes.</p>\n<h3>The budget inversion problem</h3>\n<p>BCG studied what actually drives AI outcomes and found a ratio that, frankly, explains most of the stalled programs I get called into.</p>\n<figure class=\"post__image post__image--center\"><img loading=\"lazy\"  src=\"https://nataliacummings.com/media/posts/2/Screenshot-2026-06-25-at-4.29.07-PM.png\" alt=\"\" width=\"1174\" height=\"286\" sizes=\"(max-width: 1920px) 100vw, 1920px\" srcset=\"https://nataliacummings.com/media/posts/2/responsive/Screenshot-2026-06-25-at-4.29.07-PM-xs.png 640w ,https://nataliacummings.com/media/posts/2/responsive/Screenshot-2026-06-25-at-4.29.07-PM-sm.png 768w ,https://nataliacummings.com/media/posts/2/responsive/Screenshot-2026-06-25-at-4.29.07-PM-md.png 1024w ,https://nataliacummings.com/media/posts/2/responsive/Screenshot-2026-06-25-at-4.29.07-PM-lg.png 1366w ,https://nataliacummings.com/media/posts/2/responsive/Screenshot-2026-06-25-at-4.29.07-PM-xl.png 1600w ,https://nataliacummings.com/media/posts/2/responsive/Screenshot-2026-06-25-at-4.29.07-PM-2xl.png 1920w\"></figure>\n<p>Look at that again, because the implication is brutal: the model and the stack it runs on account for less than a third of what determines success. The other 70% is organizational and operational: the connective tissue between \"the technology works\" and \"the business gets value from it.\"</p>\n<p>Now think about how most enterprises actually allocate AI budget. The majority goes to models, compute, and engineering talent. A fraction goes to data governance, success criteria, evaluation, and process change. The budget is allocated almost perfectly inverse to what the research says drives outcomes. I have yet to walk into a stalled program where this wasn't true.</p>\n<p>And the cost of getting the order wrong compounds. Without governance infrastructure, organizations get locked into whatever they deployed first, paying last quarter's prices for last quarter's capabilities, because switching costs pile up faster than anyone budgeted. The companies I advise that treat governance as an acceleration investment keep the optionality to move when something better arrives. The ones that don't are stuck.</p>\n<h3>What the 12x companies actually built</h3>\n<p>The companies pushing 12x more AI to production built three things:</p>\n<p><strong><span class=\"fm\">Agent identity.</span></strong> Every agent has its own identity, so you always know which agent took an action and who owns it. And when an agent acts for a person, it inherits that person's access limits, so it can never reach data the person couldn't reach themselves. When that employee leaves, their agents don't keep running on orphaned access. When leadership asks \"what's running and who owns it?\", there's an answer.</p>\n<p><strong><span class=\"fm\">A control plane.</span></strong> One layer that knows what agents exist, what tools and data they can reach, what they did, and what they cost. Gartner named this category the \"agent management platform\" in March 2026. It removes the security-review bottleneck, provides the visibility to measure value, and gives you the inventory to manage sprawl.</p>\n<p><strong><span class=\"fm\">Evaluation gates with success criteria.</span></strong> Not \"test it before launch\" but \"define what success looks like, measure it continuously, and flag when it degrades.\" This is the line between agents that provably deliver and agents that run forever because nobody knows whether to keep them.</p>\n<p>None of these slow you down. They remove the friction that creates the failure modes in the first place.</p>\n<h3>The question for your next board meeting</h3>\n<p class=\"font-claude-response-body break-words whitespace-normal\">If you're the leader in that Tuesday QBR, watching demo after demo impress the room and move nothing on the P&amp;L, be precise about what you're seeing. The technology works. What's missing is the operating model around it, the connection that turns a working demo into a result the business can count. That should be reassuring, because an operating model is something you own and can change.</p>\n<p class=\"font-claude-response-body break-words whitespace-normal\">It's also what now separates the companies winning with AI from the ones quietly losing money on it. Ernst &amp; Young surveyed 975 executives at billion-dollar companies last year and found that firms with real oversight, meaning live monitoring and an actual governance committee, were far more likely to report revenue growth and cost savings than firms without it. Governance is what lets the value through.</p>\n<p class=\"font-claude-response-body break-words whitespace-normal\">So the question becomes one of timing. You can build this layer now, while your agent count is small, the architecture is still reversible, and the registry that tells you what's running costs almost nothing. Or you can build it later, after finance flags a charge nobody can explain and security is chasing an agent nobody owned, when the same work runs five to ten times the cost on a board deadline. The capability you end up with is identical. Only the price and the pressure change.</p>\n<p class=\"font-claude-response-body break-words whitespace-normal\">That's the whole decision. Every enterprise lands in the same place eventually. What you're choosing today is whether you get there by design or under duress.</p>\n<blockquote>\n<p>\"The companies governing their agents today are the same ones that governed their data a decade ago. The discipline that won them the last platform shift is winning them this one.\"</p>\n</blockquote>\n<h3>Three things to do Monday morning</h3>\n<p><strong><span class=\"fm\">1. Build an agent registry this quarter.</span> </strong>You can't govern what you can't see. One registry, one identity per agent, one named owner. This is a weeks-level effort, not a years-level program.</p>\n<p><strong><span class=\"fm\">2. Define success criteria before deployment, not after.</span> </strong>What does this agent need to deliver? How will you measure it? What does degradation look like? Make it a launch prerequisite.</p>\n<p><strong><span class=\"fm\">3. Fund governance from your AI budget, not your compliance budget.</span></strong> The budget owner sets the pace. If governance lives in compliance, it moves at compliance speed and fixes only one of the three failure modes. Put it where the urgency is.</p>\n<p><em>All thoughts, ideas, and opinions expressed here are my own.</em></p>\n<h3><span style=\"font-weight: 400;\">Sources</span></h3>\n<ol>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Databricks, </span><i><span style=\"font-weight: 400;\">The State of AI Agents</span></i><span style=\"font-weight: 400;\">, January 2026 (12x production rate with governance, 6x with evaluation tools, 20,000 organizations analyzed)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">OutSystems, </span><i><span style=\"font-weight: 400;\">2026 Agentic AI Production-Scale Survey</span></i><span style=\"font-weight: 400;\">, Q1 2026 (~1,900 IT leaders)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Codiste, </span><i><span style=\"font-weight: 400;\">Enterprise Agentic AI Adoption Report</span></i><span style=\"font-weight: 400;\">, May 2026 (83% funded, 41% reached production)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Gartner, </span><i><span style=\"font-weight: 400;\">Predicts 2025: AI Agent Project Cancellations</span></i><span style=\"font-weight: 400;\">, June 2025 (40% cancelled by end of 2027)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Gartner, </span><i><span style=\"font-weight: 400;\">Agent Management Platform</span></i><span style=\"font-weight: 400;\"> category definition, March 2026</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Boston Consulting Group, </span><i><span style=\"font-weight: 400;\">From Pilot to Scale: The AI Value Equation</span></i><span style=\"font-weight: 400;\">, 2024 (10% algorithms / 20% technology / 70% data and process change)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Gravitee, </span><i><span style=\"font-weight: 400;\">2026 Agentic AI Security Report</span></i><span style=\"font-weight: 400;\"> (incident rate and detection-to-containment gap among ungoverned agents)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">IDC, 2024 (60% of organizations will fail to realize AI value by 2027 due to governance failures)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">PwC, </span><i><span style=\"font-weight: 400;\">2025 Global CEO Survey</span></i><span style=\"font-weight: 400;\"> (56% report zero financial ROI from generative AI)</span></li>\n<li style=\"font-weight: 400;\" aria-level=\"1\"><span style=\"font-weight: 400;\">Transcend, </span><i><span style=\"font-weight: 400;\">AI Development Lifecycle Governance Survey</span></i><span style=\"font-weight: 400;\">, 2025 (93% hit data quality or governance blockers)</span></li>\n<li aria-level=\"1\">Ernst &amp; Young, <em>EY survey: AI adoption outpaces governance as risk awareness among the C-suite remains low</em>, September 2025 (975 executives surveyed on AI governance)\n<p class=\"cmp-hero--press-release__heading\"> </p>\n</li>\n</ol>",
            "image": "https://nataliacummings.com/media/posts/2/Screenshot-2026-07-09-at-11.20.33-AM.png",
            "author": {
                "name": "Natalia Cummings"
            },
            "tags": [
            ],
            "date_published": "2026-06-25T19:38:25+02:00",
            "date_modified": "2026-07-24T19:35:13+02:00"
        }
    ]
}
