Skip to main content
    All posts
    7 min read

    AI Budgets Are Up, Returns Are Flat. In Other News, the Sky Is Still Blue!

    By Matt Dornfeld

    Bain & Company just released a survey of 951 companies, and the headline is worth sitting with for a minute: most firms saw AI cost savings land in the 0–10% bucket rather than the double-digit reductions they'd projected—in fact, only about 4% cleared 30%. Their conclusion was blunt. "The technology worked. The value didn't arrive."

    That's a significant statement from a firm that advises the companies making these decisions, and if you're a CEO or CFO looking at your AI spend heading into the back half of 2026, it should prompt a different question than the one most executives are asking. The question most people are asking is "how do we get more out of our AI tools?"... but the better question is "are we structured to capture value at all, given how we've built our AI stack?"

    This is a game of business transformation, not vanity metrics and performative salutes during internal meetings. I think most companies have built their AI stack the wrong way for what they're actually trying to do. And the fix is less about spending more, and more about a specific architecture decision that very few executive teams have had an honest conversation about yet.

    What "The Technology Worked" Actually Means

    When Bain says the technology worked but the value didn't arrive, they're describing something I've seen consistently across the companies I work with. The models are good. The pilots look great. The demos are genuinely impressive. And then reality sets in... as the cost-per-output at production scale doesn't pencil out the way anyone expected.

    Here's what tends to happen: a company identifies a high-frequency workflow, builds a proof of concept calling a frontier model (GPT-5+, Claude, Gemini, take your pick), demonstrates meaningful time savings in controlled conditions, then rolls it out and watches the token costs scale with usage in ways nobody fully accounted for. Meanwhile, hyperscalers are on track to spend well over $600 billion on infrastructure in 2026, up more than 60% year over year—and roughly three-quarters of that is going to AI—and they're going to recoup that investment through enterprise pricing. The companies passing that cost through to you are not losing sleep over your ROI math. So the problem has two sides: margin erosion from inference costs on one end, and productivity gains that don't survive contact with real operational complexity on the other.

    What Pinterest Just Did (And Why It's Relevant to B2B Leaders)

    Pinterest CTO Matt Madrigal described something recently that I think belongs in every executive AI conversation right now. His team took an open-source model called Qwen3-VL, stripped out its vision encoder layer entirely, and rebuilt it with proprietary embeddings trained on Pinterest's own data. The result was a 90% cost reduction and 30% accuracy improvement over calling a frontier model directly... I mean, wowza.

    His explanation of why it worked is the part worth unpacking: "If you've got really unique data that you can then fine-tune an open source model with, data quality will, frankly, outweigh or overcome model size." Pinterest has 620 million monthly users and visual interaction data that no frontier model vendor has access to. They turned that data into a competitive moat in their inference layer, and it showed up immediately in both performance and margin.

    Now, Pinterest is a consumer tech company, and you might be thinking "that's not my world." But the underlying logic applies directly to any business sitting on proprietary operational data, which is most of the companies I come across in the wild. Put bluntly, aggregate, owned data is a valuable asset. And right now, most companies are treating it like context they throw into a prompt, rather than something they can use to build a model that performs better and costs less.

    The Conversation Most CFOs Haven't Had Yet

    The typical enterprise AI budget conversation right now goes something like this: here's our seat license spend for Copilot or Gemini for Workspace, here's our API spend for the applications we've built on top of frontier models, and here's what we expect to save in labor. All that said, the conversation that needs to happen is fundamentally different. IMO, it's moreso about whether the company's data assets are being structured to train and fine-tune purpose-built models, or whether they're just fueling someone else's model with context at every API call. Those are very different cost structures at production scale, and they produce very different performance curves over time.

    This isn't an argument for every company to stand up a machine learning team and start training models from scratch. I'm not standing atop a mountain yet claiming everyone should go back to their old, on-prem servers... but I'm not not saying some of this either. The open-source model ecosystem has made fine-tuning more accessible than it's ever been, and the models you'd be customizing (Qwen, Llama, Mistral derivatives) are genuinely excellent starting points. The decision is about whether you treat your proprietary data as a fine-tuning asset, or continue paying inference costs on a model that doesn't actually know your business.

    Where to Start

    For most executive teams, the practical path forward starts with an honest audit of two things: where your AI spend is actually going today, and which workflows have the highest call volume against external models. High-frequency, high-volume use cases are where the economics of custom tooling show up most clearly. A customer service workflow running 50,000 inferences a month is a much better candidate for fine-tuning than a quarterly board report summary. If you can identify two or three of those high-frequency workflows, and you have meaningful proprietary data that's relevant to them, you have the building blocks for a very different cost structure.

    The second question is data readiness. Fine-tuning only works if your data is clean, structured, and representative of the outputs you're trying to get the model to produce. This connects back to the data layer work I talk about in other contexts... companies that have invested in data governance and clean infrastructure get compounding returns (yay!). Companies that haven't find that every AI initiative, including custom tooling, stalls out at the data step.

    The Bain numbers shouldn't be read as "AI doesn't work." They should be read as a signal that the default approach, which is buy seats, call APIs, and hope the productivity math works out, has a structural ceiling. The companies that close the gap are the ones that treat their data as a proprietary asset and build their AI stack around it.

    If your AI spend is growing and your margins aren't moving, let's talk through the architecture decisions that close that gap. You can reach me here.

    Frequently Asked Questions

    Want to talk through what this looks like for your org?

    Share where you are and Matt will dig into it with you directly.

    Get in Touch