image

Back to all articles

The Operating Model Thesis

8

min to read

The Audit-First Doctrine: AI Engagements Should Start with Audit, Not Build

Why do AI transformations fail when the technology arrives before the operating model does. Most projects break not on the model, but on data, workflows, and decision rights nobody mapped first. Carbon's view: audit before you build, or the AI just runs the dysfunction at speed.

The instinct in most AI transformations is to start with the technology. Identify the use case, select the platform, begin the build. In their case, the operating model gets addressed too late, usually when something breaks

A company's operating model, meaning the combination of process, data, decision rights, and dependencies that determines how work actually gets done, has to be audited and rebuilt before any AI system is built on top of it. Skip that step, and the AI system inherits whatever was already broken in the operating model, then runs it at speed.

Carbon's position is that the operating model has to be audited and rebuilt before any AI system is designed, not as a precondition that slows things down, but as the step that determines whether what gets built will still be running in six months.

Where the bottleneck actually sits

The AI industry frequently compares this moment to the shift to cloud computing over a decade ago. The comparison holds in one respect: the technology itself is no longer the constraint. The data most companies run on was never built for AI. Workflows shaped around human handoffs lack the structure automated systems need to run reliably, and pilots succeed under conditions production will never offer, which is why most of them never make the jump.

We made the change because we kept watching good builds fail for reasons that had nothing to do with the code. A model goes into production, produces clean and confident output for a few months, and then a number stops reconciling. Every time we trace it back, the problem is waiting in the same place: the business the software was built to serve does not run the way anyone described it.

Research from MIT's Project NANDA found that just 5 percent of AI pilot programs were extracting real, measurable value. The researchers attributed the shortfall to a persistent learning gap: most of the AI tools studied did not retain feedback or adapt to context, and struggled to integrate into the workflows they were dropped into. That finding deserves a caveat. The report reflects the AI tooling available in mid-2025, and some of the integration gaps it describes may narrow as agent memory and context-handling improve. The underlying condition it points to looks more durable: a tool cannot integrate cleanly into a workflow that nobody has mapped accurately in the first place.

Other research points the same direction. Gartner projects that 60 percent of AI projects lacking AI-ready data will be abandoned, tying much of the failure rate in agentic AI specifically to escalating costs, unclear business value, and inadequate risk controls. A Boston Consulting Group survey of 2,360 executives found that half say their own jobs are on the line if their company's AI investments fail, while a separate BCG study found only about one in four companies have found a way to turn AI experimentation into measurable value.

Why an audit, not a vision

Most consulting engagements in this space open with a vision exercise: workshops on the future state, a roadmap, a deck describing what the organization could become once AI is embedded across its workflows. There is nothing inherently wrong with that approach. It is simply sequenced too early, because a vision built on top of an unexamined operating model inherits every flaw in that model without anyone noticing until later.

Carbon decided to change that approach. Every AI Pod begins with discovery and scoping; pressure-testing a business case if the client already has one, or auditing the business to find the highest-value AI workflows and build the case from scratch if they don't. What gets built gets defined here, and so does how it gets measured. Operating models drift under normal business pressure, in every organization, regardless of how carefully they were designed, which is why a structured discovery phase finds real, useful information almost everywhere it looks. Skipping it removes the one step that would have surfaced that information before a system was built on top of it.

What discovery has to establish

Discovery is the most reliable way to see what a company actually runs, as distinct from what it assumes it works.

Consider a pattern that shows up often enough in these engagements to be treated as a template rather than an exception. A mid-sized company's claims or order-processing workflow is documented once, in a manual nobody has opened in years. In practice, the work has been quietly modified for two years by the three people who actually run it, because the documented version stopped working cleanly after a system migration that was never fully resolved. A decision formally assigned to one team has, in practice, moved to someone else's desk, because that person is the only one who understands the exceptions that come up weekly. None of this shows up in a demo. It shows up months into a deployment, when a system trained on the documented process starts producing confident, well-formatted output that is wrong in ways nobody catches until a number fails to reconcile.

Discovery has to establish four things before anything gets redesigned.

  1. The first is process reality measured against documentation. Every organization has an official version of how work gets done, and a version people actually execute under deadline pressure. Building against the official version in the scenario above would lock in a process the three people who actually run the workflow have already learned not to trust, while their real process kept running unnoticed underneath the new system.
  2. The second is data lineage and ownership. It is common for three departments to keep three versions of the same number, each correcting for the others' gaps, with no one certain which version is current. Discovery traces where the data originates and which version the business actually relies on.
  3. The third is decision rights. Every process has a formal owner and an actual one, and the two diverge more often than most leadership teams expect. Discovery identifies who genuinely holds approval authority today, and where that authority needs to sit once the process changes.
  4. The fourth is dependency mapping. No process operates in isolation. Changing one workflow shifts load, timing, or expectations onto whatever sits downstream of it. Mapping those knock-on effects here, rather than after a change has shipped, keeps a different team from absorbing disruption nobody warned them about.

Reset the operating model, then engineer the data

Carbon's methodology sequences the rebuild in a specific order: workflows, processes, and governance get redesigned before the data does.

The order can look counterintuitive next to the instinct to fix data first, since bad data is usually the most visible problem in any audit. Building a data pipeline before the operating model is redesigned, though, means engineering it against a process that is likely about to change. A workflow rebuilt around clear ownership and simplified handoffs often needs a different data structure than the broken version of that workflow did. Redesigning the operating model first establishes what the business needs from its data, so the engineering work that follows only needs to be done once.

For decades, enterprise data was collected for dashboards, audits, and compliance reporting, rarely for AI systems that need it structured and reliable in real time. Once the operating model has been rebuilt around how AI actually needs to work, that data gets engineered into a form the system can use without inheriting the confusion of three departments each keeping their own version of the same number.

Only after both of those steps does the system get built. Measurement runs from day one of that final phase, using a baseline that was fixed during discovery, before any system was designed. Cost reduction, execution speed, and output quality all get reported against those agreed metrics. An AI system built ahead of this sequence inherits the underlying damage at scale. The instinct to deploy fast and iterate, often treated as a virtue in software development, works against operating model problems specifically. Iteration is useful when the foundation is sound and the goal is refinement. When the foundation itself is the problem, each iteration further automates the dysfunction.

There is a sharper finding in the MIT research about who runs this kind of rebuild. It found a wide gap in outcomes between AI efforts built internally and those implemented through external partnership. Tools purchased from specialized vendors, or built through partnerships, succeeded about 67 percent of the time. Internal builds succeeded roughly a third as often. Internal teams understand their own business deeply, but that depth is not the same as having run this exact sequence across other operating models, under other conditions, and that prior experience is often what determines whether the sequence gets followed under pressure or quietly skipped when a deadline tightens.

A finite engagement, a permanent system

This connects to a question Carbon treats as central to organizational design in the AI era: what should a company own outright, and what should it depend on a partner to run.

A company that does not understand where its operating model is underperforming cannot make an informed decision about what to rebuild internally. what to embed AI into, and what to hand to a partner structured to own the outcome. Discovery is what makes that decision possible in the first place.

It is also why the engagement itself is built to end. A Carbon AI Pod closes when the transformation holds, measured against the business case agreed to at discovery, with the client owning and operating the system outright from that point forward. There is no recurring fee tied to Carbon's continued involvement and no platform lock-in requiring Carbon's software to keep the system running. The engagement is finite; what it leaves behind is not.

Audit first. Build to last.

A completed engagement built on this doctrine has a specific, repeatable shape: a defined business case agreed before a single system is designed, an operating model rebuilt with clear ownership, data engineered to support what the business actually needs, and a system deployed against a baseline set before any of it began.

The research reviewed here points in one direction. Outcomes track sequencing. Discovery gives a company visibility into its starting conditions before committing a budget it can't get back, and it puts that company in the bracket the data shows is more likely to succeed.

Our objective isn't to create another dependency the business has to manage. It's to produce an operating model and an AI system the business owns and runs independently. The engagement is built to be finite, with one clear goal: a business running on a foundation stable enough to hold once we're gone.

Carbon builds nearshore engineering Hubs and embeds AI capability inside client operations, for scaling technology companies and PE-backed organizations. Operational infrastructure, built to last.

Own the Build™

image

The Carbon Team

Published on

Own the build.
Start with a conversation.

Get in touch

Own the build. Start with a conversation.

Get in touch