👋 Hi, it’s Kunal, and welcome to the Insider Growth Group newsletter, our bi-weekly deep dive into the hidden playbooks behind tech’s fastest-growing companies.
Our mission is simple: We help you create a roadmap that boosts your key metrics, whether you're launching a product from scratch or scaling an existing one.
What We Stand For
Actionable Insights: Our content is a no-fluff, practical blueprint you can implement today, featuring real-world examples of what works—and what doesn’t.
Vetted Expertise: We rely on insights from seasoned professionals who truly understand what it takes to scale a business.
Community Learning: Join our network of builders, sharers, and doers to exchange experiences, compare growth tactics, and level up together.
Introduction
Cursor, Claude Code, Replit, and Lovable were all built on models that would embarrass you today. Go back and use a GPT-4-era model for anything serious and you will wince. Yet somehow, the lack of intelligence in the models did not slow any of them down.
The usual explanation is that the models caught up. A better one is that those products compensate for the model’s intelligence with the user’s.
A coding tool is iterative. You try something, it is wrong, you nudge it, you go again, and the human absorbs the model’s mistakes as a matter of course. A phone call has no such moment. Neither does a bank wire that has already left. Build a one-shot experience and you are betting everything on the model plus the harness you put around it, because there is no human iteration to fall back on.
So the human in the loop is not a safety net you bolt on at the end. It is a quantity you are spending, and you can run out of it. On the one hand, you can leverage human intelligence to compensate for lack of model intelligence. However, if you rely too heavily on it, it can have adverse consequences.
That is one of several answers to a question Kunal Datta has had to face twice, in two industries where being wrong has real consequences: How do you know your AI is actually working, when working is not obvious from the outside? Both times he was answering it before the vocabulary existed for what he was doing.
At PG&E he ran the organization behind an AI wildfire inspection suite where every feature was assigned a level of automation, zero through five on the scale used for self-driving cars (image below), and every step up that scale had to be cleared with a regulator. It was filed as a patent in 2019, before GPT-3 shipped. At Unit21 he is Chief Product Officer. The company started seven and a half years ago building software for compliance teams to build, test, and iterate on their own detection logic and workflow, no engineers required. Today their AI Agents for financial crime are the main things they sell, scored against years of human investigations since before most teams had heard the word eval, and running in production at more than 200 global banks, fintechs and credit unions.
In both places, almost nothing that worked came down to model quality. Not one of the seven learnings below is “get a better model.” They are about what your users already produce, how much you are allowed to automate and when, what your reviewer can absorb before they stop looking, and what you measure when the thing you actually care about cannot be measured. Both teams arrived at most of them by getting them wrong in production first.
This is how to build an AI product when the stakes are high. Each of the seven is worked through twice below, once on a power grid and once on a payment rail, with the specific calls Datta would make again and the ones he would not.
Your users may already be labeling your data as a side effect of their job. Go and collect it.
Inspectors drew boxes around defects. Compliance teams graded finished investigations. Neither was built as training data and neither cost anything.Decide how much each feature is allowed to decide, and give every level of autonomy its own targets.
Most teams report a single accuracy number for a product that needs six.A tool people trust completely is as dangerous as one they ignore.
Too many flags and reviewers stop looking. Too few mistakes and they stop checking. Both end with a person who is no longer really there.You may never be able to measure the thing you actually care about. Pick your stand-in carefully and watch how it moves.
Nobody can count the fires that did not start, or the crime that nobody found. That doesn’t mean that there aren’t metrics you can target.Fix the scaffolding before you touch the AI.
The true constraint is almost never model accuracy, and the slowest step is rarely the one your team is most excited to automate.Your model and your users fail the same way, through context overload.
Hand either one too much and it stops discriminating. Humans have a context window too.Narrowing is not a compromise. Rather, it’s an unlock.
The version that found product-market fit was the most constrained one, even if it’s not the version with the best initial story. Generality can come back later, once the initial value is unlocked in a constrained use case.
Introducing Kunal Datta
Kunal Datta is Chief Product Officer at Unit21, an AI risk platform used by more than 200 global banks, credit unions, fintechs, and crypto companies to detect and investigate financial crime. He went from Senior Product Manager to Head of Product to CPO in under four years.
Before Unit21 he led Checkout at Fast, and before that built a prepaid card product from zero to one at a payments startup in San Francisco, spending time in Ukraine, Vietnam, Turkey, and the Middle East enabling access to online payments in primarily cash-based economies.
As a Fulbright Scholar he lived in villages in northern India researching rural electrification, installing GSM-connected solar nanogrids with prepaid payment infrastructure built in, and studying why the government’s electrification claims did not match reality on the ground.
The question underneath that research, how you actually deliver something as trust-critical as electricity and know that you have, is what drew him to Pacific Gas & Electric in 2016. He spent five years there. He started as a business analyst, moved into a product manager role created specifically to prove that AI was worth investing in, and eventually led a thirty-person global product, design, and engineering organization that built the utility’s first AI-driven aerial wildfire inspection system, now used to inspect electrical infrastructure at scale. He is the first named inventor on the two patents covering it.
He is a Stanford graduate in Civil & Environmental Engineering and Music, Science, and Technology. He teaches yoga and writes a Substack on Indian philosophy.
He’s one of a small number of product leaders who’s had to answer the same fundamental question twice, in two industries where the wrong answer has real consequences: how do you know your AI is actually working, when “working” isn’t obvious from the outside?
Case Study #1: How do you know you prevented a fire?
Context
On 8 November 2018 the Camp Fire destroyed Paradise, California and killed 85 people. A corroded C-hook on a transmission tower had failed, dropping a live line into dry brush.
What followed was not a normal corporate year at PG&E. Protests outside the offices and inside them. Erin Brockovich coming into the building to read out the names of the dead. The CEO resigned that January, and the company filed for Chapter 11 about two weeks later. It did not emerge from bankruptcy until July 2020, which means almost everything described below was built by a team working inside a bankrupt utility.
Datta is unequivocal about the people who did it. Among the smartest, highest-agency and most mission-driven he has worked with anywhere, before or since, which is not what most of tech likely assumes about their local utility.
He had arrived two years earlier, onto a four-person team sharing a single cubicle. It was the budding Data and Analytics group, run by a visionary director with one ruthlessly focused idea: the obstacle to modernizing the utility was not strategy, it was data. In roughly two years those four people shipped a computer vision model that read gas meters from a photograph, another that spotted cross-bores, which are gas lines accidentally drilled through sewer pipes and left there until a plumber’s auger finds them, a natural language model that predicted excavation strikes - when someone digging accidentally hits a gas line and leads to an explosion, usually fatal - from call-center transcripts, anomaly detection over grid sensor data, a graph database that let safety engineers search years of incident reports for patterns nobody had connected, and more. This was 2016 to 2018, before anyone called it AI.
That cubicle eventually became a 200-person organization building data products, mobile apps, LiDAR-based vegetation management and, under Datta, AI-driven wildfire safety inspection. His own org ran to about 30 people across product, design, engineering, data engineering, data science, and ML engineering.
The arithmetic they faced was the actual problem. PG&E is the largest utility in the United States: roughly 2.5 million distribution poles and 120,000 transmission structures. Inspection ran at about 50,000 assets a year with 120 inspectors. Call it two percent of the fleet, and at the current rate, would cost tens of millions of dollars per year. Everything else sat on cycles of one, two, five, sometimes twenty years depending on where it stood. Drought was pushing high fire threat districts wider every season, and the most experienced inspectors were aging out faster than apprenticeships could replace them. One of them, still working in his late eighties, had helped build Saudi Arabia’s first electrical system.
The problem was becoming increasingly urgent, and so the vision was to move from periodic inspection to something closer to continuous. Here is what that took, in the order it happened.
Fix the scaffolding before you touch the AI
In December 2018, a month after the fire, PG&E stood up an aerial inspection program. Fly drones and helicopters, capture high-resolution imagery, let inspectors review the equipment remotely. Datta was brought over from the innovation center shortly after, with one designer, to sit with inspectors and work out where it hurt.
The program’s first version ran on nine tools. Excel to track the work, GIS for asset data, a NAS drive for storage, Word for findings, Windows Image Viewer to look at photographs, a PDF form, a desktop notepad, a paper manual so nobody forgot the process, and physical hard drives in labeled envelopes.
What replaced that stack over the next few years came to be called Sherlock - the product suite inspectors now use to review imagery, named because inspectors said they wanted to be detectives, not clerks. Its computer-vision engine, which comes in a few sections down, is called Waldo, as in Where’s Waldo.
Image caption: Remote inspection process pre-Sherlock
Flying the drone was the fast part. Images came back on hard drives carried physically into the San Ramon office, uploaded into a particular folder structure, then sealed into a labeled envelope and locked away for the record. Inspectors pulled folders down to their own machines and paged through them one at a time, deciding by eye whether a piece of equipment was a fire risk. Drones covered only a subset of structures, because covering more was too slow.
The team built a few small models to prove feasibility of the big vision and sell the story, then left them alone. Process-mapping the full workflow, seven or eight discrete steps each owned by a different person, showed that the constraint was data ingestion, not model accuracy. So the first real production investment went there.
By late 2021 five processes were fully automated. Inspection time had gone down by more than 50%. Time to inspection - the gap between a photograph existing and a human looking at it - had reduced by 85%.
None of that was the model.
Decide how much each feature is allowed to decide, and give every level of autonomy its own targets
Before any computer vision reached production, the team graded inspection automation from level zero through five, borrowed directly from the scale used for self-driving cars. Level one flags standard items. Level two flags standard items and failures. Level three passes only bad images to a human at all. Level five has no human.
Every step up that ladder had to be cleared with a regulator, one rung at a time. The framework was filed as part of a patent in 2019, before GPT-3 shipped.
What made it work in practice is the part almost nobody copies. Each feature carried its own target. High precision for imagery QA. High recall for prioritizing the inspection queue. Medium on both for the inspection profile. Better than human performance for anything approaching selective inspection, which nobody had reached. One product, six surfaces, six different definitions of good.
Narrowing is not a compromise. Rather, it’s an unlock.
They started on standard items. Overview shots, asset tags, right of way, access paths. Paperwork, essentially, and deliberately so: low risk meant the team could learn on it without a mistake becoming a fire. Failures came later, once the ladder had been climbed and there was evidence to climb it with.
Image caption: An inspector using Sherlock to inspect a transmission structure.
Your users may already be labeling your data as a side effect of their job
Inspectors were drawing boxes around defects because that was the job. Nobody had to be asked, and nobody had to be hired separately to do this. All the team had to do was build a product that would naturally capture the labels as a part of the core job done by its users.
Image caption: Screenshot of Sherlock, showing how an inspector can put a box around a piece of equipment - a part of their normal job. The product forced good labeling practices, including dropdowns on the left for standardization of data in a pre-LLM world.
Those boxes became training labels. The models trained on them produced predictions that fed back into the review tool, which produced more boxes. The team then automated the labeling jobs, the labeling QC, the generation of first-pass models and the deployment of them. The virtuous cycle of data was the core thing to build - the model was an outcome.
Image caption: The virtuous cycle of data.
A tool people trust completely is as dangerous as one they ignore
Only a small percentage of A-tags, the most serious findings, were being caught at imagery QA - the first of several human eyes to see a given image. The rest surfaced later, during the actual inspection or QC steps later on in the flow. Catching more of them earlier meant fixing problems sooner, which meant less time for a defect to become a fire.
So: precision or recall? You cannot have both, and the training data made the tradeoff sharper, because real failures are rare and labeled examples were thin.
Image caption: A primer on precision vs. recall.
After considerable discussion with senior leadership and with regulators, everyone agreed on recall. Missing something is a fire.
They released it to a small group rather than to every inspector, because Datta had been reading the literature on automation bias and expected the second-order effects to be strange.
Imagery QA time went up sharply. Inspectors handed a flood of flags that were mostly wrong got fatigued, stopped trusting the tool, and the share of A-tags caught at that stage fell to by about a third. This was below where it had been with no model at all.
The vocabulary matters here, because two opposite things can go wrong. Disuse comes from mistrust: the reviewer ignores the tool or switches it off. Misuse comes from overtrust, and splits again into errors of omission, where the model misses something and the human does not notice, and errors of commission, where the model says something and the human simply believes it. Training reduces commission errors. It does not reduce omission errors. That one has to be designed for.
Switching to a high-precision model more than 2x the rate of findings at that first stage from the initial baseline.
This is why target per-feature exists rather than a single accuracy target. A false positive costs fatigue and inspection time. A false negative can cost a fire. And being right too often is its own failure, because it produces complacency. There is a rate in between those, and it is a fact about your human users rather than about your model.
Your model and your users fail the same way
Hand an inspector two hundred images with every marginal defect flagged and they behave exactly like a model given too much context. They stop discriminating. The fix in both cases is the same: decide in advance what this particular task is allowed to see.
You may never be able to measure the thing you actually care about
There is no counterfactual for a fire that did not happen. So the team built stand-ins and created a process to measure them that added value along the way.
For example, give two inspectors the same tower without telling either that anyone else is looking, then compare what each one found. Run it across a few hundred structures and the gaps stop looking random. Certain cohorts consistently under-flagged vegetation encroachment. Others missed guy wires, the support cables that hold a structure upright.
The proxy had found a training problem in the humans, not a defect in the model. That is the sign it was working.
Where it ended up
Six distinct profiles sat underneath, because imagery QA, drone data ingestion, the inspector, post-inspection QC, supervisors and search all needed different views of the same work.
While Datta was there the program covered PG&E’s transmission towers. It has since expanded to distribution, the wooden poles that run down ordinary streets, and PG&E’s public wildfire mitigation filings track that scale-up as a primary mode of inspections today.
If you are building an agent from scratch
Pitch coverage, not efficiency, to the people who will use it. The business case included cost, and the program did eventually save real money. But what inspectors - and perhaps more importantly, regulators - heard was “inspect more, inspect more often.” A tool that closes a safety gap earns buy-in. The same tool pitched as an efficiency play invites resistance before anyone has tried it.
Stage the rollout on purpose, not by accident. The team caught the trust collapse while it was still small enough to reverse, and they caught it because they had already staged the deployment and were already measuring find rate at each stage.
Write down what “good” means per feature before you build any of them. One accuracy number for a product with many surfaces is how you optimize one surface into the ground without noticing. Consider the consequences of the metric - you get what you measure, but you may not get what you want.
Case Study #2: Effectiveness > Efficiency
Human trafficking. The online sexual exploitation of children. Fentanyl. Firearms. Scams that empty an elderly person’s savings account. Political corruption.
None of that happens for free. All of it has to be funded, and the money has to move, which means it moves through banks, credit unions, fintechs, crypto exchanges and payment processors. That is why anti-money laundering regulation exists at all. Following the money is the most reliable way to find the crime behind it, so the institutions that move money are required by law to go looking.
Which makes this a detection problem, and detection is a data problem.
Here is how hard it gets. Think about laundering money through eBay. Set up a few accounts. Have one list a rare Pokémon card. Have the others bid it up, higher and higher, until the winning bid is a million dollars. The sale goes through, money moves, and on paper it looks like two parties trading a collectible.
Now try to catch that as eBay’s compliance team. The transaction record shows a sender, a receiver, an amount and a timestamp. It does not show what was sold, and it certainly does not show who else was bidding, which is the only thing that would tell you the auction was staged.
A bank sees standard shapes: a card swipe, a wire, an ACH transfer. Anyone else moving money, and the definition is broad enough to cover most of fintech and all of crypto, generates transactions that look nothing like that, and the signal that matters is often not in the transaction at all.
Unit21 was built on that gap seven and a half years ago, on the premise that a compliance team should be able to build, test and iterate on their own detection logic against whatever shape their data actually arrives in, without engineers. Built for humans, in other words. Every part of it assumed a person doing the looking.
Then the models got good enough to ask how much of that looking something else could do. Seven versions later, AI Agents are the main thing the company sells. The first four did not work, and they are more instructive than the ones that did. Here is what it took, in the order it happened.
Image caption: Network analysis in Unit21, showing a fraud ring.
Fix the scaffolding before you touch the AI
The first attempt, in early 2023, was the obvious one. Dump a batch of transactions into a GPT model and ask it to find the fraud. It was close to random.
But something surfaced. The model could write SQL and Python, not just describe answers in English. That became a strategic bet rather than a feature. Every frontier lab is pouring effort into code generation, because that is the road to the thing they are all driving at, so a problem framed as code generation improves every time a new model ships, whether or not you do anything. Frame it any other way and every gain has to come out of your own engineering.
The bet was sound and the product still failed, which is the next section. But none of it would have worked without the seven years underneath. The rules engine already accepted arbitrary data shapes. Investigations were already structured. Filings were already tracked. None of that plumbing was built in anticipation of AI. It was built because people needed to do the work, and it happened to record every human action in a shape you could score against later.
A tool people trust completely is as dangerous as one they ignore
The second attempt was called Ask Your Data. Ask a question in plain English about anything in the system, whether it’s individuals, business entities, logins, password changes, accounts, transactions, and more, and get an answer back by having the model write code that then got run.
Image caption: A screenshot of an early version of the Ask Your Data product at Unit21 - the first production use case of AI back in 2023, since deprecated.
It failed on trust.
One user asked for a data export the product did not support. Since the whole thing worked by writing and running code, handing back a file was a plausible-looking thing to do, so the model did not say the feature was missing. It produced a download link. The user replied that the link was broken, so it apologized and produced another one. Datta was reading these exchanges manually and watched it happen six times before stepping in. The user was remarkably patient. The model was confidently, helpfully wrong, over and over.
One answer like that costs more than ten good ones earn.
Nothing caught it, because there was nothing to catch it with. This was early 2023 and the word eval was barely in circulation, so the team worked it out from first principles. If you are going to put a model in front of someone, you ought to check what it does before you do. They started scoring real analyst questions against expected answers in a spreadsheet. That spreadsheet is the ancestor of everything that followed.
They pointed it back at Ask Your Data and quality went up. It still was not enough to trust, which is its own lesson: an eval tells you how good something is, not whether good is good enough for what you are asking it to do.
Being wrong here is not symmetrical and neither direction is abstract. A false positive is somebody trying to send money home for a family member’s hospital bill, blocked because their name resembles a name on a list and nothing in the system could tell the difference. A missed case is a report that never gets filed, an investigation that never starts, somebody still being trafficked six months later. And a system people trust completely is worse than either, because then nobody is checking in either direction. Both failures keep happening. They just happen quietly.
Your model and your users fail the same way, through context overload
The third attempt used a feature that already existed. Checklists let investigators work through a customizable list of steps on a case, some with free-text fields, a bit like filling in a form. The idea was to have the model fill them in: checklist question as the prompt, case data as the context.
Image caption: A screenshot of the checklist version of Unit21’s AI product, from early 2024, since deprecated.
It failed badly, for two reasons the team only had names for later.
The first was prompt engineering, or the absence of it. The people writing checklist questions had written them for other humans, who arrive with years of context and know what “check for structuring” means in practice. A model does not.
The second was context engineering. The volume of data attached to a given case was simply too much, and you could watch the model get confused and answer worse as more of it arrived. Exactly what had happened to PG&E’s inspectors, on the other side of the screen.
Narrowing is not a compromise. Rather, it’s an unlock.
Worth remembering what Unit21 is for. The company exists because compliance teams hold data nobody else holds, in shapes nobody anticipated, and need to monitor it their own way. Flexibility is not a feature here. It is the product.
Which is why Ask Your Data went straight at it. Ask anything, about anything. That was the right destination and the wrong first move, because nothing existed yet that could tell an analyst whether a given answer was worth acting on.
So the next attempt went the other way. A fixed library of known-good investigation types. Behavior deviation analysis. On-chain transaction analysis. Document analysis. Each with a pre-defined slice of the data, only what that task needs, and its own tested prompt structure. Every one had to earn its place before reaching a customer, scored against years of already reviewed investigations. Instead of one model that would attempt anything, many narrow paths that had each been proven separately.
Image caption: Unit21’s first AI Agent product that found PMF, back in January of 2025. This screenshot shows the configuration page with the out-of-the-box task library, built upon years of human investigations as eval sets.
That is the version that found product-market fit, in January 2025.
But narrow was never the destination. It was the detour that built what customizability had been missing. Once every task was validated against history as a matter of course, the machinery for validating any task existed, so the flexibility came back properly. Customers now build their own agentic tasks, backtest them against their own history, and deploy only what clears. They are building and validating their own evals inside the product, which is precisely what Ask Your Data never had. The next step extends the same move to whole agents.
Image caption: The latest in Unit21’s AI Agent creation - custom task creation. A manager or an admin can create “tasks” for an AI Agent to complete. A second agent will look through the available data, build up a knowledge graph mimicking humans hopping between screens, and iterate with you on the form of the task based on knowledge of FinCEN advisories, historical enforcement actions, and more institutional knowledge of financial crime. Below is the ability to backtest and create your own eval set within the product itself, to ensure the output works at scale - all self-service within the Unit21 product.
Your users may already be labeling your data as a side effect of their job
Here is why any of that scoring was possible.
Unit21 has a QA product. Compliance teams have used it for years to randomly sample completed investigations and judge whether they were done well. Nobody built it as training infrastructure. It existed because regulators expect quality assurance and analysts need feedback. Years of that is a labeled evaluation set nobody commissioned and nobody paid for.
Better than that, it is one customers pay for. The labeling is revenue-positive rather than a cost line, which is the difference between a flywheel you can afford to keep turning and one you cannot.
This is also what teams underestimate when they decide to build their own. Engineers shadow a compliance team for six months and build something that works, in the sense that it produces output. What they do not have is any way to find out whether the output is right. Building the model was never the hard part. Starting from zero labeled investigations is.
Decide how much each feature is allowed to decide, and give every level of autonomy its own targets
Automation is graded zero through five, and it is set per queue, by the customer.
At level zero there is no AI. At level one the analyst kicks off the review themselves. At level two the review runs automatically as the alert arrives. At level three the system recommends a decision, against criteria the customer configures and tests, but takes no action on it. At level four it closes false positives on its own, and only those. Level five, full automation, does not exist yet.
Image caption: Unit21’s Configure Actions panel, where a customer sets the automation level per queue.
Level one is the rung that gets cut in most roadmaps. Starting the review by hand is slower, and it exists on purpose: an analyst who reads the machine’s conclusion first tends to stop forming their own.
Detection climbs a separate ladder. It began at no AI, then writing a rule from a plain-English description, then insights, which tell you what a rule and its data are actually doing and what investigation outcomes suggest changing. Now there is an optimization agent that names the specific rule, writes it, tests it and reports back. The customer still tests it themselves and still puts it live themselves.
Two ladders, because detecting and investigating are different jobs with different costs of being wrong. Neither is near the top of its own.
You may never be able to measure the thing you actually care about
Ask what “working” means for financial crime detection and the answer gets uncomfortable fast.
How well you can measure recall depends on which side of the house you are on. Fraud gives you some signal back. People dispute transactions, chargebacks arrive, someone works out they were scammed and calls their bank. Late and partial, since plenty of victims never report and some never realize at all, but a feedback channel exists.
Anti-money laundering has none. Nobody launders money and then complains that you failed to notice. The people actually harmed sit several steps back from the transaction and have no idea it happened.
So Unit21 measures true positive rate instead: of all the alerts a rule generates, how many turn out to be actual money laundering, fraud or other bad activity. A precision measure on a fixed population, honest about being a stand-in.
Over roughly a year of production use, that rate came in 5x higher than before. The important detail is what did not change. The rules are the same. The alerts are the same. Nothing about what gets flagged moved, which makes the comparison close to controlled: the only variable is how thoroughly each alert gets worked.
Which means the old number was never measuring detection quality. It was measuring how much investigation capacity analysts had. Some unknown share of what this industry calls false positives were real cases nobody had time to finish.
That was not the goal. Going in, the assumption was efficiency, with large outsourced review teams cut roughly in half. What arrived was coverage. Same caseload, more crime found inside it.
Where it ended up
More than a hundred leading global banks, fintechs, crypto companies and credit unions run Unit21’s AI Agents in production today, on cases involving human trafficking, fentanyl trafficking, scams and elder abuse.
Underneath, different tasks route to different models, picked on benchmarks that keep moving, because whatever model leads on optical character recognition is not what leads on behavioral analysis, and neither stays in front for long. Model selection turned out to be a per-task decision with a shelf life, not a procurement decision you make once.
If you are building an agent from scratch
Your buyer, your user and the person who signs off are three different people. A Head of Fraud or a Chief Compliance Officer buys it. An analyst uses it. An examiner has to accept that the process is defensible, which is a separate question from whether it works. PG&E had the same shape: inspectors as users, an executive team holding budget, unions with a view, a regulator approving every rung. The constituency most teams build nothing for is the one that can quietly make the product unusable.
Iteration only helps if your user will iterate. Every version that failed here asked an analyst to behave like a developer: phrase a question well, read a wrong answer, work out why, try again. Analysts have fifty alerts in front of them and a queue that does not care. Out-of-the-box tasks worked because the iterating had already happened on Unit21’s side before anything reached a customer.
Prove it against history before you ask anyone to trust it live. Every task type was backtested against verified data before it reached a single customer. Proof before promotion.
🔥 Insider Growth Playbook: Building AI When Being Wrong Has a Cost
Seven learnings, two industries, one grid and one payment rail.
Here is how to run each of them against whatever you are building.
1. Go and find where your users are already labeling your data
PG&E’s inspectors drew boxes around defects because that was the job. Unit21’s compliance teams graded completed investigations because regulators expect it. Neither was built as training infrastructure and neither cost anything.
Where to look. A moment where a person already makes a judgment, as part of their job, that somebody else could later check. Not a survey, not a labeling exercise you commission. A support agent closing a ticket with a resolution category. A sales rep marking a deal won or lost, and why. A moderator actioning or clearing content. An underwriter approving or declining with a reason code. Anyone accepting, editing or rejecting a draft.
The test has three tiers. If you would have to pay someone to produce it, this is not it. If it is already being produced and you are dropping it on the floor, it is. And if a customer would pay you for the product in which they produce it, you have stopped looking at a data strategy and started looking at a business.
Two things to get right. Capture the correction, not just the decision, because a human editing a model’s output is a far better label than a human agreeing with it. And capture it in a structured field, not free text nobody will parse.
If you are early enough to have no history, this is the highest-leverage item on your roadmap and it will not look like AI work to anyone reviewing it.
2. Build your ladder, and give every rung its own target
The half-day exercise. Take one feature. Write five rungs, where each is defined by who carries the consequence rather than by how capable the model is. Then mark three things: the rung you are actually on, the rung your marketing implies, and the rung your roadmap quietly assumes. Those are usually three different rungs, and the gap between the first two is where trust goes to die.
Then give each rung a target, per feature. PG&E ran high precision for imagery QA, high recall for queue prioritization, medium on both for the inspection profile. One product, six surfaces, six definitions of good. If you are reporting a single accuracy number for a product with several surfaces, you are optimizing one of them into the ground without noticing.
Include one rung that is slower than the rung below it. Unit21’s level one makes the analyst start the review by hand, so they form a view before the machine hands them one. If nothing in your ladder trades the user’s time for keeping their judgment intact, you have built a capability ladder rather than a trust ladder.
3. Aim for the right amount of trust, which is not the maximum
Two opposite things go wrong, and they have names. Disuse comes from mistrust: the reviewer ignores the tool or turns it off. That is what happened when PG&E flooded inspectors with flags. Misuse comes from overtrust, and splits into errors of omission, where the model misses something and nobody notices, and errors of commission, where the model says something and the human believes it.
Training reduces commission errors. It does not reduce omission errors. That one has to be designed for, not taught away.
Three things to do.
Write evals for the failure you have not imagined. Almost every eval set tests whether answers are correct. Almost none test whether the model will claim a capability the product does not have, which is exactly what took down Ask Your Data, and it is the failure that destroys trust fastest because the user cannot tell it is happening.
Make the override path as easy as the accept path. If accepting is one click and overriding takes five, you have not built a human in the loop, you have built a rubber stamp with extra steps.
Find the rate in between by running something. PG&E ran experiments on a number of features, measured by miss rate between inspector and QC, because nobody could reason their way to it in a room.
4. Name your unmeasurable, then design the stand-in
You cannot count the wildfires you prevented. You cannot count the crime nobody found.
Three questions, none of them about your model.
What is the thing you actually care about, stated plainly, that you cannot measure? Say it out loud. Most teams never have, which is how the proxy ends up chosen by accident.
What is your stand-in, and what would make it lie to you? Unit21’s true positive rate moves when investigation depth changes. That is the point there and would be a serious bug somewhere else.
Is your metric measuring your system, or measuring how much attention your people have left? Unit21’s moved several times over without a single rule changing. That was the whole discovery.
Then build the check. PG&E gave two inspectors the same tower without telling either that anyone else was looking, ran it across thousands of structures, and found cohorts consistently missing vegetation and guy wires. The proxy surfaced a training problem in the humans rather than a defect in the model. That is the sign it is working.
5. Fix the scaffolding before you touch the AI
PG&E mapped seven or eight process steps and found the constraint was data ingestion, not model accuracy. Inspection time halved and time-to-inspection dropped by most of itself before the interesting model work started.
Run this in a week. Diagram every step of the current process and name who owns each one. Time them. The slowest step is rarely the one your team is excited to automate. Then ask what you would gain by fixing only the slowest, least interesting step, before any model ships.
And if you do not have years of structured history, note what the plumbing at both companies actually was. Nobody built it for AI. It was built because people needed to do the work, and it recorded them doing it. You are recording something right now. Whether it is in a shape you could score against in three years costs almost nothing to decide today and is close to unrecoverable later.
6. Your reviewer has a context window too
Checklist autofill broke because the whole case file went to the model. High-recall Waldo broke because every marginal flag went to the inspector. Same failure, two kinds of processor, one fix: decide in advance what this particular task is allowed to see.
Run it on your backlog. For every job you want an agent to do, write down the exact slice of data it needs. If you cannot, you do not have a task yet, you have a wish. Then run the same test on the human screen next to it, because a reviewer buried in context degrades the same way and almost nobody instruments for that.
7. Ship narrow, and let your evals set the ceiling
The version that found product-market fit was the most constrained one. PG&E started on paperwork-type items on purpose, because low risk meant they could learn without a mistake becoming a fire.
Reframe the sequencing question. Most teams ask how flexible version one should be. Ask instead what you would need to be able to prove before flexibility is safe to offer, then build the narrow thing that produces that proof. Narrow is not the product. It is the instrument.
Before a task type ships, it needs: a pre-defined slice of data it is allowed to see, a backtest against already-verified outcomes rather than a vibe check, a named owner who can say in one sentence what good looks like for that task specifically, and a rollback plan for when the backtest score drifts.
Two things that sit underneath all seven
Iteration only helps if your user will iterate. Coding tools survived weak models because developers try, fail, adjust and go again. Most people are not like that, and it is not a failing. An analyst with fifty alerts wants the thing to work. So ask whether your user will iterate, not whether they could. On a Tuesday, with a queue in front of them. If the answer is no, somebody still has to do the iterating, and it can be your team or a builder-type inside the customer, but it cannot be the person you are building for.
Effectiveness first. Efficiency follows. Nobody on the receiving end of either of these programs wanted efficiency. Inspectors wanted to catch more defects. Investigators want to find more crime. BSA officers, examiners and the people who sign filings are measured on whether the bad thing was found, not on how fast the queue cleared. Offer them a tool that does the same work with fewer people and you get a polite meeting and no adoption. Both programs moved when the framing changed to effectiveness, and the efficiency arrived anyway. Lead with it and it reads as a headcount plan. Let it follow and it reads as a bonus.
Who all of this is actually for
Somewhere behind both of these systems is a person who will never know either of them exists.
At PG&E it is whoever lives in the house that did not burn down. At Unit21 it is whoever got money home for a hospital bill, and whoever is not still being trafficked six months from now because a report was filed in time. None of them asked for any of this. None of them will ever see a dashboard.
Unit21’s product principles have a line about that, and it is not the one most people expect. Not empathy. Compassion.
Do not simply take on the pain of the person at the end of the line - work out what you can actually do about it.
Want help applying this to your product?
Let’s talk


















