Before you read this post I wanted to note that I do not cover Objectives in this article. The reason is that without context it’s difficult to provide an example that would resonate broadly in a short blog post. So this article focuses solely on key results. Don’t forget to write an aspirational and inspirational objective to go along with your AI-tool OKRs.
Earlier this year Josh Seiden and I worked with an AI team at a big bank in Europe. As you might imagine we spent a good chunk of time with them discussing how they’re measuring the success of their new AI-powered services. Many of these tools were internally-facing, targeting the staff of the company as the target audience. “Achieve 95% accuracy in all output.” The team was all bought into this number because it felt rigorous, the kind of number a serious product team commits to.
It’s also an output. And once again, even with AI products, we’re seeing the same mistake teams have been making with OKRs for fifteen years.
I’ve argued for a long time that a key result has to be a measure of human behavior. Not an output, not a task, not a feature you shipped but rather what a customer does differently because of the work you put out into the world. AI doesn’t change that rule. If anything it makes the rule harder to follow, because the thing your feature produces is probabilistic and you don’t fully control it. You can’t put “ship the summary” on a key result, and you can’t honestly commit to a single accuracy number either, because the model’s behavior will drift from user to user and from prompt to prompt. The goal moves where it always should have been in the first place, onto the human on the other side of the output. Here’s how that plays out across three types of key results, all of them measures of behavior.
The outcome KR: what users do after the AI answers
The first and most important question is deceptively simple. Once your feature hands the user an output, what do they do next?
That next action is your real signal. When the internally-facing bot provides valuable, actionable and accurate data, the user takes the data, exports it, shares it or integrates it into their deliverable and moves on with their day. When it doesn’t work, you see that in the behavior too. They regenerate the summary two or three times hunting for a better version (a re-query can certainly be a bad sign). They abandon it and open a blank doc to type their own notes. They export it somewhere else to quietly rewrite it before anyone sees it. Each reaction to the output is indicative of the value of the output to that specific user with that specific query.
The outcome KR names the good behavior and puts a number on it. For example, it can be, “increase the share of meetings where the user sends or shares the AI created data without rewriting it from 40% to 65%.” As always, we are looking for the formula of, Who does what by how much? In this case,
Who: the user.
Does what: acts on the output without redoing it.
By how much: From 40 to 65 percent of the time.
You can pair that with a metric that keeps an eye on the bad behavior, like cutting the rate of double-regenerations from 22% to under 10%, but the headline is always a human doing something that proves the feature earned its place in their workflow.
The calibration KR: turning accuracy into a human behavior
I suspect this is where most teams want to sneak the accuracy number back in, and I get why. Quality matters. A feature that invents data or analysis nobody needs is worse than no feature at all. The calibration KR exists to hold that quality bar. The trick is to measure the bar the same way, through what the human does.
Ask the question directly. If the output is genuinely accurate and high quality, what does the user do differently? In practice they stop checking it. They stop correcting it. They accept the data output as written instead of opening the full transcript to verify every line, and they take any action items that come out of the data without editing them first. Given every user, regardless of prompt, is getting a different response, this is a way to make sure that, across the board, users are seeing answers they find to be accurate.
Once again, we can set a key result here based on this behavior. Reduce the percentage of sessions where the user opens the transcript to fact-check the data from 60% to 30%. The quality threshold is still there measured as the moment the user decides your output is good enough to use as-is. That decision point is an observable behavior, and for the most part it’s the best proof of quality that there will be an ROI on this AI-powered feature.
The trust KR: measuring reliance, not sentiment
Trust is the one everybody wants to measure with a survey. Aside from being notoriously difficult to do well, oftentimes the data from surveys is not very reliable nor indicative of actual behavior. People tell you they trust the tool and then quietly check everything it does. Real trust isn’t a feeling ultimately but rather a change in what someone is willing to hand over. So watch what they hand over.
A user who trusts the internal data analysis tool starts using it for the meetings that matter, the client calls and the exec reviews, not just the internal standups. They turn on auto-send to the team instead of reviewing every summary in private first. Their corrections drop, not because they got lazy, but because they stopped expecting to find mistakes. Each of those is a behavior you can observe and measure.
The cleanest trust KR is watching whether your users override the output. Keep the rate at which users manually override, delete, or rewrite the AI’s proposed analysis and next steps under 8%. You can set an expansion behavior next to it, like growing the share of users who enable auto-share to their team from 15% to 40%. Both answer the same question: when people believe the system, what do they let it do that they used to do themselves? That behavior is the manifestation of trust in the product.
Putting the three together
Put the three and you get an OKR for an AI feature that never once measures the output of the AI. One key result for what users do with the output, one for the behavior that high quality produces, one for what trust lets them hand over. The feature’s accuracy, its model, its calibration, all of it still matters enormously, but as the driver of human success, not the measure of it.
If you’re staring at an AI feature this quarter and the only number you can think to commit to lives inside the model, you’re measuring the machine instead of the human. Try this instead: write down the single most valuable action a user should take after your feature responds, and ask how much more often you can get them to do it. Start there. It’s always been the person on the other side of the product that tells us whether we’ve built something valuable, not the product itself.






Leave a Reply