How to write OKRs for an AI product

An example of key results for an AI product

Before you read this post I wanted to note that I do not cover Objectives in this article. The reason is that without context it’s difficult to provide an example that would resonate broadly in a short blog post. So this article focuses solely on key results. Don’t forget to write an aspirational and inspirational objective to go along with your AI-tool OKRs. 

Earlier this year Josh Seiden and I worked with an AI team at a big bank in Europe. As you might imagine we spent a good chunk of time with them discussing how they’re measuring the success of their new AI-powered services. Many of these tools were internally-facing, targeting the staff of the company as the target audience. “Achieve 95% accuracy in all output.” The team was all bought into this number because it felt rigorous, the kind of number a serious product team commits to.

It’s also an output. And once again, even with AI products, we’re seeing the same mistake teams have been making with OKRs for fifteen years.

I’ve argued for a long time that a key result has to be a measure of human behavior. Not an output, not a task, not a feature you shipped but rather what a customer does differently because of the work you put out into the world. AI doesn’t change that rule. If anything it makes the rule harder to follow, because the thing your feature produces is probabilistic and you don’t fully control it. You can’t put “ship the summary” on a key result, and you can’t honestly commit to a single accuracy number either, because the model’s behavior will drift from user to user and from prompt to prompt. The goal moves where it always should have been in the first place, onto the human on the other side of the output. Here’s how that plays out across three types of key results, all of them measures of behavior.

The outcome KR: what users do after the AI answers

The first and most important question is deceptively simple. Once your feature hands the user an output, what do they do next?

That next action is your real signal. When the internally-facing bot provides valuable, actionable and accurate data, the user takes the data, exports it, shares it or integrates it into their deliverable and moves on with their day. When it doesn’t work, you see that in the behavior too. They regenerate the summary two or three times hunting for a better version (a re-query can certainly be a bad sign). They abandon it and open a blank doc to type their own notes. They export it somewhere else to quietly rewrite it before anyone sees it. Each reaction to the output is indicative of the value of the output to that specific user with that specific query.

The outcome KR names the good behavior and puts a number on it. For example, it can be, “increase the share of meetings where the user sends or shares the AI created data without rewriting it from 40% to 65%.” As always, we are looking for the formula of, Who does what by how much? In this case,

Who: the user. 

Does what: acts on the output without redoing it. 

By how much: From 40 to 65 percent of the time. 

You can pair that with a metric that keeps an eye on the bad behavior, like cutting the rate of double-regenerations from 22% to under 10%, but the headline is always a human doing something that proves the feature earned its place in their workflow.

The calibration KR: turning accuracy into a human behavior

I suspect this is where most teams want to sneak the accuracy number back in, and I get why. Quality matters. A feature that invents data or analysis nobody needs is worse than no feature at all. The calibration KR exists to hold that quality bar. The trick is to measure the bar the same way, through what the human does.

Ask the question directly. If the output is genuinely accurate and high quality, what does the user do differently? In practice they stop checking it. They stop correcting it. They accept the data output as written instead of opening the full transcript to verify every line, and they take any action items that come out of the data without editing them first. Given every user, regardless of prompt, is getting a different response, this is a way to make sure that, across the board, users are seeing answers they find to be accurate. 

Once again, we can set a key result here based on this behavior. Reduce the percentage of sessions where the user opens the transcript to fact-check the data from 60% to 30%. The quality threshold is still there measured as the moment the user decides your output is good enough to use as-is. That decision point is an observable behavior, and for the most part it’s the best proof of quality that there will be an ROI on this AI-powered feature.

The trust KR: measuring reliance, not sentiment

Trust is the one everybody wants to measure with a survey. Aside from being notoriously difficult to do well, oftentimes the data from surveys is not very reliable nor indicative of actual behavior. People tell you they trust the tool and then quietly check everything it does. Real trust isn’t a feeling ultimately but rather a change in what someone is willing to hand over. So watch what they hand over.

A user who trusts the internal data analysis tool starts using it for the meetings that matter, the client calls and the exec reviews, not just the internal standups. They turn on auto-send to the team instead of reviewing every summary in private first. Their corrections drop, not because they got lazy, but because they stopped expecting to find mistakes. Each of those is a behavior you can observe and measure.

The cleanest trust KR is watching whether your users override the output. Keep the rate at which users manually override, delete, or rewrite the AI’s proposed analysis and next steps under 8%. You can set an expansion behavior next to it, like growing the share of users who enable auto-share to their team from 15% to 40%. Both answer the same question: when people believe the system, what do they let it do that they used to do themselves? That behavior is the manifestation of trust in the product.

Putting the three together

Put the three and you get an OKR for an AI feature that never once measures the output of the AI. One key result for what users do with the output, one for the behavior that high quality produces, one for what trust lets them hand over. The feature’s accuracy, its model, its calibration, all of it still matters enormously, but as the driver of human success, not the measure of it.

If you’re staring at an AI feature this quarter and the only number you can think to commit to lives inside the model, you’re measuring the machine instead of the human. Try this instead: write down the single most valuable action a user should take after your feature responds, and ask how much more often you can get them to do it. Start there. It’s always been the person on the other side of the product that tells us whether we’ve built something valuable, not the product itself. 

Books

Jeff Gothelf’s books provide transformative insights, guiding readers to navigate the dynamic realms of user experience, agile methodologies, and personal career strategies.

Who Does What By How Much?

Lean UX

Sense and Respond

Lean vs. Agile vs. Design Thinking

Forever Employable

3 responses to “How to write OKRs for an AI product”

  1. Love the rigor — the “95% accuracy” trap is the right thing to kill.
    Where I get stuck is the absolute framing. “A key result has to be a measure of human behavior” fits user-facing features, but breaks when success means the absence of behavior.
    Let’s take a spam filter: the better it works, the less the user does — they never see the caught mail, never act, never know it ran. The only honest measure is the output. And the costliest error, a real email wrongly filtered, is invisible to any behavioral metric — you can’t observe someone not-acting on mail they never received. I suppose you could re-label a false positive as a cost in human terms, but that feels like a re-labeling exercise for the sake of re-labeling. Another example would be where the closest human behavior is way downstream – sure, related but too far removed to be an effective metric, so proximity matters.
    Makes me think the rule might be less “measure the human” and more “measure as close to the value as you can get” — usually the human, sometimes the output. Would you carve out an exception here, or is there a behavior I’m missing?

  2. […] has low learning value, even if customers end up liking it. This is the same discipline behind writing outcome-based OKRs. We’re asking what will change in our customers’ behavior, not what we’ll have […]

  3. […] i.e., the share of AI outputs a user modifies or discards before using. I wrote about this from the OKR angle earlier this year. The same logic applies to feedback: what the user does next is the […]

Leave a Reply

Your email address will not be published. Required fields are marked *