How to write an outcome-based OKR when your user is an AI agent

How to set an OKR for an AI agent (hint, keep it human)

In the webinar Josh Seiden and I ran on September 10 about what product management looks like in 2027, the top-voted question came from a product manager who had clearly been waiting for their opportunity to ask us this. In Who Does What By How Much? we are unambiguous that a key result has to measure a change in human behavior. The key word here is human. People have to change what they do to indicate they have gotten any value from our product. This product manager wanted to know how that holds up when the thing consuming your product is increasingly an autonomous agent rather than a person. Another attendee typed into the chat that this was already happening on their team. I gave them a short answer in the chat (hey, attention span is….what it is).

So here is the longer answer, along with how I would rewrite I almost every agent-facing OKR I have been shown this year. In this week’s newsletter and a recent blog post on writing a definition of done for an AI feature, I called out that when measurement gets hard, teams quietly remove the customer from the goal and put the machine in their place, and because the machine is so much easier to measure, nobody objects.

An agent has behavior. It has no stake.

The reason agent metrics feel like outcomes is that agents genuinely do behave. They call your endpoints, they retry, they give up, they escalate to a human, and they pick a competitor’s integration over yours. All of it observable with a precision that human behavior rarely offers you. That precision is seductive and it is also where the trouble starts, because an outcome, the way Josh and I define it, is a measurable change in human behavior that creates value. An agent isn’t human, has no needs, no budget and no Tuesday morning that got better or worse. It cannot be a beneficiary.

Somebody, though, always is. There is a human who deployed the agent and is accountable for what it does on their behalf, and there is usually a second human further down the line who receives the agent’s work and has to accept it, repair it or throw it out. Those two people have behavior worth changing, and they are where your key results live. The agent sits in the middle as a courier, and how well the courier does its job is a system health metric you should watch on a dashboard, where health metrics belong.

Write the key result one level up

Take a look at this example and rewrite to get a sense for what I’m suggesting.

Objective: Become the default booking layer for autonomous travel agents.
KR1: Grow agent-initiated API calls from 1.2M to 5M per month.
KR2: Cut median response time from 900ms to 200ms.
KR3: Raise the successful request rate from 92% to 99%.

Every one of those is measurable, every one will move if the team works hard, and not one of them tells you whether any traveler had a better trip. This is the glorified task list with a confidence score attached that I complain about constantly, and it is the same mess that AI makes faster instead of fixing. Move each key result one level up, to the human on either end of the agent, and it turns into this.

Objective: Travelers get trips booked through their assistant that they actually take.
KR1: Increase the share of agent-booked itineraries the traveler accepts without editing from 48% to 70%.
KR2: Reduce the share of agent-booked trips that someone on our support team has to repair from 22% to 8%.
KR3: Increase the share of travelers who leave their assistant on auto-book after the first time it gets something wrong from 30% to 55%.

The second set is harder to instrument, which is the point. Each one names a person doing something differently. The first measures the traveler trusting the output enough to leave it alone, which is the same acceptance signal I described in how to write OKRs for an AI product. The second measures your own ops people doing less cleanup, which is a cost your agent integration either creates or removes. The third is the one I would fight hardest for, because reliance after a failure is the real test of whether people believe in your product, and it is not visible nor tracked in every KR in the first set.

The question to ask before you write a single key result

Take whatever OKR your team has written for its agent-facing work, go through it line by line, and ask of each key result whose behavior it predicts. If you can name the person, the traveler who stopped editing the itinerary or the support rep who stopped repairing bookings, you have an outcome. If the only available answer is the agent, you have a health metric, which is a fine thing to have and measure. It’s a poor thing to tell your team is the definition of success this quarter.

None of this is exotic, and it does not need a new framework, an acronym or a maturity model. It needs you to keep asking who does what by how much even when the doing is delegated to software acting on somebody’s behalf. Run the rework above on your current OKRs this week and let me know how many of them survive it.

Books

Jeff Gothelf’s books provide transformative insights, guiding readers to navigate the dynamic realms of user experience, agile methodologies, and personal career strategies.

Who Does What By How Much?

Lean UX

Sense and Respond

Lean vs. Agile vs. Design Thinking

Forever Employable

Leave a Reply

Your email address will not be published. Required fields are marked *