I take every August off. It’s a luxury and I’m lucky to be able to do it. This is the first year I’ve come back from my summer vacation and had to check in on my various Claude tasks, agents, chrons and other activities it was supposed to do while I was busy NOT doing them. I was excited by the promise of autonomous AI agents doing work “while I was asleep” or, in this case, literally in China.
What I found when I unlocked my laptop for the first time in nearly 4 weeks (yep, though I had my iPad with me for, you know, “emergencies”) was nothing short of a sh!tshow. Claude had tried to run the various tasks on schedule but every single one had broken down. Some had authentication issues that couldn’t be resolved. Others didn’t have the previous week’s data to build on and so had no idea how to continue. And, my favorite, it’s autonomous deviations from the plan in order to somehow placate its need for task completion. It literally began to build a parallel system that duplicated the one I had in place when I left because it couldn’t get answers for the original system. I spent an entire day working to replace and reconcile, not to mention reauthenticate everything to get it back to some kind of working shape.
The worst part is that the work didn’t get done. While no one except me was looking for the output of these various tasks and agents, the fact is I expected to come back and hit the ground running. To say that wasn’t the case, is, well, an understatement.
If this is the future, we’re not quite there yet. We keep hearing talk of putting a human in the loop. I couldn’t agree more given my experience after a month away. The trick, of course, is to make sure that the human is in the right part of the loop and that they have what they need to critically judge everything that comes through, rather than just clicking “Approve” every time Claude asks for it.
The reason most human-in-the-loop checkpoints turn into mindless rubber stamps
Since we can’t trust the robots to fully run the world just yet (nor should we ever, right?) we need to insert humans into the processes we’re building in ways that not only take advantage of our unique human abilities but allow for that human intervention to have an impact. Unfortunately, today most human-in-the-loop situations fail for at least three reasons:
- The human reviewing work gets the output to review, not the rationale behind it. By the time the human you built into the process gets the material they need to check over, many of the interesting and impactful decisions have already been made. They can only react to the output in its current state, not the choices that were made to get there.
- The checkpoint comes too late in the process. More often than not, the human reviewer is brought in at the end of the production process, when most of the work has been done (see #1 above) and the tokens have been spent. At this point, when faced with a choice between “approve” the work or “start over” where the cost can be anything from an hour to a week or longer, the human reviewer is really left with only one choice – approve.
- Approval is the fast, rewarded path. AI only made our feature factories faster and more efficient. The incentive for most teams is still velocity and throughput. Time for reviewing and course correction was already scarce. Now, in many cases, it almost never happens. If the reviewer doesn’t have the time or isn’t incentivized to do a thorough review, the process is dead on arrival. The checkpoint becomes a formality, a rubber stamp.
- Too many checkpoints on too much output. With the added velocity and throughput, the human in the loop now has to go through many review cycles per day. At some point they learn how to skim and, inevitably, the important review is missed as the reviewer is overwhelmed with the sheer quantity of work they have all of a sudden.
When you start to see these patterns over and over, it makes you stop and wonder whether the current implementation of human-in-the-loop solves the original problem it was intended to solve. In other words, did we add the human to the loop to feel safer or did we add them to make important decisions? I suspect it was the latter but the reality, as it currently stands, favors a “feeling” of safety rather than the ability to deliver that level of confidence.
How to test for a real “human-in-the-loop” checkpoint
If you already have these checkpoints in place or are about to place them, here are three questions to help you ensure they’re adding value rather than just ticking a box:
- Can your reviewer name what they would change when presented with something that doesn’t look right and why they would make those changes? If not, and the only thing they can report is “doesn’t look right” rather than “here’s what I would do differently here…” then your reviewer isn’t much more than a proofreader.
- Does your reviewer have what they need to figure out why something doesn’t look right and what changes they would make? When presenting any output for review, your AI agents should also share what they didn’t do, what they discarded and the rationale behind it. This gives your human reviewer a sense of why the output looks the way it does and what may need to come back into the picture.
- What happens downstream from the checkpoint? If the work proceeds either way you don’t have a human-in-the-loop check point. You have a human whose job is to notify the next human that the work is on the way.
The bottom line is this. If the reviewer can’t say what they would change, why they would change it and what impact it may have downstream from them, this checkpoint is much more noise than it is signal.
Where to place the human checkpoint in an AI workflow
Here is the scenario that my Claude agents are designed to do:
- The agent scans all the content I’ve previously posted, what’s being discussed in public forums and social media and trending topics that I typically care about which it then consolidates into specific themes
- The agent then drafts content experiment ideas that would take advantage of unowned gaps in the content matrix it has developed
- I pick one idea to pursue on a weekly basis
- The agent then drafts an outline and offers suggestions for SEO/AEO optimization, internal linking, etc
Where would you put the human in the loop in this workflow? Most would say that if you had one checkpoint you’d put it right at step 4. Why? Because Step 4 is where the production of the artifact begins. It feels responsible to review it here.
The reality though is that putting the human checkpoint between steps 2 and 3 makes more sense. It’s at this point that only one idea moves forward while others get discarded. Understanding why certain ideas are promoted versus others is a critical moment for human intervention. It’s also the “last cheap moment” in the workflow. Everything that comes downstream from this is heads-down creative work and effort to get the idea out the door. This is the last opportunity for a human to change the outcome, rather than the wording of what goes out.
This extrapolates to most other scenarios. Put the human-powered review checkpoint at the last cheap moment in your workflow, not the last actual moment. Ask yourself and your team where reversing course gets expensive, and put the human in the loop one step before it.
Building a more effective human-in-the-loop checkpoint this week
Here’s how to operationalize this idea this week with your team. Take one AI workflow your team is using, perhaps the one it runs most often right now. Write down its execution steps. Then, as a team, identify the step at which the process gets expensive if you have to reverse course. Put your human checkpoint one step before it.
Next, run the three questions above on all the other approval steps you already have. Delete the ones that fail. Every checkpoint that fails the test gets you back to a place where that attention is more valuable and useful.
The net result here is a more efficient human-in-the-loop process that actually prevents your team and your agents from going off the rails too quickly. And, perhaps, if you’re lucky it gets them to a point where eventually they can do most of the work on their own. However, it’s always worth identifying where the course correction gets expensive and planting your human there.
If you had to look at your automated AI processes today and identify that last cheap moment, where would it be? Share your thoughts in the comments.






Leave a Reply