Launchfiles logoLaunchfiles

BlogHuman-in-the-Loop vs. Human-as-QA: Why Most HITL Workflows Are Broken

Owned essay

Human-in-the-Loop vs. Human-as-QA: Why Most HITL Workflows Are Broken

Don't answer a stream of mid-task pings. Approve one finished package with the evidence attached.

Human-in-the-Loop vs. Human-as-QA: Why Most HITL Workflows Are Broken

A founder opens the chat window where their agent has been running a distribution spend analysis. Dozens of messages wait.

Many are questions.

The CSV has a column called campaign_name and one called campaign. Which is canonical? Should the test campaigns be excluded from the retargeting set? Each question is reasonable in isolation. Each one also pulls the founder out of the pricing call they were on, forces them to reacquire context lost an hour ago, and extracts a decision made on incomplete information. By the time the stream slows, the founder is no longer directing work. They are doing the agent's job for it, one ping at a time.

This is the approval ping trap. It feels like control. It delivers the opposite. The founder becomes a full-time QA reviewer, verifying each micro-step instead of deciding on outcomes. The agent never finishes anything, because it is never allowed to finish anything. And the work that does get done gets redone, because the approvals that shaped it were given without the full picture.

The decision every founder faces is whether to trust agent work without work orders, gates, and receipts. Most teams default to one of two bad options. They either babysit the agent through constant check-ins, or they let it run loose and hope the output is sound. The first burns attention. The second burns rework. There is a third path, and it starts with a different unit of work.

The Approval Ping Trap

The synchronous chat model has a structural flaw that no amount of prompt engineering fixes. It asks for approval at the wrong granularity. The agent treats every fork in the road as a decision point, and every decision point as a reason to interrupt. The founder, wanting to stay in control, answers. Each answer feels like governance. In aggregate, it is abdication.

Consider what happens to context across that long message thread. By mid-morning, they have answered several more questions about data sources and attribution rules. They approve something that contradicts their earlier intent. The agent, dutiful and literal, executes. The resulting analysis is internally consistent and substantively wrong.

The cost is not just the founder's time. It is the quality of the decisions made under interruption. A founder who is constantly being pulled into low-context approvals is a founder who is making premature calls. They are approving actions without seeing the shape of the whole package. They are signing off on steps without knowing how those steps fit the mission. The agent gets its checkpoint. The founder gets a headache and a rework cycle.

Some founders push back that more check-ins mean more control. The flaw in that reasoning is the assumption that each ping carries enough context for a good decision. It does not. A mid-task question about a CSV column is not a strategic checkpoint. It is a request for the founder to do the agent's data-cleaning work. Answering it feels like oversight. It is actually just labor.

Why Human-in-the-Loop Degenerates into Human-as-QA

The mechanism behind the trap is simple. When the only way to gate an agent is to interrupt it, the founder becomes the approval bottleneck for every action. The agent cannot proceed without a green light, so it asks for one constantly. The founder, to keep things moving, becomes a reviewer of process rather than a decider of outcomes.

This is the degeneration from human-in-the-loop to human-as-QA. The product landing on human-in-the-loop AI agents is the same job: keep the founder on the decision, not on every ping. The loop was supposed to keep the human informed. Instead, it makes the human responsible for verifying each step the agent takes. The founder is no longer asking whether the outcome was achieved. They are asking whether the agent formatted the CSV correctly. That is a catastrophic misallocation of founder attention.

The alternative is not fewer checkpoints. It is better-structured ones. Work orders define the scope of what the agent is allowed to touch. Approval boundaries mark the specific points where human judgment is required. Receipts make every claim the agent makes auditable after the fact, the same record in why agent work needs receipts. The founder does not approve the process. They approve the package.

This shifts the founder's role from QA to approver. A QA reviewer checks the work. An approver checks the outcome against the mission and decides whether to proceed. The difference is the difference between inspecting every brick and approving the building. One is exhausting. The other is governance.

The Asynchronous Bounded Package Model

The Asynchronous Bounded Package model inverts the chat-ping dynamic. The agent receives a work order that defines the scope, the data sources, and the expected deliverable. It then executes to a verified checkpoint. When it hits that checkpoint, it does not send a stream of questions. It assembles a single decision card.

That decision card is the core artifact. It contains the outcome, the key findings, and the visible evidence for each claim. It shows what changed, what was analyzed, and what the agent recommends. The founder reviews one high-context summary instead of dozens of low-context pings. They see the whole shape of the work before they make a single call.

Receipts sit underneath the decision card. Every claim in the package carries a receipt that can be checked. If the agent says the distribution spend on one channel underperformed, the receipt shows the source data and the calculation. The founder does not have to trust the agent's summary. They can verify it. This is what makes the output founder-trustworthy. It is not trust based on the agent's confidence. It is trust based on inspectable evidence.

The agent executes within the bounds of the work order. If it needs to go outside those bounds, it does not just do it and ask forgiveness later. It flags the boundary and waits. But that flag is rare, because the work order was designed to cover the scope. The founder sets the boundaries once, up front, instead of policing them in real time.

Vibe Coding vs. Receipts: The Trust Gap

Vibe coding is the name for the other failure mode. The agent is given a mission and set loose. It produces output with no work order, no gates, and no receipts. The output might be brilliant. It might be a hallucination wrapped in confident prose. The founder cannot tell the difference without auditing everything, which defeats the purpose of having an agent.

Vibe coding skips proof and burns rework. The agent publishes its analysis. The founder reads it, spots a suspicious claim, and has to trace it back to the source. There is no receipt. The founder has to redo the work to verify the work. That is not delegation. That is doing the job twice.

The trust gap between vibe coding and receipts is the gap between hope and verification. A founder who runs vibe coding is hoping the agent got it right. A founder who runs receipt-based work knows they can check. The first is a leap of faith. The second is a decision.

Some founders argue that vibe coding is faster because there are no checkpoints to slow the agent down. The speed is illusory. The agent finishes its pass quickly, but the founder then spends hours auditing the output because nothing is verifiable. The rework and the auditing consume more time than the checkpoints ever would. The receipt is not overhead. It is the thing that makes the agent's speed usable.

What receipts can and cannot prove

Receipts are evidence, not guarantees. This distinction matters. A receipt proves that work was done and that a claim is checkable. It shows the source, the calculation, and the result. It does not prove that the outcome was the right one for the business. It does not promise that the distribution spend will hit a target. It does not guarantee a return.

This is the boundary of what receipts can do. They make the agent's work auditable. They let a founder verify that the agent did what it said it did. They turn trust from a feeling into a process. But the founder still has to make the judgment call about what the evidence means for the mission. A receipt can show that a channel spent more and converted less. It cannot tell the founder whether to cut the channel or change the creative.

If receipts do not guarantee outcomes, why trust them at all? Because auditability is the basis for trust. You cannot trust a black box. You can trust a system where every claim has a receipt you can pull. The receipt does not remove the need for founder judgment. It gives the founder the raw material to exercise that judgment well. The alternative is trusting a vibe, which is not a strategy.

From QA to approver: a founder's workflow

Picture the concrete difference. A founder needs a distribution spend analysis to decide where to put the next month's budget. In the chat-ping model, they spend the morning answering questions about date ranges and data columns. In the bounded package model, they write a work order. The order specifies the analysis window, the channels to include, and the decision the analysis is meant to inform.

The agent executes. It pulls the spend data, runs the analysis, and assembles a decision card. The card shows the spend by channel, the conversion signal for each, and a recommendation for the next allocation. Each number on the card has a receipt. They do not have to take the agent's word for anything.

The founder reviews the single card. They see that one channel is consuming budget with weak conversion signal. They approve the recommendation to shift spend. The founder made the strategic call. The agent did the legwork. The receipts made the legwork verifiable.

This sounds like more overhead to some founders. The work order takes time to write. The checkpoint takes time to reach. But the overhead is front-loaded and finite. It replaces the open-ended overhead of constant interruption. The founder spends twenty minutes defining the work order and saves hours of mid-task pings. The trade is structural, and it favors the approver.

Building Trust with Receipts, Not Vibes

The choice is not between trusting agents and distrusting them. It is between two kinds of trust. One is blind, based on the hope that the agent got it right. The other is structured, based on receipts that make every claim checkable. Vibe coding offers the first. Work orders, approval boundaries, and receipts offer the second.

The goal is to move the founder from the role of QA reviewer to the role of approver. A QA reviewer is stuck in the process, verifying each step. An approver stands above the process, deciding on outcomes. The bounded package model makes that shift possible. It gives the founder a single decision card, backed by receipts, instead of a stream of low-context questions.

Receipts and approval boundaries make agent output founder-trustworthy. They do not guarantee outcomes. They do not remove the need for judgment. They make the work inspectable, which is the only durable basis for trust. A founder who can verify the work can let the agent run. A founder who cannot is stuck babysitting or hoping.

If you are tired of being the approval bottleneck, the next step is to look at how you structure the work itself. The question is not whether to check in more or less. It is whether you are reviewing a process or approving an outcome. Reading every line of the agent's diff is the other face of the same tax — agent review fatigue is that bottleneck on the code path.