The Judgment Line
Why the question is not whether AI can do the work, but whether it may do it alone
The problem
This is written for people who run mid-size companies. Fifty to two thousand employees, a real P&L, no chief AI officer and no data science team. When something goes wrong, you are the one who answers for it.
You are probably already convinced that AI matters. Most owners and operating executives I meet are. You have seen the demos, read the books, watched a competitor announce something. What you most likely do not have is a single AI system running in your operations that your team trusts and you would defend in front of a customer.
That gap is the problem. It is not a technology problem, because the technology works. It is that nobody has told you how to decide.
Three questions sit between conviction and deployment.
Which work do we automate first? You have dozens of workflows: quotations, invoices, inspection reports, customer enquiries, supplier documents. Some are good candidates. Some will consume six months and produce nothing. A vendor demo does not tell you which is which, because the vendor is showing you what their product does, not what your business needs.
Where do people stay involved? Everyone repeats the phrase “human in the loop” and nobody defines it. Which loop, doing what? Reading every output before it goes out, or checking one in twenty afterwards? The answer differs by workflow, and deciding it requires a rule you do not currently have.
How do I govern this well enough to sign off on it? If an automated quotation goes out with the wrong price, you carry that, not the vendor. So before you deploy anything you need to know what the system does on its own, what a person checks, and what evidence exists afterwards when a customer or an auditor asks what happened.
Without answers to these, companies do one of two things. Most do nothing, because the risk is unquantified and inaction feels safe. A smaller number automate something that should never have been automated, get burned in front of a customer, and quietly retreat. The first group concludes AI is not ready. The second concludes it does not work. Both conclusions are wrong. What was missing in each case was a way to decide.
This essay describes the way I use to decide. I call it the Judgment Line, and the name refers to a boundary you draw inside a workflow: on one side, work a machine may complete on its own; on the other, work a person must complete. Drawing that boundary deliberately, before deployment, is the whole method.
Why the usual approaches fail
Four approaches dominate. Most readers will have tried at least one, and each fails in a way worth naming.
The vendor demo. A product is shown, it performs well, and a workflow gets automated because that product handles it. Nobody asks whether that workflow should have been automated, because the vendor has no reason to raise the question and the buyer has no framework for raising it.
Licenses for everyone. Chatbot access is distributed across the company on the theory that people will find their own uses. Some do. But nothing becomes a system, no output is supervised, and the company acquires a large amount of unmonitored AI use it can neither see nor defend.
The consulting roadmap. Eighteen months, a transformation office, a rebuilt operating model. This is written for enterprises with a transformation budget. A two-hundred-person firm cannot execute it and should not try.
The feasibility study. The most rigorous of the four and the most instructive failure. It assesses each candidate workflow on a single question: can AI do this work? That question is real. But it is the half that improves every year, and it is almost never the half that causes damage.
Consider what actually happens when an AI deployment embarrasses a company. Almost always, the technology performed as designed. It drafted the quotation competently and sent it. The failure was that it should never have been sending anything to a customer without a person reading it first. Capability was assessed and approved. Whether the machine should have been acting alone was never assessed at all, because the study had no place to record that answer.
That is the gap the Judgment Line fills.
The core idea
The method rests on separating two questions that are usually collapsed into one.
Can the machine do this work? This is a capability question. Technology answers it, and the answer gets better every year.
May the machine do this work alone? This is a permission question. Technology has nothing to say about it. It is answered by asking who gets called when the output is wrong.
Keeping these apart is the entire idea. Everything else in this essay follows from it.
Three things come out of the separation. It explains the failures above: capability said yes, permission was never asked. It produces three possible outcomes for a task rather than two, and the middle one, where the machine prepares work and a person completes it, turns out to be where most real work in a mid-size firm belongs. An assessment that asks only about capability cannot see that middle category, because it sorts every task into automatable or not, and the largest opportunity falls between the two. And because permission changes slowly while capability changes yearly, the separation is the only reason an assessment you make today survives the next model release.
The whole method on one page
Everything that follows is in this diagram. Read it top to bottom: a workflow enters at the gate, its tasks pass through six questions, and each task exits into one of three destinations. The heavy vertical mark near the bottom is the Judgment Line itself.
The rest of this essay walks through that diagram, one band at a time: the five-step method it sits inside, then the gate, the six questions, the three destinations, and finally the line.
The method
The Judgment Line is drawn inside a five-step method. Each step ends with a document a leadership team can hold, which is what separates a method from a framework.
In the diagram, the five steps run along the bar at the very bottom. Everything above that bar is the first step.
Map. Take one workflow, put four people who actually do that work in a room, and spend ninety minutes on it. Break the workflow into individual tasks, ask six questions of each task, and sort the tasks into three groups. The output is a line map: what machines finish, what machines prepare, what people own, for that one workflow.
Rank. Map several workflows, then order them by which to automate first. The ranking uses volume, ease of verification, tolerance for error, how much of the needed information is already written down, and how heavily the workflow is audited. The output is a backlog in which the first project was chosen rather than volunteered.
Design. For each task the machine will handle, decide how closely it is supervised: every output read before it goes out, or exceptions routed to a named person, or a sample audited periodically. The output is a document stating who checks what, how often, and what happens when they find a problem.
Govern. Write down what makes the automation defensible: what the system logs, how samples are drawn, how data is handled, and a one-page policy the whole company can read. The output is one page per automation, which is what lets you answer a board member or a regulator without convening a committee.
Run. Deploy within ninety days. If waiting for a system integration would delay that, start with a person doing the handoff manually, because a working governed process proves value faster than a perfect one proves anything. Then revisit the line every quarter, because capability changes and the map goes stale.
The rest of this essay covers Map, which is the step companies most often skip. They begin at Run with a product they liked, discover in the fourth month that nobody ever agreed where the boundary was, and end up in one of the two failure modes described above.
The gate: how often does this work happen?
This is the band at the top of the diagram, marked GATE, and it disqualifies more candidates than everything below it.
Ask how often the work happens. Not whether a machine could do it, but whether anyone should care if a machine did. A task performed twice a year can pass every test in this essay and still be a poor investment, because building and supervising the automation costs more than the hours it saves, and it never runs often enough for anyone to learn whether it is working properly.
Automation earns its supervision cost through repetition. So the workflows that deserve ninety minutes on a whiteboard are the ones your people touch constantly: the quotations, the invoices, the inspection reports, the enquiries that arrive every day.
The six questions
These are the three boxes in the middle of the diagram. Take a workflow that passes the gate, break it into individual tasks, and ask six questions of each task, in the three groups shown.
Group one: can the machine do it?
The left-hand box in the diagram.
1. Repeatability. Has this task been done a hundred times in essentially the same way? Machines handle the hundredth instance of a familiar pattern well and handle genuine one-offs badly.
2. Verifiability. Can someone check the output faster than they could have produced it? A drafted quotation can be checked in two minutes. If checking takes as long as doing, automation saves nothing.
3. Boundedness. Is everything the task requires written down somewhere? If the real procedure lives in an experienced employee’s head, the task is not yet automatable. The knowledge has to be captured first.
Three yes answers mean a machine is able to do the work. These three questions get easier to answer yes to with every model release, which is exactly why they are the wrong foundation for a governance structure.
Group two: may it act alone?
The middle box, drawn with a heavier border because these are the questions that place the line.
4. Accountability. When this goes wrong, does someone outside the company expect a named person to answer for it? A machine can produce a tax filing indistinguishable from the accountant’s. But when the tax office has a question, it calls the accountant who signed it. Machines can prepare work that someone is accountable for. They cannot take on the accountability itself, and no vendor contract transfers it to them.
5. Relationship. Does the value of this task depend on a person doing it? An apology delivered by an automated system is not a cheaper apology. It is an insult.
These two questions place the line. They change only when law, professional obligation, or client expectation changes, which is to say rarely. No model release alters who signs an audit opinion, or whether your largest customer wants a machine handling the relationship.
Note that the two groups are independent. A task can pass all three capability questions and still be forbidden to run alone. A task can carry no accountability at all and still resist automation because the knowledge it needs was never written down. The most expensive mistakes come from treating a yes in one group as a yes in the other.
Group three: how bad is wrong?
The right-hand box, set apart because it does a different job.
6. Consequence. If a bad output reaches the outside world and nobody catches it, what breaks? Answer in one sentence, and have the people who own the work answer it.
This question does not place the line. It sets how closely the machine’s work is checked once the line is drawn. Two tasks can sit on the same side of the line and deserve completely different supervision: a wrong internal summary and a wrong customer quotation are not the same event.
This one sentence is the whole of risk assessment in the method, deliberately. Question four already captures regulatory and financial exposure by asking who answers. Question five already captures reputational exposure by asking who notices. What those two miss is magnitude, and one honest sentence supplies it. A full risk register scoring every task on every dimension produces a document nobody reads and a workshop nobody finishes.
Three destinations
These are the three boxes below the questions in the diagram, arranged left to right in order of decreasing machine autonomy, as the label on that row notes.
The machine finishes the work. Capability says yes and permission raises no objection. Invoice matching, data extraction, routine categorization. The machine completes these tasks, and people check a sample of the output on a schedule rather than reading every item. Question six sets how large that sample is.
The machine prepares, a person finishes. The machine is fully capable, but a person must perform the final act. The tax filing is drafted; the accountant reviews and signs. The difficult customer letter is drafted; the manager edits and sends. The performance review is assembled from the year’s notes; the supervisor decides the rating and holds the conversation.
A person does the work; the machine only briefs them. Permission says never alone. The apology call, the safety judgment on a factory floor, the price commitment that binds the firm. Here the machine’s role is preparation: the briefing before the call, the history before the decision.
The middle destination deserves attention, because it is where most real work in a mid-size firm belongs, and it has a plain description everyone understands immediately: the machine produces the first draft, and a person is the second.
What the line actually marks
In the diagram this is the heavy vertical mark sitting on the dashed track beneath the three destinations. Notice where it falls: between the first destination and the other two. It marks one thing, and that is who completes the work.
On one side, the machine completes the work itself. On the other side, the machine takes the work as far as the boundary and a person completes it. That is all the line means, and it explains why being on the human side of the line does not mean AI is prohibited there. AI works right up to the line in all three destinations. The line does not stop automation. It specifies what supervision a yes requires.
The dashed track and the small arrows on either side of the mark are there for a reason. The line’s position is not fixed, and this matters more than anything else in the method. It is drawn per task, and it falls in different places in different industries, because regulation and customer expectation differ. The same drafting task sits on the machine side in a marketing agency and on the human side in an audit firm, and both are correct. This is why no vendor can draw your line for you, and why one workshop on a single real workflow with the people who run it is worth more than a generic assessment of the whole company.
The line does move over time, but in one direction and for one reason. Better technology answers the capability questions more often, which lets more work sit on the machine side. The permission questions hold the line in place, because they are not affected by better technology at all. So anchor your governance on the permission side, which is stable, and let each year’s models extend the machine’s reach within that same structure. Companies that do this stay calm through every product launch. Companies that anchor on capability redraw everything annually and mistake activity for progress.
Where to start
Start where mistakes are cheap, not where the pain is worst. The instinct is to point AI at the workflow that hurts most. That is almost always wrong, because the most painful workflow is usually painful precisely because it is tangled with judgment, exceptions, and history. The right first candidates are high in volume, forgiving of error, and easy to check: the ones where question six has a boring answer. Early success there buys the organizational trust that harder workflows will require.
Before any of it, capture what your veterans know. Sit with your most experienced people and record how they actually solve things: how this machine fault was diagnosed, why that supplier is handled differently, what the real sequence is when the standard process does not fit. Nothing executes, nothing reaches a customer, nobody’s job is threatened, and the people whose knowledge is captured experience it as recognition rather than replacement. It is also the highest-return preparation you can do, because it produces exactly the written-down knowledge that question three demands and that every later automation depends on. Leadership teams ask me for this once they see it. Almost no company does it first.
What the line is for
The purpose of drawing the line is not to remove people. It is to move their attention. When machines complete the hundredth repetitive instance and prepare the first drafts, human judgment relocates to where it was always more valuable: the exceptions, the relationships, the decisions someone must answer for, the second draft that makes the first one safe.
For the people inside the workflow, that is a promotion, and it should be described as one, out loud, in terms of tasks rather than jobs. The companies that handle this well do not have smaller teams a year later. They have teams doing recognizably more valuable work, supervising a layer of machine effort beneath them.
One more thing, about cost.
Every company eventually learns where its line sits. Some learn it in a conference room: four people, a whiteboard, ninety minutes of argument about one workflow, before anything is deployed. Others learn it the day an automated system sends a wrong quotation to their biggest customer, and the managing director spends a week on the phone repairing a relationship that took nine years to build.
Both companies end up with the same knowledge: which parts of that workflow a machine may finish, and which parts a person must own. Only the price is different.
Draw yours on purpose.
---
The Judgment Line is the method behind the assessment work I do through GradTensor, where we help mid-size businesses put AI into their operations under their own control. I write about it, and about reliable AI more broadly, at Trust and Reason.
The Judgment Line™ is a trademark of GradTensor (Sudhanva Labs LLP). © 2026 Prabhu Eshwarla.



