By Dr. Kevin Shepherdson, CEO and Founder of Straits Interactive
In Part I, we examined why “human-in-the-loop” is often used too casually in AI governance discussions.
The phrase sounds reassuring. It suggests that a human is still involved, still exercising judgment, and still able to prevent harm. But as we discussed, the mere presence of a human does not automatically make an AI system safe, fair, accountable, or well governed.
Human oversight only becomes meaningful when the human is qualified, informed, empowered, supported, and properly positioned in the workflow.
1. A human who does not understand the domain may miss serious errors.
2. A reviewer who lacks AI literacy may over-trust a polished but inaccurate AI output.
3. A manager who has no access to the system’s evidence, prompts, logs, or sources cannot properly verify what happened.
4. A person who is expected to review hundreds of AI outputs under time pressure may become fatigued and simply rubber-stamp the system’s recommendations.
In other words, human oversight can fail when it is treated as a slogan rather than a control.
Part II moves from the problem to the design question. If “human-in-the-loop” is not always enough, then what kind of oversight should organisations use? The answer depends on risk.
An AI tool used to brainstorm internal ideas does not require the same level of oversight as an AI system used in hiring, finance, healthcare, legal work, cybersecurity, or public services.
A chatbot that drafts a response is not the same as an agent that sends the response, updates a customer record, and triggers a workflow. A low-risk productivity tool does not require the same governance model as an autonomous agent with access to systems, data, and APIs.
This is why organisations need to move beyond one generic phrase and adopt a risk-based oversight model. For GRC, legal, privacy, audit, and compliance professionals, this distinction is especially important. They do not need to become AI engineers, but they do need enough AI literacy to ask the right oversight questions: what the AI is doing, what risk it creates, where the human is positioned, what evidence the reviewer sees, and whether the reviewer has the authority to intervene.
The key question is no longer simply: Is there a human in the loop? The better question is: What type of human oversight is appropriate for this AI system, given the level of risk, autonomy, reversibility, and potential harm?
This brings us to the next principle.
The Oversight Model Must Match the Risk
The mistake many organisations make is treating human oversight as one generic control.
It is not. The right oversight model depends on the risk level, context, reversibility, potential harm, and degree of AI autonomy.
1. A low-risk AI tool that helps summarise internal meeting notes may not require human approval for every draft. But it still needs clear rules on confidentiality, accuracy, and retention.
2. A recruitment screening tool requires much stronger oversight because it may affect employment opportunities and fairness.
3. A medical advice chatbot, legal research assistant, financial eligibility tool, or cybersecurity agent requires even stronger controls because errors may have serious consequences.
4. An autonomous agent that can update systems, send messages, execute code, or trigger actions requires a different level of governance altogether.
A simple way to think about it is this:
The more consequential the AI system, the more deliberate the oversight design must be.
The Human Must Be Qualified
This is one of the most overlooked issues. A human reviewer is not automatically qualified just because they are human.
If an AI system produces a medical recommendation, the reviewer must understand the medical context. If it produces legal analysis, the reviewer must understand the legal issue. If it screens candidates, the reviewer must understand fair employment practices. If it generates code, the reviewer must understand security, architecture, dependencies, and the application context.
A human who lacks the relevant domain expertise may not detect subtle errors. They may approve an answer that sounds plausible but is wrong. They may miss bias. They may fail to notice missing context. They may assume the AI is more reliable than it is.
This is why AI oversight must be role-based and competency-based.
The question is not: Did a human review it? The question should be: Was the right human qualified to review it?
The Human Must Understand AI Limitations
Domain expertise alone is not enough. A human reviewer must also have basic AI literacy.
They do not need to be a machine learning engineer. But they should understand that GenAI systems can hallucinate, produce biased outputs, overstate confidence, omit important information, misread context, mishandle prompts, and generate inconsistent responses.
They should also understand automation bias: the tendency for humans to over-trust machine outputs, especially when those outputs appear polished, confident, or data-driven. This matters because AI outputs often look more authoritative than they really are.
A human reviewer who does not understand AI limitations may become too trusting. Instead of challenging the output, they may simply endorse it. That is not oversight. That is delegation disguised as review.
The Human Must Have Time
Human oversight consumes time and attention. This is obvious, but often ignored.
If a person is expected to review hundreds of AI-generated outputs each day, the quality of oversight will decline. If a manager is asked to supervise ten autonomous agents while still doing their normal job, the burden may become unrealistic. If a compliance officer receives too many AI-generated alerts, alert fatigue will set in.
This is how cognitive offload can become cognitive overload. AI is supposed to reduce human burden. But poorly designed AI systems can increase it by forcing people to check, correct, interpret, and recover from machine-generated errors.
The organisation must therefore ask:
1. How many outputs must the human review?
2. How complex are the outputs?
3. How much time is needed per review?
4. How often are exceptions triggered?
5. Is the review workload sustainable?
6. What happens when the reviewer is unavailable?
7. Are reviewers trained and supported?
If the human does not have time to review properly, the control is not real.
The Human Must Have Access to Evidence
A human cannot meaningfully review an AI output if they only see the final answer. They need access to evidence. Depending on the use case, this may include:
1. the original prompt;
2. the system instruction;
3. the source documents;
4. the data used;
5. the model output;
6. the confidence level, where available;
7. the tool calls made by the agent;
8. the search results relied upon;
9. the assumptions made;
10. the logs and audit trail;
11. the reason for escalation.
This is especially important for agentic AI. If an agent searched the web, which sources did it use? If it updated a customer record, what data did it rely on? If it generated code, which files did it modify? If it summarised a contract, which clauses did it miss? If it escalated a case, what triggered the escalation?
Without provenance, the human reviewer is blind. A blind reviewer cannot provide meaningful oversight.
The Human Must Have Authority
Human oversight also requires authority. If the human cannot stop the AI, reverse the decision, override the output, demand more evidence, pause the workflow, or escalate the issue, then they are not truly in control.
They are merely observing.
This is a common problem in organisations where AI is embedded into fast-moving workflows. A human may technically be included in the process, but the system, business pressure, or workflow design makes it difficult to challenge the AI.
For example:
1. the AI recommendation is already pre-filled into the system;
2. the workflow pushes the human toward approval;
3. rejecting the AI output requires extra justification;
4. the human does not want to slow down the process;
5. management expects productivity gains;
6. the reviewer is junior while the system is treated as authoritative.
In such cases, the human may be socially, operationally, or technically discouraged from intervening. Meaningful oversight requires the right to say no.
The Human Must Be Positioned at the Right Point
Timing matters. A human review after harm has occurred is not the same as a human review before harm occurs.
For high-risk use cases, human intervention must happen before the AI output affects a person or triggers an irreversible action.
For example, reviewing a rejected job applicant after the rejection email has been sent is too late. Reviewing a harmful customer response after it has been published may be too late. Reviewing an unsafe code change after deployment may be too late.
In lower-risk contexts, post-action monitoring may be acceptable. But this must be a deliberate risk decision, not an accidental design choice.
Organisations should ask:
1. Should the human review happen before action?
2. Can the AI act first and be reviewed later?
3. Is the action reversible?
4. What harm could occur before review?
5. What threshold triggers escalation?
6. What must always require approval?
Human oversight must be placed where it can still change the outcome.
The Human Must Not Be Overwhelmed by Agents
Agentic AI makes this even more important. When organisations deploy multiple agents, the human may move from using AI to supervising AI.
That can be powerful if the agents are well designed. A capable operator can oversee multiple bounded, observable, and well-governed agents. But if the agents are poorly coordinated, opaque, unreliable, or constantly escalating, the human becomes a babysitter.
This creates AI fatigue. The human is no longer being augmented. They are assisting the assistant. This is why agent oversight requires:
1. clear agent roles;
2. limited permissions;
3. escalation rules;
4. logs and provenance;
5. exception dashboards;
6. cost visibility;
7. agent performance monitoring;
8. the ability to pause or disable agents.
One human can supervise multiple agents only if the supervisory model is designed. Otherwise, multi-agent automation becomes multi-agent chaos.
Human Oversight Must Be Tested
Another weakness is that organisations often declare human oversight in policy but do not test whether it actually works.
A control that is not tested is an assumption. Organisations should test human oversight by asking:
1. Can the reviewer detect incorrect AI outputs?
2. Can the reviewer identify bias?
3. Can the reviewer challenge unsupported recommendations?
4. Can the reviewer find the source evidence?
5. Can the reviewer override the system?
6. Can the reviewer handle the volume of reviews?
7. Can escalation happen within the required time?
8. Are exceptions logged and reviewed?
9. Are false approvals being tracked?
This turns human oversight from a statement into an operational control.
From Slogan to Control
“Human-in-the-loop” should not be used as a slogan. It should be translated into specific control design. A more mature organisation would not simply say: We have human-in-the-loop.
It would say: For this AI use case, the human oversight model is human-in-the-loop because the system influences employment outcomes. The reviewer is a trained HR professional. The AI output cannot be acted upon until reviewed. The reviewer can see the source data, scoring rationale, and audit trail. The reviewer can override the recommendation. All overrides and approvals are logged. The process is reviewed monthly for bias, accuracy, and workload.
That is a real control.
Compare that with: The system is safe because a human reviews it. That is not enough.
A Practical Checklist for Meaningful Human Oversight
Before relying on human-in-the-loop as an AI control, organisations should ask:
1. Qualification
Is the human reviewer qualified to assess the AI output in this domain?
2. AI Literacy
Does the reviewer understand the limitations of GenAI and the risks of automation bias?
3. Time and Capacity
Does the reviewer have enough time to perform meaningful review?
4. Evidence Access
Can the reviewer see the data, sources, prompts, logs, and reasoning trail needed to verify the output?
5. Authority
Can the reviewer stop, override, reverse, delay, or escalate the AI action?
6. Positioning
Is the human placed before the point of harm, or only after the decision has already affected someone?
7. Independence
Can the reviewer challenge the AI without pressure to approve quickly?
8. Escalation
Is there a clear path for uncertainty, disagreement, or high-risk outputs?
9. Accountability
Who is responsible for the final decision and for monitoring the effectiveness of the review process?
10. Testing
Has the human oversight process been tested under realistic conditions?
If these questions cannot be answered, the organisation should be careful about claiming that human-in-the-loop is an effective control.
Final Thoughts
Human oversight remains essential in AI governance. But it must be designed, not assumed.
The phrase “human-in-the-loop” should not be used as a comfort blanket. It should not be a memorised answer. It should not be a box to tick when someone asks about AI risk.
A human in the loop is not automatically a safeguard. A qualified, informed, empowered, evidence-supported, and properly positioned human can be. That is the difference between governance theatre and real AI governance.
In the age of GenAI and agentic AI, the question is no longer whether there is a human somewhere in the process.
The real question is: Can the human meaningfully control the loop before harm occurs?
This is Part 2 of a two-part story. For Part 1, please click here.