AI explanations can weaken independent judgment

228 experienced evaluators assessed nearly 50 submissions to an MIT innovation challenge. The experiment, conducted by researchers affiliated with Harvard Business School, MIT, and the University of Washington, found a clear risk: an AI explanation can make a recommendation more persuasive without making the resulting human decision more accurate.

The researchers compared three conditions. Evaluators reviewed proposals without AI assistance, received an LLM pass-or-reject recommendation without an explanation, or received the recommendation with a written rationale. Four human experts provided the benchmark for classifying decisions as correct, false positives, or false negatives.

The unexpected result came from the explanations. Black-box AI recommendations improved alignment with the expert benchmark. Adding a narrative explanation did not. The effect was particularly important when the LLM recommended rejecting a proposal. A persuasive explanation increased agreement with these rejection decisions and substantially increased false negatives. Evaluators therefore rejected more proposals that the expert panel considered worth advancing.

The mechanism matters for executives. LLMs produce fluent and coherent explanations that resemble expert reasoning. People can interpret those qualities as evidence that the reasoning itself is sound. The researchers describe this as an “illusion of explanatory depth.” A reviewer can feel that they understand why a decision is correct while having limited visibility into how reliable the reasoning actually is.

This changes how enterprises should think about explainable AI. More explanation can influence human behavior as much as it improves transparency. In uncertain decisions, a polished rationale may reduce the motivation to conduct an independent check. The key design objective is therefore to preserve human verification.

For high-stakes workflows, companies should test explanations as part of the complete decision system. Measure whether reviewers catch AI errors, whether false positives or false negatives increase, and how often people successfully override incorrect recommendations. An explanation feature has value when it improves the quality of the combined human-AI decision.

People show high deference to AI recommendations

Evaluators accepted LLM recommendations 67% of the time overall. Agreement with both black-box and narrative AI decisions was roughly 75%. Agreement in the human-only condition was 54%. These results show how strongly an AI recommendation can shape a subsequent human decision.

This matters because enterprise AI increasingly operates before a human makes the final call. LLMs can screen proposals, summarize evidence, rank options, and recommend an outcome. Keeping a human formally responsible for the final decision does not guarantee independent oversight. If the reviewer routinely adopts the model’s recommendation, the AI has substantial practical control over the outcome.

The real management problem is therefore effective human oversight. Organizations need to know when employees independently assess an AI output, when they simply approve it, and when overriding it improves the decision. The study calls the last behavior a “productive override”: a person independently checks a persuasive AI recommendation and correctly intervenes.

This distinction should influence governance. A high human approval rate is a weak measure of successful human-AI collaboration. Agreement can reflect an accurate model, excessive trust in the model, or both. Leaders need outcome measures tied to decision quality. These include false-positive and false-negative rates, successful overrides, and performance against a credible benchmark.

The appropriate balance also depends on the cost of error. Early-stage innovation screening is especially sensitive to false negatives because rejecting a promising proposal can remove future options. Compliance, fraud, or quality-control processes may place greater value on conservative decisions. Companies should set review rules around these asymmetric costs before deploying AI recommendations at scale.

The practical conclusion is clear. Human oversight must be designed as a measurable control rather than assumed from the presence of a human reviewer. AI systems should make it possible and worthwhile for people to question recommendations, verify important claims, and record why they disagree. That creates a stronger basis for scaling AI-assisted decisions while retaining meaningful executive accountability.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Black-box AI recommendations can outperform narrative explanations

More transparency did not produce better decisions in this experiment. Researchers affiliated with Harvard Business School, MIT, and the University of Washington found that simple AI pass-or-reject recommendations improved evaluators’ alignment with the four-expert benchmark. Adding an LLM-generated explanation removed that improvement.

The difference was strongest when the model recommended rejection. An accompanying explanation made evaluators more willing to reject a proposal. That reduced false positives, where evaluators advanced ideas the experts would reject. It also “substantially” increased false negatives, where evaluators rejected ideas the experts believed should advance.

This result challenges a widely used assumption about explainable AI. Many organizations treat explanations as a way to increase transparency and support better oversight. An explanation can achieve the first goal while weakening the second. A convincing narrative changes how people respond to a recommendation, even when its reasoning is wrong.

Early-stage innovation screening makes this problem especially important. The information available is limited, uncertainty is high, and both types of error carry costs. Advancing a weak project can waste capital and management attention. Rejecting a strong project can eliminate a valuable future opportunity. Google Glass and Amazon’s Fire Phone illustrate projects that advanced but ultimately failed. Xerox’s termination of early Ethernet and PostScript projects illustrates the potential cost of rejecting innovations too early.

For executives, explanation design should therefore follow the economics of the decision. A company should define the cost of false positives and false negatives first. It can then test whether AI explanations change those error rates in an acceptable direction. A lower overall error rate can still be undesirable if the system concentrates errors in the category with the highest business cost.

Black-box recommendations also require careful governance. The study supports their value in this specific early-stage screening experiment. It does not establish a universal preference for opaque AI. Organizations still have requirements around auditability, accountability, regulation, and model validation. The practical lesson is narrower and more useful: explanation quality should be measured through its effect on human decisions rather than assumed to improve them.

AI explanations can suppress productive human overrides

The value of human oversight appears when a person detects an AI error and corrects it. The researchers call this a “productive override.” Their experiment found that narrative explanations can suppress this behavior by making model recommendations easier to accept.

LLMs are unusually effective at producing fluent, coherent, and credible-sounding arguments. Those properties can influence evaluators even when the underlying recommendation is incorrect. The researchers found that people can rely on these surface signals and develop an “illusion of explanatory depth.” They feel they understand the basis for a decision despite having limited insight into the model’s actual reasoning.

This creates a practical control problem. A workflow may require human approval for every AI-assisted decision and still produce weak oversight. A reviewer who reads a persuasive rationale and routinely accepts it provides limited independent verification. Formal human involvement and effective human judgment are separate conditions.

For business leaders, productive override rate is therefore a useful governance concept. Teams should test whether employees identify deliberately introduced or historically known model errors. They should examine how often overrides improve outcomes, which recommendation types receive the least scrutiny, and whether adding explanations changes those behaviors.

Interface design can help preserve independent analysis. Organizations can ask reviewers to record an initial assessment before showing the AI recommendation. They can require verification for decisions above defined financial, compliance, safety, or strategic thresholds. Systems can also present uncertainty or competing reasons for acceptance and rejection, approaches specifically suggested by the researchers for future AI-assisted decision systems.

The objective is calibrated reliance. Employees should use AI when it contributes useful information and challenge it when evidence points elsewhere. That requires training, system design, and performance measures that reward correct decisions and effective verification. An approval button alone provides little evidence that meaningful human oversight occurred.

Negativity bias makes AI rejection recommendations more influential

The experiment found an asymmetric effect when LLMs recommended rejecting an innovation proposal. Evaluators became especially likely to follow a rejection when the model provided a written explanation. This reduced false positives, but it also “substantially” increased false negatives. More promising proposals were rejected relative to the four-expert benchmark.

The researchers connect this result to negativity bias, the human tendency to give negative information greater weight than positive information. Rejection can also feel easier to justify under uncertainty. It preserves the status quo, avoids an immediate resource commitment, and limits exposure to the consequences of approving a weak proposal.

LLM explanations can reinforce these tendencies. A model can instantly produce a coherent case against a proposal, giving the evaluator a ready-made justification for rejection. Fluency and apparent expertise can increase the argument’s persuasive force even when the recommendation is wrong. This combination of human bias and machine-generated persuasion can systematically shift decisions toward caution.

That shift has real economic consequences in innovation portfolios. False positives consume money, time, and management capacity. False negatives eliminate projects before their potential becomes clear. The researchers cite Google Glass and Amazon’s Fire Phone as examples of projects that proceeded and failed. They cite Xerox’s termination of early Ethernet and PostScript projects as examples of potentially valuable innovations being abandoned.

Executives should therefore define their tolerance for each error type before introducing AI into screening. An organization pursuing innovation may rationally accept some false positives to preserve exposure to high-value opportunities. An AI system that lowers false positives while sharply increasing false negatives could work against that strategy.

This also affects how performance should be reported. A single accuracy measure can conceal important changes in the composition of errors. Management dashboards should separately track false positives, false negatives, rejection rates, and successful human overrides. Those measures make it possible to see whether AI is improving selection or systematically pushing the organization toward greater risk aversion.

AI explanations should be designed around the decision and its error costs

The same explanation design can produce different business outcomes across different workflows. Researchers affiliated with Harvard Business School, MIT, and the University of Washington argue that organizations should account for the nature of the task and the cost of AI errors when designing recommendation systems.

Early-stage innovation screening provides a clear example. Evaluators have limited information and many uncertain options. Narrative explanations can discourage independent judgment and suppress productive overrides in this setting. A simple recommendation may leave more room for reviewers to investigate the proposal themselves.

Other workflows have different objectives. The researchers identify quality control, compliance screening, and fraud detection as contexts where LLM explanations could support conservative human decisions. Greater willingness to investigate or reject a suspicious case may fit the organization’s desired risk posture. The appropriate design therefore depends on which mistakes carry the greatest operational, financial, legal, or strategic cost.

Binary recommendations are also only one design choice. The researchers propose testing systems that present reasons both for accepting and rejecting an option. Another approach is to disclose model uncertainty according to a fixed threshold. Interfaces can also explicitly invite disagreement, giving users a clearer role in challenging an AI recommendation.

Testing is essential before deployment. Enterprises should measure how an AI system changes human behavior and final outcomes. Relevant measures include accuracy against a credible benchmark, false-positive and false-negative rates, productive overrides, and human-AI agreement. Testing should also determine whether explanations help reviewers detect errors and understand when the model is unreliable.

The decision stage matters as well. The researchers identify later-stage evaluation as an area for further testing. At that point, decision-makers may have fewer options, more information, and stronger incentives to verify AI output. Narrative explanations could therefore have a different effect from the one observed during early innovation screening.

For C-suite leaders, the design principle is straightforward. Treat an AI explanation as part of the decision process and evaluate its behavioral impact. The researchers describe explanations as “behavioral interventions whose effects depend on how evaluators process information under uncertainty.” That framing puts the focus on measurable outcomes: whether the combined human-AI system makes better decisions for the specific risks the business needs to manage.

AI decision systems should be designed to encourage critical evaluation

AI-assisted decision systems need mechanisms that preserve independent human judgment. Researchers affiliated with Harvard Business School, MIT, and the University of Washington found that persuasive LLM explanations can reduce productive overrides. Reviewers become less likely to correct an erroneous recommendation when the model provides a convincing rationale.

The design problem extends beyond model accuracy. The interface determines what information a reviewer sees, when they see it, and how easily they can challenge the recommendation. A highly accurate model can still produce damaging outcomes when its errors receive too little scrutiny. Human behavior therefore becomes part of the system’s performance.

One option proposed by the researchers is to present contrasting narratives. The system could provide reasons to accept a proposal alongside reasons to reject it. This gives decision-makers evidence for competing outcomes and can encourage them to evaluate the underlying case more independently.

Uncertainty disclosures offer another option. Rather than reducing every prediction to a binary pass-or-reject decision, systems could disclose uncertainty according to predefined thresholds. A low-confidence recommendation can then trigger additional review. Organizations can set those thresholds according to the financial, regulatory, operational, or strategic consequences of an error.

Interfaces can also explicitly invite disagreement. A reviewer who challenges an AI recommendation should have a clear process for recording the reasoning and escalating the decision when appropriate. Organizations can then analyze overrides to identify recurring model weaknesses, poorly calibrated confidence, and situations where human expertise consistently adds value.

The sequencing of information also deserves testing. Organizations can evaluate workflows in which reviewers make an initial assessment before seeing the AI output. This approach can preserve an independent judgment that management can compare with the model’s recommendation. Differences between the two assessments become useful signals for deeper review.

These controls should produce measurable outcomes. Executives can monitor false positives, false negatives, successful overrides, human-AI agreement, and performance against a credible benchmark. High agreement alone provides limited evidence of quality. The experiment showed roughly 75% agreement with both black-box and narrative LLM decisions, while narrative explanations still produced problematic changes in decision behavior.

The researchers also recommend studying explanation systems at later stages of decision-making. Evaluators may then have more information, fewer remaining options, and stronger reasons to verify the model’s output. Those conditions could change how AI explanations affect human judgment.

The management objective is effective human-AI collaboration. As the researchers concluded, organizations should treat AI explanations as “behavioral interventions whose effects depend on how evaluators process information under uncertainty.” For executives, that means testing the full decision workflow. Model output, explanation design, human review, escalation rules, and error costs all need to be measured together before an AI-assisted process is scaled.

Recap

AI explanations create a management problem when persuasive language starts to substitute for independent verification. In the study, 228 experienced evaluators showed that a convincing rationale could increase false negatives and suppress useful human overrides. More explanation did not automatically produce better decisions.

For executives, the key metric is the performance of the full human-AI workflow. Track false positives, false negatives, productive overrides, and outcomes against credible benchmarks. Define which errors carry the highest business cost before deciding how much explanation, uncertainty, or human review a system should provide.

This has direct implications for AI governance. Human approval is meaningful only when reviewers have the information, incentives, and authority to disagree. Interfaces should expose uncertainty where useful, support additional verification for high-cost decisions, and make disagreement easy to record and escalate.

AI can improve decision-making at scale. That value depends on preserving independent judgment where uncertainty remains. The strongest systems will make AI recommendations useful without making them difficult to challenge.

Alexander Procter

August 31, 2026

13 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.