
Key takeaways on IBM AI judgment gap
- 1,500 CHROs and 8,800 employees
- April-June 2026 across 21 and 28 geographies respectively
- 60% of employees worry AI is eroding their skills
What changed and why it matters
The IBM AI judgment gap identifies a practical adoption problem: executives want employees to supervise, validate and override AI, but employees do not rank that judgment nearly as highly. In IBM’s survey, 71% of chief human resources officers called those oversight skills essential, while 29% of employees ranked judgment as important. That 42-point difference makes training design—not tool access—the urgent operational question.
The release is a concrete new research event, published September 21. IBM says its Institute for Business Value surveyed 1,500 CHROs and 8,800 employees between April and June 2026. The executive and employee samples covered 21 and 28 geographies respectively. The numbers are therefore broad, but they remain survey responses commissioned and interpreted by IBM; they should not be read as causal proof that AI itself weakens reasoning.
The sharper warning comes from the employee side. Sixty percent said they worry AI is eroding their skills, with critical thinking cited most often as declining. At the same time, IBM reports that 57% of CHROs put critical thinking among the most important workforce capabilities and 48% included human judgment. The tension is not whether judgment matters. It is whether day-to-day incentives, training and job design make that importance visible.
The operational test
Governance is also fragmented. IBM says 46% of organizations do not involve the CHRO when AI strategy is defined, even though workforce redesign lands squarely inside the HR remit. It also reports that 43% of employees think they would carry the blame when an AI-assisted decision goes wrong. That combination can encourage passive use: workers are held accountable for outputs without being given explicit authority or practice to challenge them.
For employers, the useful response is to translate human oversight into observable work. A review policy should name the decisions that require human validation, the evidence reviewers must inspect and the point at which a worker can stop or override an automated flow. Training can use realistic failure cases rather than generic prompting lessons. Measurement should track catches, escalations and corrected decisions, not merely seats activated or time saved.
This approach also connects to broader agent safety. Our guide to why agent permissions need hard boundaries shows that autonomy is safest when systems have clear limits. The lesson from how trust checks fail in automated systems is similar: a confident output is not a substitute for verification.
IBM’s study does not supply a universal benchmark for good judgment, and the independent coverage largely reports the same survey rather than reproducing it. The defensible conclusion is narrower. Companies scaling AI need an explicit human-judgment layer, and HR cannot be treated as a downstream communications function after the technical strategy is set.
Facts table
| Item | Verified fact | Source |
|---|---|---|
| Study sample | 1,500 CHROs and 8,800 employees | IBM Newsroom |
| Fieldwork | April-June 2026 across 21 and 28 geographies respectively | IBM Newsroom |
| Leader priority | 71% of CHROs called supervising, validating and overriding AI essential | IBM Newsroom |
| Employee ranking | 29% of employees ranked judgment as important | IBM Newsroom |
| Skills concern | 60% of employees worry AI is eroding their skills | IBM Newsroom |
Frequently asked questions
What is the IBM AI judgment gap?
It is the difference between the 71% of surveyed CHROs who prioritize supervising and overriding AI and the 29% of surveyed employees who rank judgment as important.
Does the study prove AI reduces critical thinking?
No. It reports perceptions and priorities from a survey; it does not establish that AI caused a measurable decline.
What should companies do next?
They can define human override points, train teams to challenge outputs, and measure judgment quality alongside adoption.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



