The Dumbbell Test: An AI Verification Case Study

The Dumbbell Test is an annotated case study of an ordinary AI chat, a question about the weight increments on an adjustable dumbbell set, reproduced in full and unedited. Chris Cook's side-by-side commentary marks where each error entered: the question asked, the answer given, or the control that was missing. It shares six lessons, each tied to a control a regulated organization would recognize, including why a model's own confidence score isn’t assurance.

Download the Dumbbell Test Case Study (PDF)

What’s inside

The PDF reproduces the full chat, word for word, beside running commentary that tags each moment as an error, an unverified answer, a prompt problem, or a control worth building. It starts with six lessons from the conversation:

  • Read the whole input.

  • A confidence score is not a control.

  • Visible activity is not verification.

  • Pushback moves the model more than evidence does.

  • Treat user-supplied facts as unverified.

  • The reviewer's expertise is the control.

Who it’s for

Audit, risk, compliance, and technology leaders in regulated organizations who are deciding how much to rely on AI-generated work, and board and audit committee members who want a concrete example of why verification controls matter. The topic is low stakes by design, but the same failures in a credit memo, a policy summary, or a regulatory filing carry real consequences.

Related resources

To discuss how these controls apply to your AI program, book a conversation.

Next
Next

Head of AI — Role Specification