Open evaluation protocol · Version 0.1

Handwritten Math Tutor Benchmark

The Amend Handwritten Math Tutor Benchmark v0.1 is a public 12-case protocol for testing whether an AI tutor judges written STEM work correctly, locates the earliest supported error, gives a useful next step, and admits when the evidence is unclear. It publishes the cases and scoring method, not performance results.

Published
August 18, 2026
Updated
August 21, 2026
Cases
12 across six STEM areas
Status
Protocol public; scores pending

No performance claim is attached to this release.

Version 0.1 publishes the cases and scoring rules so the evaluation can be inspected before results are reported. It is a small diagnostic set, not evidence of broad classroom effectiveness.

The starter set

Cases that reward evidence, not confidence.

ID

ALG-01

Skill

Linear equations

Visible work

3(x − 2) = 12 → 3x − 2 = 12

Expected behavior

Locate the distribution error: −2 should also be multiplied by 3.

Failure condition

Accepts the line or comments only on the eventual answer.

ID

ALG-02

Skill

Equivalent forms

Visible work

2(x + 3) = 2x + 6

Expected behavior

Mark the transformation correct and avoid inventing an error.

Failure condition

Rejects a mathematically equivalent expression.

ID

ALG-03

Skill

Ambiguous handwriting

Visible work

x² + 5x + 6 = (x + 2)(x + ?)

Expected behavior

Ask for clarification when the final factor is not legible.

Failure condition

Confidently assumes a symbol without enough visual evidence.

ID

CALC-01

Skill

Chain rule

Visible work

y = (x² + 1)³ → y′ = 3(x² + 1)²

Expected behavior

Identify the missing inner derivative and prompt for d/dx(x² + 1).

Failure condition

Calls the derivative complete or gives only the final expression.

ID

CALC-02

Skill

Product rule

Visible work

d/dx[x² sin x] = 2x cos x

Expected behavior

Identify that the factors were differentiated together instead of using both product-rule terms.

Failure condition

Focuses on notation while missing the mathematical error.

ID

CALC-03

Skill

Indefinite integrals

Visible work

∫ 2x dx = x²

Expected behavior

Recognize the antiderivative and ask about the constant of integration.

Failure condition

Labels the mathematical core wrong or ignores the missing constant entirely.

ID

PHY-01

Skill

Units

Visible work

v = d/t = 100 m / 20 s = 5 m

Expected behavior

Locate the unit error in the final line: velocity should be expressed in m/s.

Failure condition

Accepts the unit or recomputes an already correct numerical value.

ID

PHY-02

Skill

Sign convention

Visible work

Up is positive; a falling object has a = +9.8 m/s²

Expected behavior

Connect the acceleration sign to the stated positive direction.

Failure condition

Corrects the sign without referring to the chosen convention.

ID

CHEM-01

Skill

Stoichiometry

Visible work

2 mol H₂ × 2 mol H₂O / 2 mol H₂ = 2 mol H₂O

Expected behavior

Accept the mole ratio and cancellation as correct.

Failure condition

Invents an error because the coefficients repeat.

ID

STAT-01

Skill

Sample mean

Visible work

mean of 2, 4, 9 = (2 + 4 + 9)/2 = 7.5

Expected behavior

Locate the first error at the denominator and ask how many observations there are.

Failure condition

Comments only that 7.5 is wrong.

ID

GEO-01

Skill

Pythagorean theorem

Visible work

legs 3 and 4 → c² = 3² + 4² = 25 → c = 5

Expected behavior

Mark the reasoning correct without adding unnecessary remediation.

Failure condition

Finds a nonexistent issue or replaces the concise solution.

ID

CAL-01

Skill

Insufficient evidence

Visible work

A partially cropped page showing only '= 12'

Expected behavior

State that there is not enough visible work to judge and request more context.

Failure condition

Infers the original problem or declares the answer correct.

Scoring rubric

Eight points per case.

Each dimension receives 0, 1, or 2 points. Evaluators score the response shown to the student, without inferring hidden system intent.

Verdict accuracy

2: correct judgment · 1: directionally right but incomplete · 0: wrong or unsupported

Error localization

2: earliest supported issue · 1: right issue, imprecise location · 0: misses or invents it

Instructional usefulness

2: smallest useful next move · 1: helpful but over-solves · 0: blocks or misleads

Uncertainty calibration

2: confidence matches evidence · 1: minor over/under-confidence · 0: confident fabrication

Reproduction procedure

Run the same test every time.

  1. 1

    Render each visible-work string by hand on a blank whiteboard without additional context.

  2. 2

    Ask for feedback on the current work using the same neutral request for every case.

  3. 3

    Record the first tutor response without retries, edits, or follow-up clarification.

  4. 4

    Have two reviewers score all four dimensions independently, then resolve disagreements by citing the rubric.

  5. 5

    Report per-dimension scores, abstentions, and case-level failures—not only one aggregate number.

Why this exists

Tutoring quality deserves a test people can inspect.

Read our evaluation principles
Handwritten Math Tutor Benchmark v0.1 | Amend