Research and evaluation

Better tutoring starts with better questions.

Amend is built around a simple standard: feedback should help a student think and continue. This page documents how we turn that standard into concrete evaluation criteria for handwritten STEM work.

Read the open benchmark

What good looks like

An answer can be right while the tutoring is wrong.

A completed solution may be accurate and still interrupt learning. We evaluate whether feedback understands the visible work, identifies a supported issue, and leaves the student with a useful next action.

Find the first supported issue

A useful tutor should locate the earliest step that evidence shows is wrong, not merely report that the final answer differs.

Preserve student reasoning

Feedback should create a productive next move—a question, hint, or correction—without replacing the student's work by default.

Treat ambiguity as ambiguity

Unclear handwriting or incomplete work should trigger clarification or uncertainty, not a confident invented interpretation.

Evaluate the whole interaction

Correctness matters alongside error location, instructional usefulness, calibration, and whether the student can continue.

Evaluation stack

Measure the path, not just the destination.

  1. 01

    Interpretation

    What written expression, diagram, or step did the system understand from the page?

  2. 02

    Verdict

    Is the visible work correct, partially correct, incorrect, or too unclear to judge?

  3. 03

    Localization

    If there is an error, which earliest line contains it and what evidence supports that call?

  4. 04

    Instruction

    Does the response offer the smallest useful next move without taking over?

  5. 05

    Calibration

    Does the response express uncertainty when the handwriting or mathematical evidence is insufficient?

Published now

Handwritten Math Tutor Benchmark v0.1

A reproducible starter set of written-work cases, expected behaviors, failure conditions, and a scoring rubric. The protocol is public; product performance scores are not yet published.

Explore the benchmark

Important limitation

AI feedback can be wrong.

Handwriting interpretation and generated tutoring are probabilistic. Amend is a learning aid, not a substitute for an instructor, and students should verify important results against course materials.

How Amend Evaluates AI Tutoring | Amend