Evaluations require the Advanced plan or higher.
How it works
- You build a set of cases: documents plus their expected values.
- You run the set. Each case runs the workflow as a test run, so nothing is delivered anywhere.
- Each case’s output is compared field by field against its labels.
- The run reports field accuracy, document pass rate, and a per-field breakdown.
Creating a set
Open a workflow, then Evaluations from the actions menu, and create a set. A set scores one Extract node. Leave the scored node empty when the workflow has exactly one Extract node; pick one explicitly when it has several.Adding cases
There are two ways to label a document. From an approved review. This is the fast path. Every correction a reviewer made is, by definition, the right answer. Open a run whose review task was approved, then Add to eval set, and pick the set. The approved values become the case’s labels. Values that do not map onto a field in the scored node’s schema are reported back as unmapped paths, so you can see exactly what the case does not cover. Through the API.POST /evaluation-sets/{evaluationSetId}/cases takes a document and the expected values
directly, for labels that did not come from a review.
A case does not have to label every field: it scores what it labels and ignores the rest.
The app builds cases from approved reviews. Labeling a document by hand, and configuring the key columns
described below, are available through the API only.
Line items
A case can label a table as well as header fields. Each row lists the expected values for the columns you care about, and rows are matched to the extracted table before the cells are compared. Rows are matched by position by default. When the extraction can reorder rows, the case can name one or more key columns (a SKU or line code) and rows are matched by those instead. Every labeled cell counts toward accuracy. A row the document should have produced but did not counts each of its labeled cells as missing. A row the extraction invented, which the case does not describe, counts against the score as well: it has no labeled cells to compare, so it costs the table one miss of its own. Three invented rows cost three times what one does, and a case cannot reach 100% with any of them present. Cases seeded from a review are matched by position: the approved table came from that same document, in that same order. Label the table, never a single column on its own. A column holds one value per row, so a label naming it directly has nothing to compare against and is left out of the accuracy rather than counted as a miss.Running a set
Press Run evaluation. Cases run four at a time, so a large set takes a few minutes. Evaluation runs are ordinary runs with one difference: they are test runs, so no callback fires, nothing is written to a connector, no email goes out, and review nodes pass straight through. You can open any case’s run from the results and inspect it like any other.What a run costs
An evaluation run bills like any other run: the AI steps it executes consume credits at the normal per-node rate. Two things keep that small:- Steps whose configuration has not changed since a previous run of the same document are served from cache and bill nothing. Re-running a set after a change usually pays only for the changed node and everything after it.
- A first run of a new set pays in full, because there is nothing to replay.
Reading the results
A case is passed when every field it labels matched, failed when at least one did not, errored
when there was nothing to score, and skipped when the run stopped before reaching it. A case errors
either because its run produced no output or because none of its labels resolve against the schema any
more, and the case says which. An errored case is left out of the document pass rate rather than counted
against it.
Fields whose labels no longer match any field in the schema are left out of the accuracy entirely rather
than counted as misses. They are listed in the per-field breakdown with no accuracy of their own, and the run
says how many there were: a stale label is a signal to
fix the set, not a failing score.
Comparing two runs
Pick a run to compare against from the run’s header. The comparison shows per-field deltas with the largest regressions first, and the cases whose result changed, with passed-to-failed first. Comparison is keyed by case and field, so it survives graph edits and schema renames. Each side reports the graph revision it scored, so you can see which two versions of the workflow you are comparing.Trends
The set page charts accuracy across the last completed runs, either overall or for a single field. A field the run never scored leaves a gap in the line rather than a zero: never measured and measured zero are different facts.Blocking activation on a passing run
A set can be marked as an activation gate. While it is, the workflow cannot be activated until that set has a completed, passing run against the exact graph revision you are activating. Saving is never blocked. You can keep editing a paused workflow freely; the gate applies only when you go live. On the workflow page and in the workflow editor, the Activate action is disabled and says why: the set has not been run against this revision, the run scored below the threshold, or the run mixed revisions. Activating from the workflows list reports the same reason when the attempt is refused. Because saving creates a new revision, activating a workflow you have just edited means saving, running the gating set against the new revision, and then activating. The pass threshold is the field accuracy the gating run must reach. It defaults to 100%, meaning every scored field in every case must match.Keeping a set healthy
- Schema drift. A case is flagged when the scored node’s schema has changed since the case was authored. It is a warning, never a block, and renames alone do not trigger it. Re-seed the case from a fresh approved review to bring its labels back in step.
- Disabling instead of deleting. Disable a case to leave it out of the next run without losing the labeling work behind it. Deleting a case also deletes its results from every past run, so a case you are still comparing against is better disabled than deleted.
- Deleting documents. A document referenced by a case cannot be deleted. Remove the case first. A golden set is a curated asset, and letting a cleanup silently shrink one would change every future number with no record of why. Retention leaves those documents alone for the same reason: a case’s document is never reclaimed by the workflow’s retention window while the case still points at it.
- Deleted runs. A case’s result outlives the run that produced it, so retention or a manual deletion never erases evaluation history.
Related
Reviews
Approved reviews are the fastest source of labels
Extract
The node an evaluation set scores
Test mode
How test runs differ from live runs
Credits
Per-node costs and worked examples