> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ingestly.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluations

> Measure a workflow's extraction accuracy against a set of labeled documents, and see what a change did to it.

An **evaluation set** is a group of labeled documents, a golden set, that a workflow is scored against. Each
**case** in the set pairs a document with the values that document should produce. Running the set produces an
**evaluation run** that scores every case and reports one number for the whole workflow.

This answers the question that guessing cannot: did that prompt edit, schema change, or precision bump make
extraction better or worse?

<Note>Evaluations require the **Advanced** plan or higher.</Note>

## How it works

1. You build a set of cases: documents plus their expected values.
2. You run the set. Each case runs the workflow as a test run, so nothing is delivered anywhere.
3. Each case's output is compared field by field against its labels.
4. The run reports field accuracy, document pass rate, and a per-field breakdown.

Labels are keyed by the field's identity in the schema, not by its name, so renaming a field in the extract
schema leaves every label attached to it.

## Creating a set

Open a workflow, then **Evaluations** from the actions menu, and create a set. A set scores one **Extract**
node. Leave the scored node empty when the workflow has exactly one Extract node; pick one explicitly when it
has several.

## Adding cases

There are two ways to label a document.

**From an approved review.** This is the fast path. Every correction a reviewer made is, by definition, the
right answer. Open a run whose review task was approved, then **Add to eval set**, and pick the set. The
approved values become the case's labels.

Values that do not map onto a field in the scored node's schema are reported back as unmapped paths, so you
can see exactly what the case does not cover.

**Through the API.** `POST /evaluation-sets/{evaluationSetId}/cases` takes a document and the expected values
directly, for labels that did not come from a review.

A case does not have to label every field: it scores what it labels and ignores the rest.

<Note>
  The app builds cases from approved reviews. Labeling a document by hand, and configuring the key columns
  described below, are available through the API only.
</Note>

### Line items

A case can label a table as well as header fields. Each row lists the expected values for the columns you care
about, and rows are matched to the extracted table before the cells are compared. Rows are matched by position
by default. When the extraction can reorder rows, the case can name one or more key columns (a SKU or line
code) and rows are matched by those instead.

Every labeled cell counts toward accuracy. A row the document should have produced but did not counts each of
its labeled cells as missing. A row the extraction invented, which the case does not describe, counts against
the score as well: it has no labeled cells to compare, so it costs the table one miss of its own. Three
invented rows cost three times what one does, and a case cannot reach 100% with any of them present.

Cases seeded from a review are matched by position: the approved table came from that same document, in that
same order.

Label the table, never a single column on its own. A column holds one value per row, so a label naming it
directly has nothing to compare against and is left out of the accuracy rather than counted as a miss.

## Running a set

Press **Run evaluation**. Cases run four at a time, so a large set takes a few minutes.

Evaluation runs are ordinary runs with one difference: they are test runs, so no callback fires, nothing is
written to a connector, no email goes out, and review nodes pass straight through. You can open any case's run
from the results and inspect it like any other.

### What a run costs

An evaluation run bills like any other run: the AI steps it executes consume credits at the normal per-node
rate. Two things keep that small:

* Steps whose configuration has not changed since a previous run of the same document are served from cache
  and bill nothing. Re-running a set after a change usually pays only for the changed node and everything
  after it.
* A first run of a new set pays in full, because there is nothing to replay.

The run reports what it actually charged when it finishes.

## Reading the results

| Metric                  | What it means                                                         |
| ----------------------- | --------------------------------------------------------------------- |
| **Field accuracy**      | Matched fields divided by scored fields, across every case in the run |
| **Document pass rate**  | Cases where every scored field matched, divided by the cases that ran |
| **Per-field breakdown** | Accuracy and mean confidence for each scored field                    |
| **Credits**             | What the run actually cost                                            |

A case is **passed** when every field it labels matched, **failed** when at least one did not, **errored**
when there was nothing to score, and **skipped** when the run stopped before reaching it. A case errors
either because its run produced no output or because none of its labels resolve against the schema any
more, and the case says which. An errored case is left out of the document pass rate rather than counted
against it.

Fields whose labels no longer match any field in the schema are left out of the accuracy entirely rather
than counted as misses. They are listed in the per-field breakdown with no accuracy of their own, and the run
says how many there were: a stale label is a signal to
fix the set, not a failing score.

## Comparing two runs

Pick a run to compare against from the run's header. The comparison shows per-field deltas with the largest
regressions first, and the cases whose result changed, with passed-to-failed first.

Comparison is keyed by case and field, so it survives graph edits and schema renames. Each side reports the
graph revision it scored, so you can see which two versions of the workflow you are comparing.

<Warning>
  A run whose cases did not all score against the same graph revision is flagged as **mixed revisions**. That
  happens when the workflow was saved while the run was in flight. Re-run the set for a comparable number.
</Warning>

## Trends

The set page charts accuracy across the last completed runs, either overall or for a single field. A field the
run never scored leaves a gap in the line rather than a zero: never measured and measured zero are different
facts.

## Blocking activation on a passing run

A set can be marked as an **activation gate**. While it is, the workflow cannot be activated until that set
has a completed, passing run against the exact graph revision you are activating.

Saving is never blocked. You can keep editing a paused workflow freely; the gate applies only when you go
live. On the workflow page and in the workflow editor, the **Activate** action is disabled and says why: the
set has not been run against this revision, the run scored below the threshold, or the run mixed revisions.
Activating from the workflows list reports the same reason when the attempt is refused.

Because saving creates a new revision, activating a workflow you have just edited means saving, running the
gating set against the new revision, and then activating.

The **pass threshold** is the field accuracy the gating run must reach. It defaults to 100%, meaning every
scored field in every case must match.

<Tip>
  A gate set with no enabled cases does not gate anything: an empty set can never run, so it could never pass.
  Marking a set as a gate is a promise to fill it.
</Tip>

## Keeping a set healthy

* **Schema drift.** A case is flagged when the scored node's schema has changed since the case was authored.
  It is a warning, never a block, and renames alone do not trigger it. Re-seed the case from a fresh approved
  review to bring its labels back in step.
* **Disabling instead of deleting.** Disable a case to leave it out of the next run without losing the
  labeling work behind it. Deleting a case also deletes its results from every past run, so a case you are
  still comparing against is better disabled than deleted.
* **Deleting documents.** A document referenced by a case cannot be deleted. Remove the case first. A golden
  set is a curated asset, and letting a cleanup silently shrink one would change every future number with no
  record of why. Retention leaves those documents alone for the same reason: a case's document is never
  reclaimed by the workflow's retention window while the case still points at it.
* **Deleted runs.** A case's result outlives the run that produced it, so retention or a manual deletion never
  erases evaluation history.

## Related

<CardGroup cols={2}>
  <Card title="Reviews" icon="user-check" href="/reviews/introduction">
    Approved reviews are the fastest source of labels
  </Card>

  <Card title="Extract" icon="table" href="/nodes/extract">
    The node an evaluation set scores
  </Card>

  <Card title="Test mode" icon="flask" href="/guides/test-mode">
    How test runs differ from live runs
  </Card>

  <Card title="Credits" icon="coins" href="/admin/credits">
    Per-node costs and worked examples
  </Card>
</CardGroup>
