> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ingestly.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Self-learning

> Turn approved reviews into examples that improve later extractions, and see whether it helps.

Self-learning is an opt-in, per-workflow feature. When it is on, every approved review of an [Extract](/nodes/extract) node is kept as an example, and later extractions in the same workflow consult the most similar past documents before they read the new one. The workflow's **Learning** tab shows what the workflow has learned and whether it is helping.

Self-learning is not part of any plan. It is granted per organization; contact us to have it turned on for yours. Without the grant the workflow editor shows no **Learning** tab and no self-learning action, and the Api refuses every self-learning request.

Self-learning is off per workflow until you turn it on. Open a workflow, then **Enable self-learning** from the [workflow editor](/workflows/editor)'s **More actions** menu. That turns capture on with the workflow's stored settings and opens the **Learning** tab, where **Settings** holds every dial.

## How it works

1. A reviewer approves a [review task](/reviews/queue) for an Extract step.
2. Ingestly keeps that extraction as a learning example: the values the model extracted, the values the reviewer approved, and every change between them, including fields the reviewer removed or added.
3. On later runs of the same workflow, the most similar past documents are handed to the model as reference data before it extracts.
4. The model still reads every value from the document in front of it. The examples show it what a document of this shape looks like after a person has corrected it; they never supply values of their own.

Examples belong to the Extract node that produced them, and they are never shared with another workflow.

Corrections feed a second, slower loop as well. Ingestly distills repeated corrections into short written [instructions](#instructions) that you approve before they reach the model, so a correction the reviewers keep making becomes a rule instead of an example.

## What reviewers see

When self-learning is on, the review panel footer says "Self-learning is on: approved values from Extract steps become examples for later runs." so a reviewer knows the approval trains the workflow.

To keep one review out of the bank, tick **Don't learn from this review** in the approve dialog. The review is approved and the run resumes exactly as usual; only the example is skipped. An approved review that was excluded reads "Excluded from learning" in the same footer, so the decision stays visible afterwards.

A review whose Review node reads a custom source or limits which fields the reviewer can edit never becomes an example either, with or without the checkbox: the reviewer vetted only part of the document, so treating the rest as confirmed would teach the workflow values nobody checked.

## Holdout

A holdout reserves a share of runs that extract with no learned context at all, so the two can be compared. Runs that received learned context are **treatment** runs, holdout runs are the control.

Set the share with **Holdout percent**. It is 0, meaning off, by default. Turn it on when you want a number you can trust for whether learning is helping, and leave it off when you want every run to use everything the workflow knows.

## The Learning tab

Once self-learning is on, the [workflow editor](/workflows/editor) carries a **Learning** tab beside **Builder**, **Documents**, **Runs**, and **Insights**. Its title line shows whether learning is on or off, a range control that sets the window every statistic covers, **Settings**, which opens the settings sheet, and a **More actions** menu with **Reset Learning**. Below the title line the tab has four sections, **Overview**, **Fields**, **Instructions**, and **Examples**. Reading the tab needs the `Core.Workflow.Read` permission; the actions and the row actions that change the bank appear only for roles holding `Core.Workflow.Update` (see [roles and permissions](/admin/roles-and-permissions)).

Turning learning off in **Settings** keeps the tab open until you leave it; it is gone from the tab bar the next time you open the workflow, and **Enable self-learning** is back in the **More actions** menu.

### Overview

The headline numbers are:

| Statistic               | What it means                                                                                                                                                                                                                                                                    |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Documents reviewed**  | Every document reviewed in the window. The treatment and holdout counts beside it are the two arms the comparison uses: a run that was meant to learn but found nothing similar is counted in the total and in neither arm, so the two arms can add up to less than the headline |
| **Live examples**       | Examples an extraction can consult right now, with the archived count beside them                                                                                                                                                                                                |
| **First-pass accuracy** | Per arm, the share of reviewed fields the reviewer left unchanged, with the sample size it is based on                                                                                                                                                                           |
| **Clean approvals**     | Documents a reviewer approved without changing anything                                                                                                                                                                                                                          |

A rate never appears without its sample size. Compare the treatment and holdout figures only once both arms have enough reviewed documents to mean something.

A daily chart plots reviewed and clean documents per arm across the window, so you can see the trend rather than one aggregate.

Below the chart, the bank table lists one row per Extract node with its live and archived example counts, how many live examples are not embedded yet, how many vendors they cover, an **Instructions** count reading approved then proposed, when the node last captured, and a status:

| Status             | What it means                                                                                                       |
| ------------------ | ------------------------------------------------------------------------------------------------------------------- |
| **Ready**          | The node is in the workflow with its current schema and nothing needs attention, whether it has examples yet or not |
| **Schema changed** | Every example was archived because the schema changed, and the node is collecting again                             |
| **Stale**          | Live examples were collected against a different schema than the node has now                                       |
| **Not in graph**   | The node these examples came from is no longer in the workflow                                                      |

Retrieval matches on an example's embedding, so an example counted under **Not embedded** is skipped until nightly [maintenance](#maintenance) embeds it from the document's text.

### Fields

The **Fields** section breaks accuracy down by field path. Each row shows how often the field was reviewed, how often it was corrected, removed, or added, its first-pass accuracy per arm with the sample sizes, and how many of its corrections are contested. Fields with the most changes come first, so the worst field is the first one you read.

A contested correction is one that contradicts another: the same original value was corrected two different ways on the same field. Nightly [maintenance](#maintenance) marks them, and a contested correction stops reaching the model.

Pick a node to scope the table to one Extract node, and click a row to chart that field's daily trend across the window.

### Instructions

An instruction is a short rule distilled from the corrections reviewers made, such as "Take the invoice number from the stamp in the header, not from the reference line." Where an example shows the model a corrected document, an instruction tells it what to do, so a correction reviewers keep repeating is handed over once instead of every time.

The **Instructions** section lists them. Pick a node to scope the table to one Extract node, and use the status toggle to move between **Proposed**, the default, **Approved**, **Rejected**, **Retired**, and **All**. Each row shows:

| Column      | What it means                                                                                               |
| ----------- | ----------------------------------------------------------------------------------------------------------- |
| Node        | The Extract node the instruction belongs to                                                                 |
| Field       | The field path the instruction applies to, or "Whole document" when it applies to the extraction as a whole |
| Instruction | The text handed to the model                                                                                |
| Support     | How many corrections back the instruction                                                                   |
| Gate        | **Not gated**, **Passed**, or **Failed**. Hover the badge for the detail                                    |
| Status      | **Proposed**, **Approved**, **Rejected**, or **Retired**                                                    |
| **Stale**   | A badge shown when the node's schema changed after the instruction was proposed                             |
| Proposed    | When the instruction was proposed                                                                           |

Each row carries these actions:

* **Approve** puts the instruction to work on later extractions.
* **Reject** leaves it out. The **Reject Instruction** dialog takes an optional reason, so whoever reads the row later knows why.
* **Retire** takes an approved instruction back out again. Approved rows only.
* **View evidence** opens a panel with the rationale and the corrections the instruction was distilled from, so you can check it against what reviewers actually changed.

The header carries two more:

* **Add Instruction** writes one yourself: choose the Extract node, an optional field path, the instruction text of 12 to 400 characters, and an optional rationale. A manual instruction is approved immediately and attributed to you.
* **Distill Now** distills on the spot, for the selected node or for every Extract node when none is selected. It reports how many instructions were proposed, or why a node was skipped: learning is off, the node is not an Extract node, it has no schema, it has too few new corrections, or no field has enough support.

Nightly, Ingestly distills every enabled workflow's Extract nodes that gathered at least 10 new corrections since their last proposal, with at least one field corrected 3 or more times across at least 2 different documents, so one document with a whole column wrong does not ground a rule on its own. Nodes below those thresholds are left alone until they collect more. Distillation is not billed.

The gate is where an [evaluation set](/workflows/evaluations) checks an instruction before you trust it. It reads **Not gated** until an evaluation set covers the node, which means the gate has no verdict to offer and the judgement is yours.

A proposal never reaches the model until you approve it. Approved instructions are added to the extraction prompt after the node's **Extraction hints**, as a `### LEARNED INSTRUCTIONS` section, and approving, rejecting, or retiring one changes that prompt exactly once, so the next run works from the new set.

Only instructions written against the node's current schema are used. Change the schema and the older ones read **Stale**: they stay on the page for the record, and they no longer reach the model.

Ask the assistant to explain a proposed instruction. On the **Instructions** section it can read any instruction listed there, including the one you have open, and the corrections it was distilled from, so it answers why the instruction was proposed, what evidence backs it, and whether to approve it. It describes what reviewers changed rather than quoting values out of your documents, and the decision stays yours: the assistant never approves, rejects, or retires an instruction.

<Note>
  An instruction never supplies a value. It tells the model where to read a value, or how to read it; the values still come from the document in front of it.
</Note>

### Examples

The **Examples** section lists the bank itself, live or archived, with the node, the vendor, the arm, the corrections the reviewer made, the page count, when it was captured, and its status. Click a row to open the drawer, which shows every correction alongside the extracted payload and the approved payload.

A vendor that came from the email sender rather than from the **Vendor key field** carries a muted **from sender** hint, in the table and in the drawer, so you can tell the two groupings apart.

* **Archive** takes an example out of later extractions without deleting it.
* **Restore** puts an archived example back. Only examples archived by a user for the current schema can be restored: an example archived automatically, or one collected against an older schema, stays archived.

### Settings

| Setting                              | What it does                                                                                                                                                                                                                                              |
| ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Enable**                           | Turns capture on for this workflow. Off by default.                                                                                                                                                                                                       |
| **PII handling**                     | **Skip** withholds values that look like personal data, **Redact** stores a placeholder such as `[EMAIL]`, **Allow** stores them verbatim. Fields marked `x-pii` in the schema, and formats such as `ssn` or `iban`, are always treated as personal data. |
| **Similar documents per extraction** | How many past documents an extraction may consult (0 to 12). 0 keeps collecting examples without using them yet.                                                                                                                                          |
| **Holdout percent**                  | Share of runs (0 to 50) that extract without learned context, so accuracy with and without learning can be compared. 0 turns holdout off.                                                                                                                 |
| **Vendor key field**                 | A payload path, such as `vendor.name`, whose approved value groups examples per vendor. When the path yields no value and the document arrived by email from an authenticated sender, the sender's domain becomes the key instead.                        |
| **Max examples per node**            | How many live examples one Extract node may keep. Nightly [maintenance](#maintenance) archives the overflow.                                                                                                                                              |
| **Max examples per vendor**          | How many live examples one vendor key may keep within a node. Nightly [maintenance](#maintenance) archives the overflow, the per-vendor cap first and the per-node cap after it.                                                                          |

**Reset Learning** archives every live example of the workflow, or of one Extract node when you scope it, and keeps the accuracy history so the numbers you have already collected survive. Because it empties the bank, the **Reset Learning** dialog asks you to type the workflow name before it proceeds.

<Note>
  Examples outlive the documents they came from: retention deleting a document does not remove what the workflow learned from it. Archive an example, or reset the workflow's learning, to make the workflow forget it.
</Note>

## Schema changes

Examples are collected against the schema the Extract node had at the time. Change that schema and the examples collected against the previous one are archived automatically, because values that no longer fit the fields would mislead the next extraction rather than help it. The node starts collecting again with the next approved review, and its accuracy history is kept. A review of a run that extracted with the previous schema and is approved after the change is still recorded, for the history, but its example stays archived and never reaches the model.

Learned instructions are treated the same way but are never thrown out: an instruction written against the previous schema stops reaching the model and reads **Stale** on the **Instructions** section, so the decision you made about it stays on the record.

## Maintenance

Every night Ingestly tidies the bank of each workflow that has self-learning on, so it stays the size the settings promise and nothing misleading reaches an extraction. None of it is billed, and none of it needs you.

The nightly pass runs these steps per workflow, in order:

1. **Archives what a schema change orphaned.** Live examples collected against a schema the Extract node no longer has, and examples whose node has left the workflow, are archived. See [Schema changes](#schema-changes).
2. **Enforces the caps.** Live examples beyond **Max examples per vendor** within a node are archived first, then those beyond **Max examples per node**. Examples the reviewer changed nothing in go first, then the oldest; the newest survive. An archived example reads **Evicted** on the **Examples** section.
3. **Marks and clears contested corrections.** When the same original value was corrected to two or more different values on one field in the last 90 days, every correction in that disagreement is marked contested. Contested corrections are hidden from retrieval and from distillation, so a contradiction never becomes a rule. The example itself stays live, and the **Fields** section shows the contested count. When the disagreement is gone from the last 90 days, because the conflicting examples were archived or aged out, the corrections still inside that window are cleared again; older contested corrections stay hidden.
4. **Embeds examples that have none.** An example captured before its text was available has no embedding, and retrieval matches on the embedding. The pass embeds it from the document's text. An example whose document has been deleted, or whose document carries no text, is skipped for good and no longer counts as **Not embedded**; an embedding that fails is tried again the next night.
5. **Trims field statistics.** Daily field statistics older than 400 days are deleted. No surface reads past 90 days, so this changes nothing you can see.

<Note>
  Maintenance archives examples, it never deletes them. An archived example stays on the **Examples** section with the reason it was archived, and it no longer reaches an extraction.
</Note>

## Personal data

**PII handling** decides what a captured example is allowed to store:

| Mode       | What it stores                                                                        |
| ---------- | ------------------------------------------------------------------------------------- |
| **Skip**   | Values that look like personal data are withheld from the example                     |
| **Redact** | Values that look like personal data are replaced with a placeholder such as `[EMAIL]` |
| **Allow**  | Values are stored verbatim                                                            |

The policy applies at capture time, so what the example holds is what the **Examples** drawer shows. Changing the mode does not rewrite examples that were already captured.

## Run detail

A run's Extract step carries a badge saying what learned context it received:

| Badge                              | What it means                                                                       |
| ---------------------------------- | ----------------------------------------------------------------------------------- |
| **Learned: N similar documents**   | The extraction consulted N past documents ("1 similar document" when there was one) |
| **Learning: no similar documents** | Self-learning is on, but the bank had nothing similar to offer yet                  |
| **Holdout**                        | The run was the control: it extracted with no learned context                       |

The badge appends ", instructions applied" when the extraction also received approved [instructions](#instructions), so **Learned: 3 similar documents, instructions applied** means the run used both. Hover the badge to read what it was given. A holdout run receives neither examples nor instructions.

## Workflow insights

A workflow's **Insights** tab carries a **Self-learning** card with the live example count, documents reviewed, an **Approved instructions** row, treatment and holdout first-pass accuracy with their sample sizes, and the clean-approval rate. **Open self-learning** switches to the **Learning** tab. When the workflow has self-learning off, the card reads "Self-learning is off for this workflow." and **Open self-learning** still opens the **Learning** tab, where the settings live. An organization without the self-learning grant sees the card without the button.

## Related

<CardGroup cols={2}>
  <Card title="Managing workflows" icon="browser" href="/workflows/managing">
    Where self-learning sits among the workflow's settings
  </Card>

  <Card title="Review queue" icon="list-check" href="/reviews/queue">
    Approving a review is what captures an example
  </Card>

  <Card title="Evaluations" icon="scale-balanced" href="/workflows/evaluations">
    Score a workflow against labeled documents instead of measuring reviews
  </Card>
</CardGroup>
