When to use extract
- You want structured fields out of a document: invoices, forms, contracts, statements, IDs.
- You want the AI to handle OCR and field-finding in one step. Extract runs OCR internally; you don’t need a separate parse node in front of it.
- You want a confidence signal per field for downstream review. Enable Field confidence with Engine 1.
- Use parse instead when you only need raw text, or classify when the goal is to sort documents rather than read fields off them.
Configuration
Engines
- Engine 1: Faster, lower cost. Best for simple documents with clear structure.
- Engine 2: More capable. Best for complex documents, handwriting, or ambiguous layouts.
Documents read as text (office and CSV) need Engine 1. They are converted to text
rather than rendered as pages, so Engine 2 has no page image to read and the step fails with a
message saying so. Engine 1 extracts from the converted text. See
supported formats.Because there are no pages to chunk, the whole of that text goes to the model in one pass. A
document that converts to more text than the model can read in one pass is refused with a message
naming its size rather than sent and failed; narrow it with a filter node, or
split the source into smaller files.
Precision
Auto (default) picks the right tier automatically per document. Small, Medium, and High pin a fixed tier, where higher is more accurate but slower and costs more credits.
With Auto the editor cannot show a single per-page cost, so the node displays a 1-6 credits/page range badge. The exact charge is shown after the run.
Chunk strategy
Long documents are split into chunks before extraction. The Chunk Strategy controls how those chunks are processed.- Parallel (default) runs chunks concurrently for the fastest wall-clock time.
- Sequential runs chunks one at a time. It’s slower but can improve accuracy on long documents with tables or array-heavy schemas, where context from earlier chunks helps the model align rows in later ones.
Custom prompt for schema generation
When using Generate to let AI create a schema from a sample document, you can provide a custom prompt to specify which fields to extract. The AI will include only the fields described in your prompt, ignoring other visible data in the document. This is useful when a document contains many fields but you only need a subset.JSON path
Leave JSON path blank to keep the full extraction output. Set it when you want to narrow the result down to a specific section of the extracted JSON. For example, use it to pull a single nested array out and pass only that array to downstream nodes.Schema field types
The schema editor accepts the JSON Schema primitives (string, number, integer, boolean, object, array) plus three semantic types: date, email, and phone. Semantic types serialize to string with a JSON Schema format, so any downstream consumer that reads the schema continues to work.
date fields take an output format on the field itself (ISO 8601, US, EU, or Long). Extract normalizes the value it reads from the document to that format, regardless of how the document writes the date. See date output formats in the schema design guide.
Unverified fields
A field is marked unverified when the matcher could not place its value on the page. The value itself is still extracted and still appears in the step output; what is missing is the highlight, because no region of the document matched the value closely enough to point at. This is independent of the Field confidence toggle: a field can carry a high confidence score and still be unverified.Review advice
Turn on Review advice and each document gets one more question answered after extraction: should a person look at this before the data is used? The check reads what the extraction already knows, every field’s value, description, confidence, whether it was found on the page and whether it is required, plus the page quality and any pages that were not read. It does not read the document text itself. The answer is published asmetadata.reviewAdvice:
Route on it with a condition edge:
{{metadata.reviewAdvice.needsReview}} is true sends the document to a review node, and a clean document goes straight on. Combine it with {{metadata.validationFailures}} is not empty when a validation node sits in between, because the check runs before validation and cannot see its failures.
When the check is unsure or unavailable, the signals decide: any unverified or contested field, any field below 60% confidence, a page below 60% OCR confidence (Engine 2 only, since it is the engine that reads the pages), or a page that was not read asks for a review.
Advanced duplicate-value disambiguation
A collapsible Advanced duplicate-value disambiguation section appears under the extract editor only when both Engine 1 is selected and OCR grounding is enabled. It tunes how the node handles fields that appear to have landed on the same source text.- Duplicate Verifier (toggle): runs a short follow-up step to disambiguate two fields that appear to have landed on the same source text. The step runs on the same built-in check the other nodes use and needs no precision setting.
- Ensemble Verification (toggle): runs the extraction multiple times in parallel and keeps the most consistent result, for when accuracy matters more than throughput.
- Max Fields per Verifier Call (number,
1to8, default4): the maximum number of colliding fields the verifier resolves in a single call. Collisions above this count are skipped.
Bounding box overlay
After an extraction run completes, the document viewer displays colored bounding box overlays on the original document showing where each field value was found:- The overlay color indicates the confidence tier: green for high confidence, amber for medium, red for low.
- Use Show all fields in the viewer toolbar to switch between drawing every field’s box on the page and drawing only the box for the field you have selected.
- Hover over a bounding box to see the exact confidence score.
- Click a bounding box to select the corresponding field in the side panel.
Office and CSV documents have no bounding boxes. They are converted to text rather than
rendered into pages, so there is no page image to draw on and nothing to ground a value to.
Extraction still runs and still reports confidence; the document viewer shows the converted text
instead of an overlay. See supported formats.
Contested bindings
A field is contested when its extracted value is printed more than once in the document and the copy its highlight landed on scored barely ahead of the next one. The value itself is settled: what the marker questions is which printed copy the highlight points at. Contested fields keep their confidence color and are drawn with a dashed outline instead of a solid one. Hover the highlight to see how many times the value appears and how close the next copy scored. The count covers every page searched for that field, not just the page you are on, so it can be higher than the number of outlines drawn. Select the field and the copies it was not bound to are outlined as well, so you can compare them without leaving the viewer. The side panel collects contested fields under Contested bindings, each with the number of times its value appears. Click one to select the field and draw its other occurrences. A contested marker is not a quality score. Confidence tells you how much to trust the value; the dashed outline tells you the position is the part worth a second look.When a step has contested fields, its metadata lists their paths under
metadata.contestedFields.
Route on it with a condition edge to send only those runs to a
review node.Linking schema fields to the document in the builder
When you open the extract node in the editor and select a document, the document preview panel shows the same bounding box overlays you see on the Runs page. Selecting a schema row in the editor draws a connector line from the row to every matching bbox on the document; for array items the line splits so each row in the document lights up. Clicking a bbox highlights the matching schema row and scrolls it into view. A small chip on each leaf row shows the match count, and a warning appears for fields that did not match anywhere in the document. If you edit the schema or prompt after the last run, the preview panel shows a stale-run banner so you know the overlays may not reflect the current configuration. Use test step to refresh them.Inputs and outputs
Allowed inputs: Trigger nodes, filter, split, classify, parse, review, loop. Output: Structured JSON data matching the configured schema.Related
Extract best practices
Engine choice, confidence patterns, and pitfalls
Schema design
How to write JSON Schema for Ingestly extractions
Review action
Pause runs with unverified fields for human review
Validation action
Enforce required fields and formats after extraction