Skip to main content
The classify node sends each document page to a specific output branch based on its content. Use it to sort incoming documents by type, topic, or any other classification criteria before applying downstream processing.

When to use classify

  • You receive mixed document types through one trigger and need to apply different processing to each.
  • Different document types call for different extract schemas. Classifying upstream lets each branch use a focused schema instead of one mega-schema covering everything.
  • You want to drop or quarantine documents that don’t fit any expected type. Use the other output for unmatched pages.
  • Use filter instead when you’re narrowing one document to a subset of its pages, and split when one upload holds several documents that each need their own run. Classify groups consecutive pages of the same type into one document, so two invoices back to back need a split step in front of it to come apart.
For mode selection and choosing sample documents, see classify best practices.

Configuration

Modes

Rules

Define conditions that match page content or page numbers. Each rule gets its own output on the node, named after the rule, and the pages it matches continue from that output. Rules accept Contains text, Not contains text, and Page numbers, and each rule needs at least one of them: pages no rule matches already leave through Other. Renaming a rule keeps its connections; removing it removes them.

By example

Add one example per kind of document you expect, and give each one a sample document: Add Sample, then Pick Documents chooses a document this workflow already processed, and Upload Files takes one file only to read it as the sample (the file is deleted once its text is read, and only the text of its first three pages is kept). AI compares every page with the examples’ samples and sends the page to the example whose sample it matches: the same kind of document, from the same issuer. A Sunfire invoice page goes to the example holding a Sunfire invoice, while a Sunfire purchase order or an Acme invoice does not. Different numbers, dates, amounts and line items do not matter. Every example needs its sample, or the step fails and names the example that has none. The page may be any page of such a document: its first page, a continuation page or its last page. A sample shows AI its first three pages, so the pages that follow a document’s first page still find their target. Word similarity first narrows the samples worth comparing, so a page that shares almost no words with any sample is not compared at all. There is no threshold to tune: AI decides, and a match it is not sure enough of is not taken. Each example gets its own output on the node, named after it. Classify only decides which output a page takes: connect any node to an example’s output to process those pages here, or a Forward to send them to another workflow. Pages that match no sample fall through to the other output so you can handle them separately. When pages keep missing their example or going to the wrong one, replace that example’s sample with a more typical document. When AI is unavailable, Classify falls back to word similarity: a page goes to the example whose sample its words are closest to, when they are close enough.

Inputs and outputs

Allowed inputs: any trigger, plus filter, split, parse, if and switch. A parse step is not required. Classify reads the text captured when the document was uploaded, and for a scanned document with no text it runs OCR once, on demand. That pass is not billed separately and is reused by a downstream extract step. Put a parse step in front only when you want its output for another reason. Outputs: one per rule (Rules) or per example (By example), plus Other for pages nothing matched. Classify groups consecutive pages that took the same output, and each group continues from its output in its own child run, so a branch only ever sees its own pages. A group whose output is not connected is not processed; if no group reaches a connected output, the step fails.

Credits

See credits for pricing. Rules mode costs 0.5 credits per page; By example mode costs 0.2 credits per page, with the AI comparison included.

Classify best practices

Pick a mode and choose a sample document for each example

Conditional routing guide

Patterns for branching workflows on document type

Forward output

Receive the pages Classify forwards

Extract action

Extract data after classifying by document type