Skip to main content
The classify node decides which downstream branch each page belongs to. Reliable routing depends on choosing the right mode (Rules vs By example), keeping categories cleanly separated, and giving each example a typical sample document.

1. Pick the right mode for the signal you have

Each mode answers a different question:
  • Rules answer “does this page contain text X (or sit on page N)?”. Cheap, deterministic, perfect when you can describe each category in one or two phrases that are reliably present.
  • By example answers “which example’s sample is this page the same kind of document, from the same issuer, as?”. Billed per page, ideal when each category is one issuer’s document type and easy to point to.
Reach for Rules first. Move to By example when:
  • The pages of one issuer’s document type vary too much for a fixed phrase (layouts shift, fields move, pages continue).
  • The category is obvious to a reader but has no phrase every page shares.
  • You have a typical document of each category to hand: each example’s sample document is what pages are compared with.
By example matches the issuer as well as the kind of document, so one example takes one issuer’s documents of one kind. Invoices from many vendors that all go the same way are a better fit for Rules.

2. Make categories mutually exclusive

A page should match exactly one category. When categories overlap, the same document falls through to whichever rule comes first, and small tweaks to one rule break others.
If two categories genuinely look similar, write the rule for the harder one first or add a disambiguating phrase to both.

3. Write criteria the way you’d brief a human reviewer

Think of each rule’s criteria as instructions to a colleague: “Look for X, Y, and Z to identify this type of document.” Specific anchors that a reviewer would notice in a few seconds tend to be the same anchors a rule can match cleanly.
The first version uses headings the document is required to print. The second matches half the documents in a typical pile.

4. Use enough categories, not too many

Too few categories: a single bucket hides important differences and forces downstream nodes to re-classify. Too many categories: rules collide, and in By example mode near-identical samples on two examples leave AI little to tell them apart by. Practical guideline: define a category whenever the downstream workflow behavior differs. If two document types go through the same extract schema and the same delivery, fold them into one category. Always include an explicit fallback (Rules) or rely on the other output (By example) so you never have to handle “didn’t match any rule” implicitly.

5. Give each example a typical sample document

By example mode has AI compare each page with each example’s sample document, so the sample is the category definition. There is no threshold to tune: a page matches when AI finds it is the same kind of document, from the same issuer, as a sample, and a match AI is not sure enough of goes to other. Each segment’s confidence in the run output is AI’s confidence for its least certain page. Pick a sample with several pages when the documents run long: AI sees a sample’s first three pages, so continuation and last pages find their example too. When an example keeps missing its pages or attracting pages that belong elsewhere, replace its sample with a more typical document.

6. Know what happens when AI is unavailable

When AI cannot be reached, By example mode falls back to word similarity: a page goes to the example whose sample its words are closest to, when they are close enough. Word similarity cannot tell two issuers’ documents of the same kind apart as well as AI does, so check the routes of a run that happened during an outage. The price per page is the same either way.

7. Add a fallback path for the other output

Pages that don’t match any category in By example mode go to the other output. Don’t dead-end this: route it somewhere actionable, for example:
  • A review node so a human classifies the page, and you replace a sample that keeps missing.
  • A separate workflow for unknown documents so they don’t pollute your main run history.
  • An alert via HTTP action so the team knows volume of unknown-type pages is growing.
Treat the other output as a feedback signal; it’s where you learn which samples need replacing or which examples are missing.

Common pitfalls

Rules and By example are different modes; only one is active at a time. If you find yourself wanting both, run a Rules classify first to handle the easy cases, then a second classify in By example mode on the other branch.
An example’s sample matches documents from its own issuer only. A second vendor’s invoices need their own example with their own sample (both outputs can lead to the same steps), or a Rules classify when every vendor’s invoices go the same way.
A sample with an unusual first page, or a one-page sample for documents that run several pages, gives AI less to go on. Replace it with a document that looks like the ones the example should take.
Rules evaluate in order; the first match wins. If two rules can match the same page, the page goes to the rule defined first, regardless of which one is “more correct”. Check your category criteria pairwise for overlaps.
Unmatched pages disappear silently if other is dangling. Always wire it to a review, a fallback workflow, or at minimum an alerting HTTP action.
Classifying upstream of extract lets you use a focused schema per type. Classifying downstream means you’ve already paid full extract cost on every document; consider re-ordering the workflow.

Classify action

Configuration reference for the classify node

Conditional routing

Patterns for branching workflows on document type

Forward trigger

Receive the pages Classify forwards

Review action

Send unmatched pages to a human for handling