1. Pick the right mode for the signal you have
Each mode answers a different question:- Rules answer “does this page contain text X (or sit on page N)?”. Cheap, deterministic, perfect when you can describe each category in one or two phrases that are reliably present.
- By example answers “which example’s sample is this page the same kind of document, from the same issuer, as?”. Billed per page, ideal when each category is one issuer’s document type and easy to point to.
- The pages of one issuer’s document type vary too much for a fixed phrase (layouts shift, fields move, pages continue).
- The category is obvious to a reader but has no phrase every page shares.
- You have a typical document of each category to hand: each example’s sample document is what pages are compared with.
2. Make categories mutually exclusive
A page should match exactly one category. When categories overlap, the same document falls through to whichever rule comes first, and small tweaks to one rule break others.3. Write criteria the way you’d brief a human reviewer
Think of each rule’s criteria as instructions to a colleague: “Look for X, Y, and Z to identify this type of document.” Specific anchors that a reviewer would notice in a few seconds tend to be the same anchors a rule can match cleanly.4. Use enough categories, not too many
Too few categories: a single bucket hides important differences and forces downstream nodes to re-classify. Too many categories: rules collide, and in By example mode near-identical samples on two examples leave AI little to tell them apart by. Practical guideline: define a category whenever the downstream workflow behavior differs. If two document types go through the same extract schema and the same delivery, fold them into one category. Always include an explicit fallback (Rules) or rely on theother output (By example) so you never have to handle “didn’t match any rule” implicitly.
5. Give each example a typical sample document
By example mode has AI compare each page with each example’s sample document, so the sample is the category definition. There is no threshold to tune: a page matches when AI finds it is the same kind of document, from the same issuer, as a sample, and a match AI is not sure enough of goes toother. Each segment’s confidence in the run output is AI’s confidence for its least certain page.
Pick a sample with several pages when the documents run long: AI sees a sample’s first three pages, so continuation and last pages find their example too. When an example keeps missing its pages or attracting pages that belong elsewhere, replace its sample with a more typical document.
6. Know what happens when AI is unavailable
When AI cannot be reached, By example mode falls back to word similarity: a page goes to the example whose sample its words are closest to, when they are close enough. Word similarity cannot tell two issuers’ documents of the same kind apart as well as AI does, so check the routes of a run that happened during an outage. The price per page is the same either way.7. Add a fallback path for the other output
Pages that don’t match any category in By example mode go to the other output. Don’t dead-end this: route it somewhere actionable, for example:
- A review node so a human classifies the page, and you replace a sample that keeps missing.
- A separate workflow for unknown documents so they don’t pollute your main run history.
- An alert via HTTP action so the team knows volume of unknown-type pages is growing.
other output as a feedback signal; it’s where you learn which samples need replacing or which examples are missing.
Common pitfalls
Rules in By example mode (or vice versa)
Rules in By example mode (or vice versa)
Rules and By example are different modes; only one is active at a time. If you find yourself wanting both, run a Rules classify first to handle the easy cases, then a second classify in By example mode on the
other branch.One example for several issuers
One example for several issuers
An example’s sample matches documents from its own issuer only. A second vendor’s invoices need their own example with their own sample (both outputs can lead to the same steps), or a Rules classify when every vendor’s invoices go the same way.
A sample that is not typical
A sample that is not typical
A sample with an unusual first page, or a one-page sample for documents that run several pages, gives AI less to go on. Replace it with a document that looks like the ones the example should take.
Overlapping rule criteria silently routes wrong
Overlapping rule criteria silently routes wrong
Rules evaluate in order; the first match wins. If two rules can match the same page, the page goes to the rule defined first, regardless of which one is “more correct”. Check your category criteria pairwise for overlaps.
other output not connected to anything
other output not connected to anything
Unmatched pages disappear silently if
other is dangling. Always wire it to a review, a fallback workflow, or at minimum an alerting HTTP action.Classifying after extract instead of before
Classifying after extract instead of before
Classifying upstream of extract lets you use a focused schema per type. Classifying downstream means you’ve already paid full extract cost on every document; consider re-ordering the workflow.
Related
Classify action
Configuration reference for the classify node
Conditional routing
Patterns for branching workflows on document type
Forward trigger
Receive the pages Classify forwards
Review action
Send unmatched pages to a human for handling