one envelope, every document.
a finished parse comes back in the same envelope, whatever went in — so the code that reads a receipt reads a 200-page filing without changing.
hand us a pdf
web upload for people; one claritize call for agents. up to 2,000 pages.
the managed pipeline parses it
layout, tables, six languages — you never see the pipeline, and you never tune it.
typed json comes back
on our engine, every value says how it was checked, such as printed on the page, confirmed by a second reading, worked out from the document's own arithmetic, disputed, not found or not checked. a value that was found also says where, and a typed result says whether it is ready and names what to look at first. a download for you; a short-lived link for your agent.
what each side actually gets.
for people
a folder portal. upload pdfs, watch them parse, download json. move and organize — a file manager, not a developer tool. you never write a line of code.
for agents
two mcp tools: claritize does the work,
claritize_check watches it. async by design —
submit a long document, keep working, harvest the result when it lands.
it arrives as a short-lived link. or ask for one field, one table or one page:
a value with how it was checked, a small table's cells (a larger one points to
its csv), or the values cited on a page.
illustrative arithmetic, tables left out — the json waits in storage; your agent reads it where tokens are free.
receipts. tables. six languages.
receipts
merchant, date, line items, tax broken out by jurisdiction, tip, total, last 4 of the card. restaurant, retail, gas, online — every flavor.
extracted as printed, in the document's own currencytables
tables come back as csv files, listed with their pages and column headers, and with titles and row labels under the default schema.
a warning when the index lists a table the csvs lack, or the reversemulti-language
english, spanish, french, german, italian, and portuguese — native, with mixed-language documents handled per page. need another? more on request. we extract in the source language — translation stays yours.
en · es · fr · de · it · ptand by extension: forms, long documents, bad scans — same call, nothing to configure. an amount we can't read comes back marked unread, and when the document's own arithmetic pins it down, that arithmetic comes with it.
the whole ledger.
the short version is on the way down: we never learn from your documents. here is the long version — what we keep, and for how long.
coming soon.
effort pricing: each document is priced from an estimate of the work, shown free before the parse, not a flat rate per page. the work comes in six kinds: pages with a text layer, scanned pages, hard scans read by a stronger model, extraction passes, second extraction attempts, and checks of a value against the page. the receipt lists the units it counted.