claritize.ai · by blockhouse ventures

documents in.
data out.

your documents and data, structured and organized — for you and your agents. you never see the pipeline.

open the portal coming soon contact us →
never trained on
scroll
how it works

a pdf goes in.
typed json comes out.

01

hand us a pdf

web upload for people; one mcp call for agents.

02

the managed pipeline parses it

layout, tables, six languages — you never see the pipeline.

03

typed json comes back

a download for you; a short-lived link for your agent, so its context stays clean.

two audiences, both first-class

one portal for people.
one mcp server for agents.

a folder portal — a file manager, not a developer tool. and two mcp tools: claritize does the work, claritize_check watches it. async by design — the result arrives as a link, and your agent can ask for one field, one table or one page.

trust

we never learn
from your documents.

claritize processes your documents to return structured data. we don't train on your content and we don't sell it. our infrastructure provider processes it for us and keeps no copy.

training = none, ever
sold = never
encryption = in transit, and at rest
your documents teach us nothing

documents in. data out.
nothing learned.

open the portal coming soon contact us →
effort pricing coming soon, no surprise bills · claritize.ai · © 2026 blockhouse ventures
the shape of a result

one envelope, every document.

a finished parse comes back in the same envelope, whatever went in — so the code that reads a receipt reads a 200-page filing without changing.

01

hand us a pdf

web upload for people; one claritize call for agents. up to 2,000 pages.

02

the managed pipeline parses it

layout, tables, six languages — you never see the pipeline, and you never tune it.

03

typed json comes back

on our engine, every value says how it was checked, such as printed on the page, confirmed by a second reading, worked out from the document's own arithmetic, disputed, not found or not checked. a value that was found also says where, and a typed result says whether it is ready and names what to look at first. a download for you; a short-lived link for your agent.

lattice-desk-invoice.pdf parsed

total due="$147.96" · printed, p.1
invoice no.="SUB-30991" · printed, p.1
tables=1 csv · 6 rows · headers + row labels
checked=20 of 20 values printed on the page
review=empty · nothing objected
ready=yes · 12 of 12 requirements proven
an illustration — a real parse of a fictional sample invoice
two audiences, both first-class

what each side actually gets.

for people

a folder portal. upload pdfs, watch them parse, download json. move and organize — a file manager, not a developer tool. you never write a line of code.

for agents

two mcp tools: claritize does the work, claritize_check watches it. async by design — submit a long document, keep working, harvest the result when it lands. it arrives as a short-lived link. or ask for one field, one table or one page: a value with how it was checked, a small table's cells (a larger one points to its csv), or the values cited on a page.

a 200-page contract in. a short answer out.
a 200-page contract, read inline≈ 80,000 tokens· per question
the same contract via claritize≈ 500 tokens· a link and a summary

illustrative arithmetic, tables left out — the json waits in storage; your agent reads it where tokens are free.

what it's exceptionally good at

receipts. tables. six languages.

receipts

merchant, date, line items, tax broken out by jurisdiction, tip, total, last 4 of the card. restaurant, retail, gas, online — every flavor.

extracted as printed, in the document's own currency

tables

tables come back as csv files, listed with their pages and column headers, and with titles and row labels under the default schema.

a warning when the index lists a table the csvs lack, or the reverse

multi-language

english, spanish, french, german, italian, and portuguese — native, with mixed-language documents handled per page. need another? more on request. we extract in the source language — translation stays yours.

en · es · fr · de · it · pt

and by extension: forms, long documents, bad scans — same call, nothing to configure. an amount we can't read comes back marked unread, and when the document's own arithmetic pins it down, that arithmetic comes with it.

trust

the whole ledger.

the short version is on the way down: we never learn from your documents. here is the long version — what we keep, and for how long.

training=none · your content never trains a model, ours or anyone's
selling=never · not sold or brokered. our infrastructure provider processes it for us, keeps no copy and doesn't train on it
encryption=tls in transit · encrypted at rest with aws kms
records=locations, not pages · our database stores where your files are, not the pages themselves, and we never log a request body — though an error can quote a fragment of what failed
storage=your file and what we made from it · the result and its table csvs, so you can fetch them again, and the page images and text the engine worked from
retention=14 days on free · counted from when a parse completes. on any plan, files expire a year after they're written, and a deleted file's last copy expires 7 days later
pricing

coming soon.

effort pricing: each document is priced from an estimate of the work, shown free before the parse, not a flat rate per page. the work comes in six kinds: pages with a text layer, scanned pages, hard scans read by a stronger model, extraction passes, second extraction attempts, and checks of a value against the page. the receipt lists the units it counted.