How StatementHarbor works

No black box: here is the exact pipeline your statement goes through — all of it on your machine, none of it on a server.

The four-step pipeline

1. Read the PDF locally

Your browser reads the PDF bytes straight from your disk using pdf.js, the same open-source engine Firefox uses to display PDFs. The file is held in memory on your device. There is no network request carrying your data — the architecture makes privacy automatic rather than promised.

2. Extract text with positions

PDFs don't store tables — they store text fragments with coordinates. StatementHarbor groups fragments into lines by vertical position and re-orders each line by horizontal position, so column structure survives: dates stay dates, amounts stay amounts.

3. Recognize transactions

Pattern matching identifies dates (supporting MM/DD/YYYY, DD/MM/YYYY, ISO, and "01 Mar 2024" styles), money values (handling commas, parentheses for negatives, and CR/DR markers), and balances. Description rows that wrap onto a second line are merged back together. Headers, page numbers, and balance-summary lines are skipped.

4. Export

Rows export as CSV (opens in Excel, Google Sheets, Numbers) or QBO (QuickBooks Online's Web Connect format). Files are built in memory and saved via your browser's normal download — again, nothing transmitted.

Verify the privacy claim yourself

We'd rather show than tell. Two independent checks:

Known limitations (the honest section)

Why this matters for bookkeepers

If you handle client files, uploading statements to third-party servers can conflict with confidentiality duties and data-protection obligations. A converter that cannot receive your files cannot breach them. That's the standard StatementHarbor was built to.

SH

Reviewed by the StatementHarbor team
Last updated September 2026