Skip to content
stmtai

Explainer

How OCR reads a bank statement, and where it fails

What happens when software reads a PDF or scanned bank statement: template OCR, generic OCR and vision models, and the inputs that trip each one up.

7 min read · Last reviewed · by the stmtai team

A PDF bank statement is either text or a picture. If the bank generated it, the characters are stored in the file and software can read them directly; the hard part is working out which characters belong to which column. If it was scanned or photographed, the file is only pixels, and something has to turn the pixels into characters before anyone can think about columns. OCR, optical character recognition, is that something.

Whether the result is usable depends more on the input than on the software. This guide explains the three ways software reads a statement, what each one gets wrong, and how to check the output so a misread does not end up in your books.

First, which kind of PDF do you have

Open the statement and try to select a line of text with the mouse. If the words highlight, it is a text PDF. If the cursor draws a box across the page, it is an image, and there is nothing to select. Scanned statements, faxed statements and photos saved as PDF are all images. A few banks also emit "text" PDFs where the characters are drawn as shapes rather than stored as letters; those behave like images too.

Text PDFs still need parsing. The text is stored as positioned fragments, not as rows, so the software has to rebuild each row from the coordinates of the fragments. That is where a wrapped description becomes a second row with no amount, or a "balance brought forward" line gets treated as a transaction.

Three approaches

Template OCR

The oldest approach. Someone measures a particular bank's statement and writes down where each column starts and ends, what the date format is and where the page footer begins. Given that bank's statement, it is very accurate, because it is not really recognising anything; it is cutting the page into known boxes.

It fails completely, not gradually, on anything the template did not anticipate: the bank redesigns its statement, the first page carries a summary block that pushes the table down, the final page has three transactions and a page of disclaimers, or the statement is from a bank nobody wrote a template for. Tools that advertise a list of "supported banks" usually work this way.

Generic OCR

Engines like Tesseract and the cloud document services from the large providers read characters wherever they appear, then a separate step detects tables and groups the characters into cells. This copes with any layout, but it does not understand what it is reading. A "1" and an "l" are just two shapes that look alike.

The characteristic errors are substitutions and losses: 0 read as O, 1 as l or I, 5 as S, 8 as B, a decimal point read as a comma or dropped altogether, and a minus sign lost because it sits outside the cell the table detector drew. The Tesseract project's own guidance is that it works best on images of at least 300 dots per inch and that accuracy falls off quickly on small type, which describes a lot of bank statements printed at eight or nine points. Generic OCR also has no idea that the last column should be a running total, so an amount misread by one digit passes straight through.

Vision language models

The newer approach shows the whole page image to a model that has been trained on images and text together and asks it for structured output: date, description, money in, money out, balance. Because the model reads in context, it knows a column headed "Withdrawals" holds debits, that "1O5.00" is 105.00 because letters do not appear in amounts, and that a line of text with no amount is the continuation of the description above it. It handles layout variety and moderate skew far better than either approach above.

Its failure is different in kind. A vision model can be confidently wrong. Where generic OCR produces a visibly broken token like "1O5.OO", a vision model may produce a clean, plausible 105.00 when the page actually says 165.00, and it may occasionally skip a row on a dense page without leaving a gap. It is also slower and costs more to run per page.

That failure mode is why arithmetic checking matters more with vision models, not less. The printed opening balance, plus every credit, minus every debit, has to equal the printed closing balance, and the running balance has to hold at every row in between. A model cannot fudge that. stmtai reads statements with a vision model and then re-adds every row against the printed balances, flagging the first row where the chain breaks; the sample result shows a flagged row so you can see what a caught misread looks like.

Inputs that defeat all three

InputWhat goes wrongWhat helps
Skewed or curled scanRows drift across column boundaries; a figure near the edge of a column lands in the neighbouring oneRe-scan flat on a scanner, not a phone; use the scanner's deskew option
Low-contrast thermal paper or a faded faxStrokes break up; 8 reads as 3, 0 as C, whole rows vanishScan in greyscale rather than black and white, at 300 dpi or more, before the print fades further
HandwritingAnnotations are read as data; handwritten figures are unreliable in every engineCover notes before scanning; type handwritten figures yourself
Merged columnsDebit and credit columns printed close together, so amounts fall into one column or swapPrefer the bank's own PDF over a scan of a printout; check that debits and credits each sum to the statement totals
Multi-line descriptionsContinuation lines become new rows with no amount, or attach to the wrong rowLook for rows with no amount, and rows with two dates
Negative amounts written as (100.00), 100.00- or 100.00 CRThe sign is lostCompare money in and money out totals to the statement summary
Repeated page furnitureHeaders, "balance brought forward" and "carried forward" lines get counted as transactionsRemove duplicate rows sharing date, amount and description
Phone photographPerspective distortion, shadow across the page, glare, moiré on printed shadingUse a scanner; if you must use a phone, lay the page flat under even light and shoot straight on

Two of these deserve a word more. Skew is the quiet one. A scan rotated by two degrees still looks fine to a person, but across a wide table the bottom rows have shifted several millimetres relative to the top, which is enough to push a balance into the credit column. Thermal receipts and faxes are the hopeless one: once the contrast is gone, no amount of software brings the strokes back, and the honest answer is to get a fresh copy from the bank.

How to check the output

None of the three approaches is reliable enough to skip this. Four checks, in order, catch nearly everything.

  1. Count the rows against the statement. Each page of a statement usually shows how many transactions it holds, or you can count them.
  2. Sum the debits and sum the credits, and compare each with the totals the statement prints in its summary. A missing minus sign or a swapped column shows up here immediately.
  3. Recompute the running balance: opening balance, plus credits, minus debits, row by row, down to the closing balance. The first row where your running figure and the printed balance part company is the row with the error.
  4. Spot-check the largest amounts and anything with the wrong number of decimal places.

This takes a few minutes for a typical statement. For a fair comparison of that time against typing the rows yourself, see converting versus typing it yourself.

Getting a better input in the first place

  • Download the statement from the bank's website rather than scanning a printed copy. The bank's PDF has a text layer and clean column boundaries. The Chase and Wells Fargo guides show where the downloads are hidden.
  • If you must scan, use 300 dpi, greyscale, and the flatbed rather than the sheet feeder for anything creased.
  • Keep the whole statement, including the summary page. The printed totals are what you check against.
  • Do not photograph a statement unless there is no alternative, and never photograph a screen.

If the statement you have is already a clean text PDF, the OCR question does not arise, and the remaining risk is the column parsing described at the top. The checks are the same either way. For what happens after extraction, when the rows go into a spreadsheet, see converting a PDF bank statement to Excel.

Questions

My scanner produces a "searchable PDF". Is that good enough?

The scanner software has run its own OCR and hidden a text layer behind the image. That layer carries whatever errors the scanner engine made, and a converter that trusts it inherits them. It is usually better to hand over the image so it is read fresh, and in either case to check the totals before you import.

Why does the same tool read one bank well and another badly?

Layout. A statement with a single signed amount column is easier than one with separate debit and credit columns printed close together, and short descriptions are easier than ones that wrap onto two or three lines. Some banks also generate PDFs where the text is drawn as outlines rather than stored as characters, which forces a text PDF to be treated as an image.

Can OCR handle statements in other languages or number formats?

Recognising the characters is rarely the problem. Interpreting them is: 1.250,00 and 1,250.00 are different numbers in different countries, and 03/04 is either March 4th or 3rd April. After converting a statement from another country, check one date you know and one amount over a thousand to confirm the format was read the right way round.

Does scanning at a higher resolution always help?

Up to about 300 dpi, yes, and that is what most OCR guidance recommends. Beyond 400 to 600 dpi there is little further gain for printed statements and the files get large and slow. Past that point contrast, flatness and even lighting matter more than resolution.

Is my statement kept by the converter after it has been read?

That depends on the tool, and it is worth reading the privacy policy rather than assuming. Some services retain uploaded documents, sometimes to train their models. stmtai deletes the file when processing ends, and its privacy page says so in plain terms.

Related guides

All guides · All banks and formats