How to Extract Data from a Scanned Invoice (PDF or Photo)

Jun 19, 2026

Try it now: convert an invoice to Excel or CSV

Upload a PDF or scanned invoice and get the vendor, dates, totals, and every line item back in seconds. Your first invoice is free.

PDF, JPG, PNG, BMP, HEIC, TIFF

Upload your receipts and invoices

Last updated June 2026.

A scanned invoice is just an image to a computer. You can see the vendor, the totals, and the line items, but your spreadsheet cannot, because there is no selectable text behind the picture. Getting that data out means reading the image, not copying from it. This guide covers how to extract data from a scanned invoice, whether it arrived as a scanned PDF or a phone photo, and how to get it into Excel or CSV without retyping a thing.

Extract a scanned invoice right now

Upload a scanned PDF or a photo of an invoice and get the vendor, dates, totals, and every line item back in Excel or CSV. No typing, no setup.

Convert a scanned invoice now

How do I extract data from a scanned invoice?

Run the scan through AI invoice OCR. The tool reads the image, recognizes the characters as text, works out the layout, and pulls the vendor, invoice number, dates, tax, totals, and line items into structured fields you can export to Excel or CSV. Upload the file, check the extracted values, and download the spreadsheet. The whole pass takes seconds per invoice.

The key difference from a normal PDF is that a scan has no text layer. With a born-digital invoice you can highlight and copy the numbers; with a scan or a photo there is nothing to select, so the software has to recognize each character from the pixels first. A purpose-built invoice tool does both steps, the reading and the field mapping, in one upload, which is why you do not need a separate OCR program and a separate parsing step.

Can you extract data from a scanned PDF invoice?

Yes. A scanned PDF is an image wrapped in a PDF container, so the data is recoverable as long as the scan is legible. AI invoice OCR reads the image inside the PDF, recognizes the text, and maps it to fields. Multi-page scanned PDFs work too, and the line items are pulled from each page, not just the first.

The catch with scanned PDFs is quality. A clean 300 DPI scan reads almost as well as a digital file. A faxed, skewed, or low-resolution scan is harder, because faint or broken characters are easier to misread. If you control the scan settings, scanning at 300 DPI in black and white or grayscale, kept straight and in focus, gives the reader the cleanest input.

What is the best way to extract data from a scanned invoice?

The best way for most teams is a dedicated invoice OCR tool that reads the scan and exports a spreadsheet in one step. It beats manual retyping on speed and accuracy, and it beats general OCR because it knows what an invoice is: it returns labeled fields and line-item rows, not a wall of loose text you still have to organize. For one-off invoices, an upload-and-download tool is fastest.

General-purpose OCR, the kind built into a scanner or a PDF reader, will give you raw text, but you then have to find the vendor, the totals, and each line yourself. That is fine for a paragraph of prose and painful for a table of charges. A tool tuned for invoices skips that cleanup. If you want the full picture of how the reading and field-mapping happen under the hood, see how invoice OCR works.

Why is scanned invoice data harder to extract?

Scanned data is harder because the computer starts from pixels, not text. A digital invoice already contains the characters; a scan or photo hides them inside an image, so the software must recognize every character before it can find a field. Anything that degrades the image, low resolution, skew, shadows, creases, or faint print, raises the chance of a misread.

Phone photos add their own problems: angle, glare, and uneven lighting. Most modern tools correct for mild skew and lighting automatically, but a sharp, flat, well-lit shot still reads more reliably than a crumpled receipt photographed at an angle. The cleaner the input, the less checking you do on the output.

How accurate is data extraction from scanned invoices?

On clean, legible scans, AI invoice OCR commonly reads core fields like vendor, invoice number, dates, and totals at 95 percent accuracy or better. Accuracy slips on poor scans, handwriting, and unusual layouts, which is why a quick human review of the money fields is still worth the few seconds it takes. The tool does the heavy lifting; you confirm the numbers that matter.

Accuracy is usually measured at the field level, meaning whether the whole invoice number or total came out right, which matters more for bookkeeping than character-level scores. For a deeper look at the numbers and what moves them, see how accurate invoice OCR is.

Can you extract line items from a scanned invoice?

Yes. A good invoice tool extracts the full line-item table from a scan, returning one row per line with description, quantity, unit price, and amount, alongside the header fields. This is the part that saves the most time, because line items are the slowest thing to retype and the easiest to fumble by hand.

Line items are also the hardest part of the page to read, since tables vary so much between vendors and a scan can blur column edges. Tools built for invoice line-item extraction handle multi-line descriptions, wrapped rows, and totals that span pages, so the itemized detail comes through instead of collapsing into the header total.

How do I convert a scanned invoice to Excel?

Upload the scanned PDF or photo to an invoice OCR tool, let it read the fields and line items, then export to Excel or CSV. The output is a spreadsheet with columns for vendor, invoice number, dates, tax, totals, and a row for each line item, ready to import into QuickBooks, Xero, or NetSuite. There is no copy and paste between a text dump and a sheet.

If you mostly work in spreadsheets, an invoice PDF to Excel converter is the direct path: the scan goes in, a clean Excel file comes out. The same approach works for other scanned paperwork too. Scanned receipts go through receipt OCR the same way, and mixed scanned documents beyond invoices can be read with a general document OCR tool.

Do I need code or software to extract scanned invoice data?

No. For everyday work you do not need to write code or install anything. A browser-based invoice OCR tool reads the scan and hands you a spreadsheet, which is all most accountants and AP teams need. You only reach for code when you want extraction to run automatically as part of a larger workflow.

If you do automate, you have options that still avoid heavy building. The same job can run through a no-code platform you already use, for example as a step in a Zapier invoice flow or a Power Automate process, or through a direct invoice OCR API when you want to send files and get structured data back on a schedule. Start by uploading manually, then automate only the part that repeats.

How can I improve scanned invoice extraction accuracy?

Feed the tool the cleanest image you can. Scan at 300 DPI, keep the page straight and flat, use good lighting for photos, and avoid shadows, glare, and creases. The clearer the characters, the fewer misreads, and the less time you spend checking the result. Quality in is the single biggest lever on accuracy out.

After extraction, a fast review pays off: glance at the vendor, the invoice number, and the totals, since those are the fields that matter most for your books. Most tools flag low-confidence values so you know where to look. Once you trust the input quality and the review habit, scanned invoices move about as fast as digital ones. To choose a tool for it, compare options in invoice OCR software, or convert your own scan in the tool above and check the columns before you commit.