Solution

Extract data from PDFs and scans – photos and handwriting included

Digital PDF, scan or phone photo: Tiro reads header fields and line items from every variant into the same table – with an example from three files.

By Christian Gerloff · Updated August 28, 2026

In short: Tiro reads header fields and line items from PDFs – digitally created, scanned or photographed with a phone – and returns them as a table. Text layer, image recognition and AI work in one step; you choose neither a document type nor an OCR mode. Below you see the same purchase order in three variants and the values that come out of each.

How Tiro reads a PDF

A PDF can contain two things: a text layer (in digitally created files – exact, every digit is a character) and a page image (scans and photos have nothing else). Classic tools force a choice: text extraction for one, OCR for the other – and fail when a file mixes both or the text layer is broken.

Tiro evaluates both together. Every page goes to the AI model as text and as image: the text makes numbers, IBANs and item numbers precise, the image supplies layout, table structure, stamps and handwriting. Out of that come the fields you defined in the inbox – as real numbers and ISO dates, not text snippets.

What people typically pull out of PDFs and scans

  • Purchase orders and order confirmations (the example below) – header data and line items into the order list.
  • Delivery notes and weigh tickets from goods receipt, often a scan or a photo taken in the yard.
  • Invoices and credit notes from suppliers in every layout.
  • Certificates, test reports, data sheets – measurements and figures from multi-page PDFs.
  • Forms and lists from line-of-business systems that only exist as PDF export.

Why not OCR or a PDF converter?

OCR programs deliver text, not data – working out “which number is the quantity?” stays with you. PDF-to-Excel converters copy the layout into cells and break every table differently. Parseur and Airparser solve this with fields instead of layout – Parseur, however, adds OCR as a separate stage (with its own price) and returns dates as text in the source format. Tiro has one processing path for every variant, no surcharge for scans, and dates and numbers arrive normalised.

Example: from document to table

A fictional document, one inbox, one run in Tiro — shown unchanged. Company, people and figures are invented.

1. The documents (3 variants)

Purchase order PO-2026-04471 of the fictional Corvid Outdoor Supply Co. as a digitally created PDF with six line items
Digital PDF · Open the example PDF
The same purchase order as a slightly skewed greyscale scan without a text layer
Scan · Open the example PDF
The same purchase order as a phone photo on a desk, with a handwritten quantity correction on position 3, a margin note and a signature
Phone photo with handwriting · Open the example PDF

2. The extracted data

Header fields once per document, the table as one row per line item — dates as dates, numbers as numbers. The first rows are expanded, the rest follows behind.

Header fields
FieldDigital PDFScanPhone photo with handwriting
PO numberPO-2026-04471PO-2026-04471PO-2026-04471
Order dateSep 2, 2026Sep 2, 2026Sep 2, 2026
Requested deliverySep 30, 2026Sep 30, 2026Sep 30, 2026
BuyerCorvid Outdoor Supply Co.Corvid Outdoor Supply Co.Corvid Outdoor Supply Co.
SupplierHalden Fabrication LtdHalden Fabrication LtdHalden Fabrication Ltd
Ship toCorvid DC North — Dock 4 1900 Ravensbrook Way Ravensbrook, OR 97015 Attn: Receiving, M. OkonkwoCorvid DC North — Dock 4 1900 Ravensbrook Way Ravensbrook, OR 97015 Attn: Receiving, M. OkonkwoCorvid DC North — Dock 4 1900 Ravensbrook Way Ravensbrook, OR 97015 Attn: Receiving, M. Okonkwo
Payment termsNet 30 daysNet 30 daysNet 30 days
CurrencyUSDUSDUSD
Total7,1057,1057,105
Handwritten notesPos. 3: qty changed 8 → 12, confirmed by phone w/ Halden 03 Sep. Please ship pos. 2 first!
Line items (18)
SourcePosItem no.DescriptionQtyUnitUnit priceAmount
Digital PDF1HF-3120-SSteel tent stake, galvanized, 30 cm40bundle38.51,540
Digital PDF2HF-4410Aluminum ridge pole, 3-section, 2.4 m120pc21.92,628
Digital PDF3HF-5077-BFolding camp table frame, powder-coated black8pc64512
Digital PDF4HF-2205Guy line tensioner, cast aluminum25bag441,100
Digital PDF5HF-6001-LLantern hook, 3-arm, stainless300pc3.15945
Digital PDF6SRV-FRTFreight & crating, DAP Ravensbrook1lot380380
Scan1HF-3120-SSteel tent stake, galvanized, 30 cm40bundle38.51,540
Scan2HF-4410Aluminum ridge pole, 3-section, 2.4 m120pc21.92,628
Scan3HF-5077-BFolding camp table frame, powder-coated black8pc64512
Scan4HF-2205Guy line tensioner, cast aluminum25bag441,100
Scan5HF-6001-LLantern hook, 3-arm, stainless300pc3.15945
Scan6SRV-FRTFreight & crating, DAP Ravensbrook1lot380380
Phone photo with handwriting1HF-3120-SSteel tent stake, galvanized, 30 cm40bundle38.51,540
Phone photo with handwriting2HF-4410Aluminum ridge pole, 3-section, 2.4 m120pc21.92,628
Phone photo with handwriting3HF-5077-BFolding camp table frame, powder-coated black12pc64512
Phone photo with handwriting4HF-2205Guy line tensioner, cast aluminum25bag441,100
Phone photo with handwriting5HF-6001-LLantern hook, 3-arm, stainless300pc3.15945
Phone photo with handwriting6SRV-FRTFreight & crating, DAP Ravensbrook1lot380380
Show all 18 rowsShow fewer rows

3. What it looks like in Tiro

The result in the browser, right after the upload. From here it is one click to Excel or CSV.

The “Purchase orders” inbox in Tiro with the three processed files – digital PDF, scan and photo – all with status “done”

How it works

  1. Create an inbox

    A name is enough, for example “Purchase orders”. On the first upload the AI proposes the fields – keep what you need and add instructions such as “if a quantity was corrected by hand, use the handwritten value”.

  2. Upload the files

    PDF, scan or photo as PDF – one by one or as a batch. You never say which variant it is; Tiro always reads text layer and page image together.

  3. Download the table

    Every document yields the same column structure – one row per line item. Excel or CSV, dates as dates, amounts as numbers; the tables of all files stack directly underneath each other.

Frequently asked questions

Does Tiro need an OCR setting for scanned PDFs?

No. Every page is evaluated as text and as image – for digital PDFs the text layer delivers exact digits, for scans and photos the image takes over. There is no switch and no document type to pick beforehand.

Does it work with phone photos of documents?

Yes, as long as the text is legible. Perspective, shadows and lighting like in the example above do not get in the way. The photo has to be uploaded as a PDF – most scanning apps on the phone do that directly.

Are handwritten corrections and notes recognised?

Yes, as far as the AI can read the handwriting. In the example the crossed-out 8 on position 3 was replaced by the handwritten 12 and the margin note was transcribed in full – because the field instructions say so. What counts is up to you in the inbox.

What is the difference to an OCR program or a PDF converter?

OCR produces text, a converter copies the layout into cells – in both cases you still have to find and assign the values yourself. Tiro returns the fields you defined, in a fixed column structure, with normalised numbers and dates.

Where are the limits?

Password-protected PDFs must be unlocked first, files are limited to 10 MB, and illegible spots – faded thermal print, blurred photos – stay illegible. Such fields come back empty, not invented; every document has a visible status.

Where are the files processed?

On servers in Germany; AI processing runs through AWS Bedrock in the EU region Frankfurt, with no training on your data. Retention is set per inbox.