PDPilot
Statement extraction

Automated financial statement extraction from PDF annual reports

Balance sheet, income statement and cash flow, read from the pages of the report as printed and standardized onto one vocabulary — deterministic, reconciled and traceable line by line.

PDPilot › Automated financial statement extraction

The input: the statement pages, confirmed by the analyst

An annual report is a long document; the financial statements are three to six pages of it. PDPilot works on those pages only. When a PDF is chosen for upload, the server proposes the balance-sheet, P&L and cash-flow page sets with a confidence level and thumbnails; the analyst confirms or corrects them, and the extraction runs on exactly the confirmed positions. A statement that continues on the following page is read as one statement.

The document records whether its pages were proposed, edited or entered manually, so the origin of every extraction is part of its provenance.

How a statement page is read

Deterministic geometry and vocabulary — the same PDF always produces the same result, and every step can be inspected.

  1. Table geometry. The label column, the value columns and their alignment are resolved from the page's word positions. Side-by-side presentations (assets beside liabilities, a pre-column of sub-amounts) are split into their parts.
  2. Year columns from the printed header. Which column holds the target year and which the comparative is read from the header as printed — never assumed from position.
  3. Unit scale from the page. "In millions", "TEUR", "in thousands of CHF" and their equivalents are read from the statement itself. The display unit is chosen separately.
  4. Number parsing. Thousands separators, decimal commas, bracketed negatives, dashes for zero and the sign conventions of expense lines are normalized.
  5. Classification onto one vocabulary. Each printed line is mapped to a standardized balance-sheet or P&L position by a word-bounded vocabulary that covers the eleven supported languages. A line on a balance-sheet page can only become a balance-sheet position, a line on a P&L page only a P&L position.
  6. Assembly and reconciliation. Sections, subtotals and totals are built from the mapped leaves. Total assets must equal equity plus liabilities; the P&L walk-down from revenue through EBITDA, EBIT and pre-tax profit must foot to net income. Printed totals validate the sums — they are never substituted for them.

What comes out

Standardized statements

One line vocabulary and order for every issuer. Expand any standardized row to see the printed lines exactly as they appear in the report; click the page link to see the spot on the page.

Quality signals

Balance identity, walk-down reconciliation, a guessed year column, an undeclared scale — each is a visible flag on the document, read before the figures are relied on.

Ratios, rating, series

Leverage, coverage and cash-flow ratios, an indicative metrics-implied rating, and the multi-year issuer view — all computed live from the same line items the document pages show.

Scanned reports and register filings

Many filings exist only as images — company-register scans, rasterized PDFs, exports with outlined fonts. For a named statement page without a text layer, PDPilot builds an OCR text layer with word geometry and reads it with the same engine as a text PDF. OCR engages only where it is needed; documents with a text layer are untouched.

Rows read from scanned pages carry an OCR tag with the page's recognition confidence, the quality panel lists the pages concerned, and the document stores the OCR provenance. OCR language packs in the product: German, English, French and Dutch. UK GAAP accounts filed as scans with Companies House are a supported class.

Structured filings: ESEF, inline XBRL, Companies House

Where a machine-readable filing exists, PDPilot can start from it instead of a PDF. The public ESEF index can be searched by LEI or company name and Companies House by company number or name; a saved filing — an inline-XBRL report, an ESEF package or an xBRL-JSON file — can also be dropped directly.

The filing's tagged facts become the same standardized spread, walk-down, ratios and rating as a PDF extraction. Leaves are summed along the filing's own calculation network, with signs from its weights and sections from its structure; tagged totals validate and are never copied; unmapped facts are excluded and reported. Every row carries an XBRL tag and the quality panel links to a fact-by-fact reconciliation.

When something looks wrong

A wrong page position is the most common cause of a strange figure — the printed page number and the position in the file usually differ. Beyond that, the analyst has three tools, all recorded:

  • Re-map or correct a line. Drag a printed line to the right standardized row, correct a value or add a missing line. The engine's reading is kept and restorable; the edit is marked permanently.
  • Guide the geometry. Show the engine where the year columns sit, which region holds the statement and which pages belong to it, then re-extract. No number is typed — every figure still comes from the document.
  • Re-run. Corrections are re-applied wherever the freshly extracted lines still match; anything that no longer fits is listed, never silently dropped.

Common questions about extraction

Is a language model involved in reading the numbers?

No. Extraction and derivation are deterministic, verifiable rules — geometry, vocabulary and arithmetic. Every figure is computed, traceable and reproducible rather than generated, which is what an audit or validation function needs.

Do I have to tell it "figures in thousands"?

No. The scale note on the statement page is read automatically. How figures are displayed — millions, thousands or as printed — is a separate choice in the interface.

Why does a total differ from the one printed in the report?

Because every total is added up from the line items shown, never copied from the printed total. A tiny gap is rounding; a larger one means a line was missed or misplaced — and it is flagged as a finding rather than smoothed over.

Which page numbers are entered?

The page's position in the PDF file (1 = the first page of the file), not the number printed in the footer. The page proposal on upload pre-fills these positions for you to confirm.

See it on your own annual reports

A live demonstration environment is available on request — bring the PDFs your analysts spread today.