Balance sheet, income statement and cash flow, read from the pages of the report as printed and standardized onto one vocabulary — deterministic, reconciled and traceable line by line.
An annual report is a long document; the financial statements are three to six pages of it. PDPilot works on those pages only. When a PDF is chosen for upload, the server proposes the balance-sheet, P&L and cash-flow page sets with a confidence level and thumbnails; the analyst confirms or corrects them, and the extraction runs on exactly the confirmed positions. A statement that continues on the following page is read as one statement.
The document records whether its pages were proposed, edited or entered manually, so the origin of every extraction is part of its provenance.
Deterministic geometry and vocabulary — the same PDF always produces the same result, and every step can be inspected.
One line vocabulary and order for every issuer. Expand any standardized row to see the printed lines exactly as they appear in the report; click the page link to see the spot on the page.
Balance identity, walk-down reconciliation, a guessed year column, an undeclared scale — each is a visible flag on the document, read before the figures are relied on.
Leverage, coverage and cash-flow ratios, an indicative metrics-implied rating, and the multi-year issuer view — all computed live from the same line items the document pages show.
Many filings exist only as images — company-register scans, rasterized PDFs, exports with outlined fonts. For a named statement page without a text layer, PDPilot builds an OCR text layer with word geometry and reads it with the same engine as a text PDF. OCR engages only where it is needed; documents with a text layer are untouched.
Rows read from scanned pages carry an OCR tag with the page's recognition confidence, the quality panel lists the pages concerned, and the document stores the OCR provenance. OCR language packs in the product: German, English, French and Dutch. UK GAAP accounts filed as scans with Companies House are a supported class.
Where a machine-readable filing exists, PDPilot can start from it instead of a PDF. The public ESEF index can be searched by LEI or company name and Companies House by company number or name; a saved filing — an inline-XBRL report, an ESEF package or an xBRL-JSON file — can also be dropped directly.
The filing's tagged facts become the same standardized spread, walk-down, ratios and rating as a PDF extraction. Leaves are summed along the filing's own calculation network, with signs from its weights and sections from its structure; tagged totals validate and are never copied; unmapped facts are excluded and reported. Every row carries an XBRL tag and the quality panel links to a fact-by-fact reconciliation.
A wrong page position is the most common cause of a strange figure — the printed page number and the position in the file usually differ. Beyond that, the analyst has three tools, all recorded:
No. Extraction and derivation are deterministic, verifiable rules — geometry, vocabulary and arithmetic. Every figure is computed, traceable and reproducible rather than generated, which is what an audit or validation function needs.
No. The scale note on the statement page is read automatically. How figures are displayed — millions, thousands or as printed — is a separate choice in the interface.
Because every total is added up from the line items shown, never copied from the printed total. A tiny gap is rounding; a larger one means a line was missed or misplaced — and it is flagged as a finding rather than smoothed over.
The page's position in the PDF file (1 = the first page of the file), not the number printed in the footer. The page proposal on upload pre-fills these positions for you to confirm.
A live demonstration environment is available on request — bring the PDFs your analysts spread today.