PDFMaple PDFMaple

Extract PDF tables to Excel: prepare and check your file

Reviewed for workflow clarityUpdated:
Editorial illustration of PDF to Excel: extract data from PDFs (best practices)

Start with a PDF that contains selectable text and clearly separated columns. Convert a small representative section first, then compare the workbook with the original. A spreadsheet that opens successfully can still contain missing cells or shifted values.

Already have an XLSX file? Follow the post-conversion cleanup checklist instead.

Choose the best source

If the original Excel or CSV file is available, use it: a PDF usually contains displayed values, not the original formulas or data types. For PDF-only records, try selecting a number and copying it. An image-only page needs text recognition before its contents can become cells.

PDFMaple attempts an OCR pass when the sampled document text is sparse. This is not a guarantee that every scanned page in a mixed document will be recognized. A clean, upright scan with legible characters is a better starting point than a compressed photograph.

Extract the relevant pages

  1. Keep the original PDF unchanged. Identify the table pages, including headings and footnotes that explain units.
  2. For a long document, use Extract Pages to make a smaller working copy. Retain enough context to understand each column.
  3. Open PDF to Excel, choose the file and download the workbook.
  4. Compare the first table, a page with a different layout, and the final table before converting a larger batch.

What the workbook contains

The current converter creates a worksheet for each PDF page. It detects tables where possible and approximates their layout. Nearby text may appear alongside the table; a heading can become a merged cell. This is an extraction of the visible page, not a reconstruction of the source workbook.

A report that spans three pages can therefore require joining three worksheets. Before joining them, check that the column order and units agree. Remove repeated headings only from the working copy; do not discard genuine data rows that happen to resemble headings.

Check extraction before editing

  • Coverage: count source rows and find the first and last record on each page.
  • Alignment: compare a row containing a blank cell. A blank amount must not cause the next value to shift left.
  • Signs and units: inspect a negative amount, a percentage and any currency symbol.
  • Wrapped labels: confirm that a description continued on a second line still belongs to the same record.
  • Footnotes: check whether a marker belongs to a number or describes the whole table.

When another source is necessary

A chart image does not contain its underlying spreadsheet values. Merged headings, rotated text and faint scans may also produce incomplete structure. If a small sample loses important values, obtain the original data or correct the extraction against the source before relying on it. Do not treat an empty cell as zero.

For confidential records, check your organization's upload rules and PDFMaple's file-handling policy before sending a document to the server.

Turn the extraction into usable data

Next, check numeric types, identifiers and totals in Excel. Keep a raw worksheet and a separate cleaned copy so every correction can be traced to the PDF.