papero-extract
Extract structured text from PDFs and office documents
TLDR
SYNOPSIS
papero-extract --versionpapero-extract extract file [-o path] [-f format] [-p pages] [--fast] [options]papero-extract batch input -o directory [options]papero-extract serve [--host address] [--port port]
DESCRIPTION
papero-extract rebuilds the structure of a PDF: reading order, tables, formulas, figures, and a bounding box for every block. It is the command-line entry point of papero (the papero-extract Python package). Two engines run on the same file. A layout engine on PDFium reads glyphs, rules, and images and rebuilds columns and tables from geometry. Apache Tika adds metadata, tagged headings, OCR, and non-PDF formats such as DOCX, PPTX, XLSX, EPUB, and HTML.extract handles one file. With no -f, the format follows the -o extension (.md, .txt, .json, .html, .csv) and otherwise is Markdown. csv writes tables only. --images crops figures, tables, and formulas to PNG; when -o is set, those files land in an images/ directory next to the output. --fast skips the layout engine and returns cleaned text from Tika.batch walks a directory (default pattern **\*.pdf) and writes one Markdown and one JSON file per document, a chunks.jsonl file, a manifest, and a fidelity report under fidelity/**. The report compares the extraction to the words PDFium reads on each page. It is a guide to documents worth reviewing, not a score against a ground-truth corpus.serve starts the REST API and the bundled browser app. The API extra must be installed (pip install 'papero-extract[api]'). POST /v1/extract accepts an uploaded file. Interactive docs are at /docs.
PARAMETERS
extract file
Extract one document. Writes to stdout unless -o is set.-o, --output path
Output file. CSV is written as UTF-8 with a BOM. Other formats are UTF-8.-f, --format format
One of markdown, text, json, html, csv. Default: inferred from -o, otherwise markdown.-p, --pages spec
Page selection such as 1-3,5,10-.--password password
Password for an encrypted PDF.--images
Crop figures, tables, and formulas to PNG.--image-scale scale
Crop resolution. The default 2 is 144 dpi.--no-tables
Skip table detection.--no-formulas
Skip formula detection.--ocr auto|off|force
When to OCR. The default auto OCRs pages that look scanned. force OCRs every page. off never does.--ocr-language langs
Tesseract language string. The default is por+eng.--no-tika
Run the PDFium layout engine only, with no Java Tika server.--page-breaks
Mark the start of each page in Markdown.--fast
Text only, via Tika. Fastest path. Layout, tables, and formulas are not rebuilt.--raw
With --fast, skip all text cleanup.--keep-headers
With --fast, keep headers and footers.--dehyphenate
With --fast, join words split by a hyphen at a line break.-w, --workers n
Worker processes. For extract, the default is the CPU count. For serve, workers are off unless this is set (exported as PTE_WORKERS).-q, --quiet
Suppress the stderr summary. Warnings are still printed.batch input -o directory
input is a directory (walked recursively) or a single file. -o is required and is the dataset directory.--pattern glob
Files to include. Default **\*.pdf**.--formats list
Per-document outputs, comma-separated. Default markdown,json.--chunk-size chars
Maximum chunk size in characters. Default 1500. Chunks break on headings and keep tables whole.serve
HTTP API and browser app. Defaults to host 0.0.0.0 and port 8000, or the PORT environment variable when it is set.--host address
Bind address.--port port
Listen port.--version
Print the package version and exit.
CAVEATS
serve fails until the API extra is installed: pip install 'papero-extract[api]'. OCR and non-PDF formats need a working Tika and Tesseract setup. --no-tika drops both and keeps only the geometry engine.Formula reconstruction is glyph-based. Fractions, roots, exponents, and indices become LaTeX. Matrices and aligned systems come out linear. A cropped image of the formula is the reliable copy when --images is on. Borderless tables with very narrow column gaps can be read as paragraphs.Word (.docx) and Excel export exist in the browser app. The CLI formats are Markdown, text, JSON, HTML, and CSV. Scanned pages are not readable with --fast. That path warns when it finds little text.
