Migrating from pymupdf¶
pylopdf is pymupdf-style, not a drop-in replacement. The data shapes that
determine migration cost — "words" tuples, "dict" structure,
search_for → list[Rect], 1-based TOC pages — match pymupdf, so most
extraction and page-management code ports with small edits. This page lists
what carries over, what changed, and what to use instead of the parts pylopdf
deliberately does not implement.
Note
pylopdf handles PDF files only. pymupdf's ability to open XPS / EPUB / images does not carry over.
Quick mapping¶
| pymupdf | pylopdf | Notes |
|---|---|---|
import fitz / import pymupdf |
import pylopdf |
|
fitz.open(path) / open(stream=…) |
pylopdf.open(path) / open(stream=…) |
same shape; password= too |
doc[i], len(doc), iteration |
same | 0-based, negative indices |
doc.metadata / set_metadata |
same | same key names |
page.get_text() |
same | options: text / words / blocks / dict |
page.search_for(t) |
same | returns list[Rect]; no quads= |
page.get_pixmap(matrix=fitz.Matrix(2, 2)) |
page.get_pixmap(scale=2) |
or dpi=144; no Matrix class |
pix.samples / width / height / stride |
same | always straight-alpha RGBA8; tobytes() → PNG |
page.get_images() / extract |
page.get_images() |
returns drawn images with bbox; JPEG passthrough |
doc.select, delete_page(s), copy_page, new_page |
same | select with repeats duplicates pages |
doc.insert_pdf(src, from_page=, to_page=, start_at=) |
same | |
doc.get_toc() / set_toc() |
same | pages 1-based (both) |
doc.save(garbage=4, deflate=True) |
doc.save(garbage=True, deflate=True, object_streams=True) |
garbage is a bool |
doc.save(encryption=…, user_pw=…) |
doc.save(user_pw=…, owner_pw=…, permissions=…) |
AES-256 only |
doc.needs_pass / authenticate() |
same | same return semantics (0/1/2/4/6) |
page.rect / rotation / set_rotation |
same | |
page.insert_image(rect, filename=) |
same | JPEG/PNG only; no pixmap= — convert via Pillow |
page.show_pdf_page(rect, src, pno) |
same | same-document overlay unsupported (copy first) |
page.insert_text(point, text, fontsize=, fontname=, fontfile=) |
same, plus fontbuffer= / fontindex= |
no font source: standard-14 / WinAnsi; a source is shaped with HarfRust and subset-embedded through krilla |
page.insert_textbox(rect, text, align=, lineheight=) |
same, plus arbitrary fontfile= / fontbuffer= |
UAX #14 CJK wrapping; negative return means no drawing |
page.add_highlight_annot(...) |
same | appearance stream always generated |
doc.embfile_add / names / get / del |
same | |
doc.get_page_labels / set_page_labels, page.get_label |
same | |
page.widgets() / widget objects |
doc.get_form_fields() / doc.set_form_field(name, value) |
document-level; NeedAppearances |
pymupdf4llm.to_markdown(doc) |
doc.to_markdown() |
built in, MIT |
Behavioral differences¶
- Coordinates are top-left-origin display space in both libraries, and in pylopdf this includes rotated pages consistently across extraction, search, drawing and rendering.
- Types:
Rectis an immutableNamedTuple(x0, y0, x1, y1, pluswidth/height). There are noPoint/Matrix/Quadclasses — APIs take plain tuples andscale=/dpi=keywords. - Stale pages: after structural changes (delete / insert / reorder),
previously fetched
Pageobjects raiseStalePageErrorinstead of silently pointing at a different page. Re-fetch withdoc[i]. - Exceptions:
PdfError(aValueErrorsubclass) is the base;PasswordError,DocumentClosedError,EncryptedDocumentError,StalePageErrorrefine it.except ValueErrorkeeps working. get_textoptions are limited totext/words/blocks/dict(nohtml/rawdict/xml). Span dicts carryfontand pymupdf-styleflags(bold/italic/serif/mono) for embedded fonts.- Multicolumn text follows deterministic whitespace gutters, reading top-to-bottom within each column and columns from left to right.
Page.find_tables()reconstructs axis-aligned bordered grids from strokes or thin filled rectangles, including rectangular merged cells. Opt in to high-confidence borderless detection withstrategy="text"; aligned multicolumn prose remains geometrically ambiguous. Useclip=to keep only complete tables in a known display-coordinate region, and inspectTable.confidence/Table.diagnosticswhen ranking borderless results.- Form filling sets values +
NeedAppearances; viewers draw the values. pylopdf's own renderer does not regenerate widget appearances. - Vertical CJK writing is detected conservatively and read top-to-bottom, with columns ordered right-to-left. Ruby, warichu and mixed-orientation typography are not interpreted.
Deliberately not implemented — use the ecosystem¶
| pymupdf feature | pylopdf answer |
|---|---|
Story API / insert_htmlbox (typesetting) |
typst via typst-py — recipe |
OCR (get_textpage_ocr, needs Tesseract installed) |
any OCR engine + insert_ocr_text_layer |
| Digital signatures | pyHanko (MIT) — recipe |
| Incremental save | not planned (qpdf/pikepdf-style rewrite philosophy); pyHanko covers the signature use case |
| Opening XPS / EPUB / CBZ / images | out of scope — PDF only |