Python document conversion engine
anyconvert converts PDF documents into DOCX, PPTX, ODT, ODP, and TXT using only the Python standard library.
--mode canvas): 1:1 pixel-accurate positioning. Renders exact coordinates, font families, embedded images, and vector artwork (native DrawingML custom geometries in DOCX, SVG paths in ODT). Ideal for flyers, forms, annotated PDFs, and presentation slides.--mode flow): Semantic reflowable document reconstruction. Automatically detects headings (H1–H6), bullet/numbered lists, tables, and reflowable paragraphs./Type /XRef), and object streams (/Type /ObjStm).FlateDecode with PNG/TIFF predictors, LZWDecode, ASCII85Decode, ASCIIHexDecode, RunLengthDecode)./Differences encodings..docx) and PresentationML (.pptx) packages with complete [Content_Types].xml and .rels hierarchies..odt) and Presentation (.odp) archives with uncompressed mimetype headers at byte offset 0.| Target Format | Extension | Flow Mode | Canvas Mode | Container Standard |
|---|---|---|---|---|
| Microsoft Word | .docx | Semantic <w:p>, <w:tbl> | DrawingML <wps:wsp>, <w:framePr> | ISO/IEC 29500-2 (OPC) |
| Microsoft PowerPoint | .pptx | Reflowable slide text | Exact coordinate shapes & text | ISO/IEC 29500-2 (OPC) |
| OpenDocument Text | .odt | Semantic <text:p>, <table:table> | Positioned <draw:frame>, <draw:path> | ISO/IEC 26300 (ODF) |
| OpenDocument Presentation | .odp | Slide content boxes | Coordinate <draw:frame> | ISO/IEC 26300 (ODF) |
| Plaintext / Markdown | .txt | Formatted Markdown text | 2D character-grid layout | UTF-8 Plaintext |
Install from PyPI:
pip install anyconvert
Or install in editable development mode:
git clone https://github.com/word-sys/anyconvert.git
cd anyconvert
pip install -e .
anyconvert provides a standalone command-line tool:
anyconvert [OPTIONS] INPUT_PDF
| Flag | Argument | Description | Default |
|---|---|---|---|
-f, --format | docx|pptx|odt|odp|txt | Target document format | docx |
-o, --output | PATH | Output destination file path | {input_basename}.{format} |
-m, --mode | flow|canvas | Layout reconstruction mode (canvas for 1:1 visual match) | flow |
-p, --password | STRING | Decryption password for encrypted PDFs | "" |
-v, --verbose | Enable diagnostic logging | False | |
--version | Print version and exit | ||
-h, --help | Show help message and exit |
1:1 Pixel-Accurate Visual Export (DOCX):
anyconvert flyer.pdf -f docx -m canvas -o flyer.docx
1:1 Pixel-Accurate Visual Export (ODT):
anyconvert document.pdf -f odt -m canvas -o document.odt
Semantic Reflowable Conversion for Editing (DOCX):
anyconvert article.pdf -f docx -m flow -o article.docx
Convert PDF Slides to PowerPoint (PPTX):
anyconvert presentation.pdf -f pptx -m canvas -o presentation.pptx
Convert Encrypted / Password-Protected PDF:
anyconvert secret.pdf -f docx -p "MyPassword123"
convert)from anyconvert import convert, ConversionMode
# Accurate visual export to Word or LibreOffice Writer
convert(
input_path="input.pdf",
output_format="docx",
output_path="output.docx",
mode=ConversionMode.CANVAS, # or "canvas"
)
# Semantic reflowable export to OpenDocument Text
convert(
input_path="input.pdf",
output_format="odt",
output_path="output.odt",
mode=ConversionMode.FLOW, # or "flow"
)
convert_bytes)Ideal for web APIs, AWS Lambda, Cloud Run, and microservices:
from anyconvert import convert_bytes
pdf_bytes = open("document.pdf", "rb").read()
# Convert in-memory without touching disk
docx_bytes = convert_bytes(
pdf_bytes=pdf_bytes,
output_format="docx",
mode="canvas",
)
pdf_to_document_ir)Inspect and manipulate document AST nodes before emitting:
from anyconvert import pdf_to_document_ir
doc_ir = pdf_to_document_ir("document.pdf", mode="canvas")
print(f"Total Pages: {len(doc_ir.pages)}")
print(f"Metadata: {doc_ir.metadata}")
for page in doc_ir.pages:
print(f"Page {page.page_number} ({page.width}x{page.height} pt): {len(page.blocks)} blocks")
[ Raw PDF Bytes / Stream ]
│
▼
[ PDF Lexer & Parser ] ───► [ XRef Resolver & ObjStm Unpacker ]
│
┌────────────────────────────────┼────────────────────────────────┐
▼ ▼ ▼
[ Decryption Engine ] [ Decompression Filters ] [ Typography Engine ]
(RC4 / AES-128 / AES-256) (Flate/LZW/ASCII85/PackBits) (SFNT / CFF / Type 1 / CMap)
│ │ │
└────────────────────────────────┼────────────────────────────────┘
│
▼
[ Content Stream Interpreter ]
(Graphics & Text Evaluator)
│
▼
[ Spatial Indexing & Recursive XY-Cut ]
│
▼
[ Semantic Feature Detectors ]
(Headings / Lists / Tables / Flow)
│
▼
[ DocumentIR Builder ]
│
┌────────────────────────────┼────────────────────────────┐
▼ ▼ ▼
[ DOCX Emitter ] [ PPTX Emitter ] [ ODT / ODP / TXT ]
(DrawingML) (DrawingML) (ISO/IEC 26300)
python3 -m pytest tests/
The codebase enforces 100% strict type safety (mypy --strict):
python3 -m mypy --strict src/ tests/
Contributions are welcome! Please read CONTRIBUTING.md for details on our architectural principles, development setup, code of conduct, and pull request guidelines.
GPL-3.0-or-later. See LICENSE for details.
Python
100.0%
Python document conversion engine
anyconvert converts PDF documents into DOCX, PPTX, ODT, ODP, and TXT using only the Python standard library.
--mode canvas): 1:1 pixel-accurate positioning. Renders exact coordinates, font families, embedded images, and vector artwork (native DrawingML custom geometries in DOCX, SVG paths in ODT). Ideal for flyers, forms, annotated PDFs, and presentation slides.--mode flow): Semantic reflowable document reconstruction. Automatically detects headings (H1–H6), bullet/numbered lists, tables, and reflowable paragraphs./Type /XRef), and object streams (/Type /ObjStm).FlateDecode with PNG/TIFF predictors, LZWDecode, ASCII85Decode, ASCIIHexDecode, RunLengthDecode)./Differences encodings..docx) and PresentationML (.pptx) packages with complete [Content_Types].xml and .rels hierarchies..odt) and Presentation (.odp) archives with uncompressed mimetype headers at byte offset 0.| Target Format | Extension | Flow Mode | Canvas Mode | Container Standard |
|---|---|---|---|---|
| Microsoft Word | .docx | Semantic <w:p>, <w:tbl> | DrawingML <wps:wsp>, <w:framePr> | ISO/IEC 29500-2 (OPC) |
| Microsoft PowerPoint | .pptx | Reflowable slide text | Exact coordinate shapes & text | ISO/IEC 29500-2 (OPC) |
| OpenDocument Text | .odt | Semantic <text:p>, <table:table> | Positioned <draw:frame>, <draw:path> | ISO/IEC 26300 (ODF) |
| OpenDocument Presentation | .odp | Slide content boxes | Coordinate <draw:frame> | ISO/IEC 26300 (ODF) |
| Plaintext / Markdown | .txt | Formatted Markdown text | 2D character-grid layout | UTF-8 Plaintext |
Install from PyPI:
pip install anyconvert
Or install in editable development mode:
git clone https://github.com/word-sys/anyconvert.git
cd anyconvert
pip install -e .
anyconvert provides a standalone command-line tool:
anyconvert [OPTIONS] INPUT_PDF
| Flag | Argument | Description | Default |
|---|---|---|---|
-f, --format | docx|pptx|odt|odp|txt | Target document format | docx |
-o, --output | PATH | Output destination file path | {input_basename}.{format} |
-m, --mode | flow|canvas | Layout reconstruction mode (canvas for 1:1 visual match) | flow |
-p, --password | STRING | Decryption password for encrypted PDFs | "" |
-v, --verbose | Enable diagnostic logging | False | |
--version | Print version and exit | ||
-h, --help | Show help message and exit |
1:1 Pixel-Accurate Visual Export (DOCX):
anyconvert flyer.pdf -f docx -m canvas -o flyer.docx
1:1 Pixel-Accurate Visual Export (ODT):
anyconvert document.pdf -f odt -m canvas -o document.odt
Semantic Reflowable Conversion for Editing (DOCX):
anyconvert article.pdf -f docx -m flow -o article.docx
Convert PDF Slides to PowerPoint (PPTX):
anyconvert presentation.pdf -f pptx -m canvas -o presentation.pptx
Convert Encrypted / Password-Protected PDF:
anyconvert secret.pdf -f docx -p "MyPassword123"
convert)from anyconvert import convert, ConversionMode
# Accurate visual export to Word or LibreOffice Writer
convert(
input_path="input.pdf",
output_format="docx",
output_path="output.docx",
mode=ConversionMode.CANVAS, # or "canvas"
)
# Semantic reflowable export to OpenDocument Text
convert(
input_path="input.pdf",
output_format="odt",
output_path="output.odt",
mode=ConversionMode.FLOW, # or "flow"
)
convert_bytes)Ideal for web APIs, AWS Lambda, Cloud Run, and microservices:
from anyconvert import convert_bytes
pdf_bytes = open("document.pdf", "rb").read()
# Convert in-memory without touching disk
docx_bytes = convert_bytes(
pdf_bytes=pdf_bytes,
output_format="docx",
mode="canvas",
)
pdf_to_document_ir)Inspect and manipulate document AST nodes before emitting:
from anyconvert import pdf_to_document_ir
doc_ir = pdf_to_document_ir("document.pdf", mode="canvas")
print(f"Total Pages: {len(doc_ir.pages)}")
print(f"Metadata: {doc_ir.metadata}")
for page in doc_ir.pages:
print(f"Page {page.page_number} ({page.width}x{page.height} pt): {len(page.blocks)} blocks")
[ Raw PDF Bytes / Stream ]
│
▼
[ PDF Lexer & Parser ] ───► [ XRef Resolver & ObjStm Unpacker ]
│
┌────────────────────────────────┼────────────────────────────────┐
▼ ▼ ▼
[ Decryption Engine ] [ Decompression Filters ] [ Typography Engine ]
(RC4 / AES-128 / AES-256) (Flate/LZW/ASCII85/PackBits) (SFNT / CFF / Type 1 / CMap)
│ │ │
└────────────────────────────────┼────────────────────────────────┘
│
▼
[ Content Stream Interpreter ]
(Graphics & Text Evaluator)
│
▼
[ Spatial Indexing & Recursive XY-Cut ]
│
▼
[ Semantic Feature Detectors ]
(Headings / Lists / Tables / Flow)
│
▼
[ DocumentIR Builder ]
│
┌────────────────────────────┼────────────────────────────┐
▼ ▼ ▼
[ DOCX Emitter ] [ PPTX Emitter ] [ ODT / ODP / TXT ]
(DrawingML) (DrawingML) (ISO/IEC 26300)
python3 -m pytest tests/
The codebase enforces 100% strict type safety (mypy --strict):
python3 -m mypy --strict src/ tests/
Contributions are welcome! Please read CONTRIBUTING.md for details on our architectural principles, development setup, code of conduct, and pull request guidelines.
GPL-3.0-or-later. See LICENSE for details.
Python
100.0%