Apache Tika as a self-contained portable executable — use it from Python and other languages without installing Java.
Python
0
5 commits
updated Oct 1, 2026
Apache Tika 4 as a portable CLI and from Python or Node.js, without installing Java.
tika-ape bundles Apache Tika 4 and its Java runtime into a portable
tika.com, then exposes document extraction through language-native bindings
generated with APEBind. No system Java,
Tika server, or JVM configuration is required.
tika.comYou can download the latest tika.com release
with tika-config.json and use it directly as an APE CLI:
./tika.com --config=tika-config.json --text report.pdf
./tika.com --config=tika-config.json --json report.pdf
./tika.com --config=tika-config.json --detect report.pdf
The standalone CLI has no JNI support, so its companion text-only profile disables PDF inline-image extraction. It still extracts text, emits document metadata as JSON, and detects MIME types without a host Java runtime. The Python binding supplies a Host Services provider for java-ape, and therefore enables the tested inline-image path by default.
Install the Python wheel directly from the release:
python -m pip install https://github.com/bear0330/tika-ape/releases/download/v0.2.7/tika_ape-0.2.7-py3-none-any.whl
Or install the Node.js package tarball:
npm install https://github.com/bear0330/tika-ape/releases/download/v0.2.7/tika-ape-0.2.7.tgz
Replace both version strings with the chosen release version. The package
contains the same portable tika.com asset as the release.
Python provides a small, tika-python-compatible surface:
import tika_ape
from tika_ape import parser
text = tika_ape.extract_text('report.pdf')
document = parser.from_file('report.pdf')
Node.js exposes the same text-oriented CLI capability:
import { extractJson, extractText } from 'tika-ape';
const text = await extractText('report.pdf');
const metadata = await extractJson('report.pdf');
Unlike standalone tika.com, the Python binding enables PDF inline-image
extraction by default because it bundles a Host Services provider. Node.js
does not yet ship that provider, so it remains text-oriented. Use Python's
text-only profile when embedded image output is unnecessary:
import tika_ape
tika_ape.configure(inline_images=False)
text = tika_ape.extract_text('report.pdf')
The Python provider also implements LCMS profile and colour conversion through
Pillow's ImageCms. For PDFBox's tested image path, it registers a generic
MaskBlit primitive and handles unclipped, maskless Src and SrcOver
compositing through BufferedImage pixels. This covers the Klook voucher PDF
used by the regression suite.
tika_ape.initVM() remains as a tika-python compatibility alias; it
configures the package and does not start a JVM.
Text, metadata, MIME detection, encoding, language detection, and normal PDF extraction run without host Java. Full JNI/AWT compatibility is not promised. The Python binding supports the tested PDFBox inline-image paths, including Klook's ICC/masked-image document. Other composite rules, clipping, and coverage-mask operations remain outside the current Host provider. OCR and media-transcoding parsers still need their respective external tools.
document_digest is a small Python program that
turns a folder of mixed documents into one Markdown corpus with source paths,
MIME types, metadata, and extracted text. It is useful as an LLM-context or
RAG-ingestion step.
evidence_copilot is an optional local AI demo.
It downloads public AI materials, uses Tika to extract them, and uses a local
embedding model to create an evidence-backed HTML briefing. It requires
sentence-transformers and downloads its model on first use.
tika.com is built with javacosmofy
from the paired java.com and java-modules.zip assets released by
java-ape. Tika needs the
java.desktop module closure.
The reviewed tika.apebind.yaml is the generated binding contract. Regenerate with APEBind:
apebind validate tika.apebind.yaml
apebind generate tika.apebind.yaml --ape tika.com --lang python -o bindings/python
apebind generate tika.apebind.yaml --ape tika.com --lang node -o bindings/node
Root tests use Python's standard library to verify the tika.com CLI itself.
Their fixtures are shared by every binding. prepare downloads the latest
release asset when the local cache is absent; set TIKA_APE_VERSION to pin a
tag or TIKA_APE_FORCE_DOWNLOAD=1 to refresh it. Each binding owns its
language API tests and materializes the root APE and shared fixtures into
ignored local staging paths before testing or packaging:
& .\scripts\test.ps1
& .\bindings\python\scripts\test.ps1
& .\bindings\node\scripts\test.ps1
On Linux or macOS, run ./scripts/test.sh, ./bindings/python/scripts/test.sh,
and ./bindings/node/scripts/test.sh. Use the corresponding binding build
scripts to produce the Python wheel or npm tarball.
All fixtures live in tests/fixtures/. Each binding preparation script copies
that directory into its ignored tests/fixtures/ staging path before tests run.
TIKA_PYTHON_NOTICE records the source and license for fixtures adapted from
tika-python.
Python
79.5%
JavaScript
13.6%
PowerShell
3.8%
Shell
3.0%
Apache Tika as a self-contained portable executable — use it from Python and other languages without installing Java.
Python
0
5 commits
updated Oct 1, 2026
Apache Tika 4 as a portable CLI and from Python or Node.js, without installing Java.
tika-ape bundles Apache Tika 4 and its Java runtime into a portable
tika.com, then exposes document extraction through language-native bindings
generated with APEBind. No system Java,
Tika server, or JVM configuration is required.
tika.comYou can download the latest tika.com release
with tika-config.json and use it directly as an APE CLI:
./tika.com --config=tika-config.json --text report.pdf
./tika.com --config=tika-config.json --json report.pdf
./tika.com --config=tika-config.json --detect report.pdf
The standalone CLI has no JNI support, so its companion text-only profile disables PDF inline-image extraction. It still extracts text, emits document metadata as JSON, and detects MIME types without a host Java runtime. The Python binding supplies a Host Services provider for java-ape, and therefore enables the tested inline-image path by default.
Install the Python wheel directly from the release:
python -m pip install https://github.com/bear0330/tika-ape/releases/download/v0.2.7/tika_ape-0.2.7-py3-none-any.whl
Or install the Node.js package tarball:
npm install https://github.com/bear0330/tika-ape/releases/download/v0.2.7/tika-ape-0.2.7.tgz
Replace both version strings with the chosen release version. The package
contains the same portable tika.com asset as the release.
Python provides a small, tika-python-compatible surface:
import tika_ape
from tika_ape import parser
text = tika_ape.extract_text('report.pdf')
document = parser.from_file('report.pdf')
Node.js exposes the same text-oriented CLI capability:
import { extractJson, extractText } from 'tika-ape';
const text = await extractText('report.pdf');
const metadata = await extractJson('report.pdf');
Unlike standalone tika.com, the Python binding enables PDF inline-image
extraction by default because it bundles a Host Services provider. Node.js
does not yet ship that provider, so it remains text-oriented. Use Python's
text-only profile when embedded image output is unnecessary:
import tika_ape
tika_ape.configure(inline_images=False)
text = tika_ape.extract_text('report.pdf')
The Python provider also implements LCMS profile and colour conversion through
Pillow's ImageCms. For PDFBox's tested image path, it registers a generic
MaskBlit primitive and handles unclipped, maskless Src and SrcOver
compositing through BufferedImage pixels. This covers the Klook voucher PDF
used by the regression suite.
tika_ape.initVM() remains as a tika-python compatibility alias; it
configures the package and does not start a JVM.
Text, metadata, MIME detection, encoding, language detection, and normal PDF extraction run without host Java. Full JNI/AWT compatibility is not promised. The Python binding supports the tested PDFBox inline-image paths, including Klook's ICC/masked-image document. Other composite rules, clipping, and coverage-mask operations remain outside the current Host provider. OCR and media-transcoding parsers still need their respective external tools.
document_digest is a small Python program that
turns a folder of mixed documents into one Markdown corpus with source paths,
MIME types, metadata, and extracted text. It is useful as an LLM-context or
RAG-ingestion step.
evidence_copilot is an optional local AI demo.
It downloads public AI materials, uses Tika to extract them, and uses a local
embedding model to create an evidence-backed HTML briefing. It requires
sentence-transformers and downloads its model on first use.
tika.com is built with javacosmofy
from the paired java.com and java-modules.zip assets released by
java-ape. Tika needs the
java.desktop module closure.
The reviewed tika.apebind.yaml is the generated binding contract. Regenerate with APEBind:
apebind validate tika.apebind.yaml
apebind generate tika.apebind.yaml --ape tika.com --lang python -o bindings/python
apebind generate tika.apebind.yaml --ape tika.com --lang node -o bindings/node
Root tests use Python's standard library to verify the tika.com CLI itself.
Their fixtures are shared by every binding. prepare downloads the latest
release asset when the local cache is absent; set TIKA_APE_VERSION to pin a
tag or TIKA_APE_FORCE_DOWNLOAD=1 to refresh it. Each binding owns its
language API tests and materializes the root APE and shared fixtures into
ignored local staging paths before testing or packaging:
& .\scripts\test.ps1
& .\bindings\python\scripts\test.ps1
& .\bindings\node\scripts\test.ps1
On Linux or macOS, run ./scripts/test.sh, ./bindings/python/scripts/test.sh,
and ./bindings/node/scripts/test.sh. Use the corresponding binding build
scripts to produce the Python wheel or npm tarball.
All fixtures live in tests/fixtures/. Each binding preparation script copies
that directory into its ignored tests/fixtures/ staging path before tests run.
TIKA_PYTHON_NOTICE records the source and license for fixtures adapted from
tika-python.
Python
79.5%
JavaScript
13.6%
PowerShell
3.8%
Shell
3.0%