Compress JSON for LLM prompts — same data, fewer tokens.
Author: Hermann Samimi
PyPI · Repository · Issues
jtoken strips JSON syntactic noise, collapses repeated booleans and nulls into summary lines, flattens nested dicts with dot notation, and supports normalization for Elasticsearch hits and MongoDB JSON. The package ships as a stdlib-first library with an optional tiktoken extra and a jtoken CLI.
[!TIP] Install
jtoken[tiktoken]when you want OpenAI-compatible token counts from the real tokenizer. The core package uses only the standard library and falls back to an estimate whentiktokenis not installed.
pip install jtoken
No extra runtime dependencies.
pip install "jtoken[tiktoken]"
Use the tiktoken extra when you want OpenAI-compatible token counts instead of the built-in estimate backend.
import jtoken
data = {
"user": "alice",
"age": 30,
"premium": True,
"verified": True,
"is_remote": False,
"trial": False,
"score": 9.5,
"referral": None,
"last_login": None,
}
text = jtoken.encode(data)
original = jtoken.decode(text)
assert original == data
dumps / loads are json-style aliases for encode / decode.
JSON
{"name": "Alice", "age": 30, "active": true, "verified": false, "ref": null}
jtoken
name: Alice
age: 30
trues: active
falses: verified
nulls: ref
Nested dicts flatten with dot notation. Booleans and nulls at any depth collapse into the same summary lines. Decode reconstructs the original nested structure.
[!NOTE] Decoding back to Mongo shell, Extended JSON, or an Elasticsearch hit envelope is only lossless when you keep the normalization context from encode (
NormalizationContext, or--context-outon the CLI) and pass it into decode (context=, or--context-in). Plainjson/pythonoutput does not need a sidecar.
Foreign document shapes can be normalized before encoding and restored after decode with a sidecar context.
flowchart LR
raw[RawDoc] --> norm[Normalize]
norm --> enc[Encode]
enc --> jt[JtokenText]
enc --> ctx[ContextSidecar]
jt --> dec[Decode]
ctx --> dec
dec --> denorm[Denormalize]
denorm --> out[Output]
import jtoken
raw_hit = {...}
normalized, context = jtoken.normalize(raw_hit, source="elastic_hit")
text = jtoken.encode(normalized)
restored = jtoken.denormalize(
jtoken.decode(text),
target="elastic_hit",
context=context,
)
jtoken encode --input-format elastic_hit -f hit.json --context-out hit.ctx.json
jtoken decode --output-format mongo_shell -f hit.jtoken --context-in hit.ctx.json
Use source= / target= in Python or --input-format / --output-format on the CLI. encode, stats, and count accept --input-format (default auto). decode accepts --output-format (default json).
Input (source / --input-format) | Use when |
|---|---|
auto | Let jtoken detect the dialect from the text or object shape |
json | Standard JSON object |
python | Same JSON parser as json |
mongo_extended | MongoDB Extended JSON with $oid, $date, $numberInt, $numberLong, $numberDouble, $numberDecimal |
mongo_shell | MongoDB shell document with ObjectId(), ISODate(), NumberInt(), NumberLong() |
elastic_hit | Elasticsearch search hit with _source (and optional fields) |
elastic_source | _source payload only, or a document wrapped as {"_source": {...}} |
Output (target / --output-format) | Use when |
|---|---|
python | Python repr (Python API default) |
json | Pretty-printed JSON (CLI decode default) |
mongo_extended | Extended JSON; requires a context sidecar for BSON-like types |
mongo_shell | Mongo shell document; requires a context sidecar for BSON-like types |
elastic_hit | Full Elasticsearch hit envelope; requires a context sidecar |
elastic_source | JSON shaped like an Elasticsearch _source wrapper |
With auto, jtoken picks mongo_shell when it sees ObjectId(...) or ISODate(...), elastic_hit when the object has a dict _source, mongo_extended when Extended JSON markers such as $oid or $date appear, and otherwise json.
Write the normalization context to a sidecar on encode (--context-out / NormalizationContext.to_dict()) and pass it back on decode when the output dialect is not plain JSON or Python. The sidecar records list paths, dotted keys, Elasticsearch envelope metadata, and MongoDB type markers in typed_values (object_id, datetime, long).
Mongo shell input is parsed as JSON after rewriting shell literals: ObjectId("...") and ISODate("...") become Extended JSON, NumberInt(n) becomes a plain integer, and NumberLong(n) becomes {"$numberLong": "n"}. On normalize, object_id, datetime, and long values are stored in the context so mongo_extended and mongo_shell output can restore {"$oid": ...} / ObjectId(...), {"$date": ...} / ISODate(...), and {"$numberLong": ...} / NumberLong(...). $numberInt, $numberDouble, and $numberDecimal are coerced to Python scalars and are not tracked in typed_values.
elastic_hit encodes the merged _source document (plus any fields values that are not already present in _source) and stores _index, _id, _version, _score, _type, and _routing in the context for lossless elastic_hit output.
echo '{"name": "Alice", "active": true}' | jtoken encode
echo 'name: Alice\ntrues: active' | jtoken decode
echo '{"name": "Alice", "active": true}' | jtoken stats
echo '{"name": "Alice", "active": true}' | jtoken count
Use -f/--file for file input. encode, stats, and count accept --input-format. decode accepts --output-format and --context-in when restoring non-JSON dialects. stats and count accept --model and --backend.
import jtoken
stats = jtoken.token_savings(data, model="gpt-4o", backend="tiktoken", json_indent=2)
print(stats)
# jtoken: 22 tokens | json: 36 tokens | saved: 14 (38.9%)
print(stats.jtoken_tokens, stats.json_tokens, stats.saved, stats.percent)
count_tokens and count_text_tokens are also available. Savings compare the jtoken representation against pretty JSON by default (json_indent=2).
Measured with tiktoken (cl100k_base) on 50 synthetic documents per shape, encoded one
document at a time (as they would be injected into a prompt). Every payload is verified
lossless (encode_document → decode_document → equal). Re-run it with the bundled script:
python3 benchmarks/benchmark.py
| Payload (50 docs) | JSON (pretty) | jtoken | Saved |
|---|---|---|---|
| Elasticsearch hits | 12,038 | 10,671 | 11.4% |
| MongoDB documents (extended JSON) | 9,437 | 7,630 | 19.1% |
| Nested API events | 12,773 | 11,083 | 13.2% |
| Total | 34,248 | 29,384 | 14.2% |
Savings depend on structure: short prose-heavy values cap the wins, while repetitive, nested machine-generated JSON (logs, monitoring, index payloads) compresses substantially better. Run the script on your own payloads before drawing conclusions.
jtoken.__version__jtoken.__author__encode(data: dict) -> strdecode(text: str) -> dictdumps / loadsparse_input(text, *, source="auto")normalize(data, *, source="auto", context=None) -> tuple[dict, NormalizationContext]denormalize(data, *, target="python", context)render_output(value, *, target="python") -> strencode_document(raw, *, source="auto", context=None) -> tuple[str, NormalizationContext]decode_document(text, *, target="python", context)count_tokens(data, *, model="cl100k_base", backend="auto") -> intcount_text_tokens(text, *, model="cl100k_base", backend="auto") -> inttoken_savings(data, *, model="cl100k_base", backend="auto", json_indent=2) -> TokenSavingsTokenSavingsjtoken_tokensjson_tokenssavedpercentNormalizationContextsource_formattarget_formattyped_valueslistsdotted_keyselasticto_dict() / from_dict()InputFormatOutputFormatJPackErrorJPackEncodeErrorJPackDecodeErrorNormalizationErrorDenormalizationErrorTokenCountErrorgit clone https://github.com/HermannSamimi/jtoken.git
cd jtoken
pip install -e ".[dev]"
pytest
pytest --cov=jtoken --cov-report=term-missing
Pull requests are welcome. Install the dev extra (see Development), run pytest, and open a PR against the default branch with a short description of the change.
If you discover a security issue, please report it privately via GitHub Security advisories for this repository rather than a public issue.
MIT — © 2026 Hermann Samimi
⭐ If jtoken saved you tokens (or money on your API bill), please star the repo — it helps other LLM developers find it. And if you measured savings on your payloads, an issue with the numbers (and shape of data) is very welcome.
Python
100.0%
Compress JSON for LLM prompts — same data, fewer tokens.
Author: Hermann Samimi
PyPI · Repository · Issues
jtoken strips JSON syntactic noise, collapses repeated booleans and nulls into summary lines, flattens nested dicts with dot notation, and supports normalization for Elasticsearch hits and MongoDB JSON. The package ships as a stdlib-first library with an optional tiktoken extra and a jtoken CLI.
[!TIP] Install
jtoken[tiktoken]when you want OpenAI-compatible token counts from the real tokenizer. The core package uses only the standard library and falls back to an estimate whentiktokenis not installed.
pip install jtoken
No extra runtime dependencies.
pip install "jtoken[tiktoken]"
Use the tiktoken extra when you want OpenAI-compatible token counts instead of the built-in estimate backend.
import jtoken
data = {
"user": "alice",
"age": 30,
"premium": True,
"verified": True,
"is_remote": False,
"trial": False,
"score": 9.5,
"referral": None,
"last_login": None,
}
text = jtoken.encode(data)
original = jtoken.decode(text)
assert original == data
dumps / loads are json-style aliases for encode / decode.
JSON
{"name": "Alice", "age": 30, "active": true, "verified": false, "ref": null}
jtoken
name: Alice
age: 30
trues: active
falses: verified
nulls: ref
Nested dicts flatten with dot notation. Booleans and nulls at any depth collapse into the same summary lines. Decode reconstructs the original nested structure.
[!NOTE] Decoding back to Mongo shell, Extended JSON, or an Elasticsearch hit envelope is only lossless when you keep the normalization context from encode (
NormalizationContext, or--context-outon the CLI) and pass it into decode (context=, or--context-in). Plainjson/pythonoutput does not need a sidecar.
Foreign document shapes can be normalized before encoding and restored after decode with a sidecar context.
flowchart LR
raw[RawDoc] --> norm[Normalize]
norm --> enc[Encode]
enc --> jt[JtokenText]
enc --> ctx[ContextSidecar]
jt --> dec[Decode]
ctx --> dec
dec --> denorm[Denormalize]
denorm --> out[Output]
import jtoken
raw_hit = {...}
normalized, context = jtoken.normalize(raw_hit, source="elastic_hit")
text = jtoken.encode(normalized)
restored = jtoken.denormalize(
jtoken.decode(text),
target="elastic_hit",
context=context,
)
jtoken encode --input-format elastic_hit -f hit.json --context-out hit.ctx.json
jtoken decode --output-format mongo_shell -f hit.jtoken --context-in hit.ctx.json
Use source= / target= in Python or --input-format / --output-format on the CLI. encode, stats, and count accept --input-format (default auto). decode accepts --output-format (default json).
Input (source / --input-format) | Use when |
|---|---|
auto | Let jtoken detect the dialect from the text or object shape |
json | Standard JSON object |
python | Same JSON parser as json |
mongo_extended | MongoDB Extended JSON with $oid, $date, $numberInt, $numberLong, $numberDouble, $numberDecimal |
mongo_shell | MongoDB shell document with ObjectId(), ISODate(), NumberInt(), NumberLong() |
elastic_hit | Elasticsearch search hit with _source (and optional fields) |
elastic_source | _source payload only, or a document wrapped as {"_source": {...}} |
Output (target / --output-format) | Use when |
|---|---|
python | Python repr (Python API default) |
json | Pretty-printed JSON (CLI decode default) |
mongo_extended | Extended JSON; requires a context sidecar for BSON-like types |
mongo_shell | Mongo shell document; requires a context sidecar for BSON-like types |
elastic_hit | Full Elasticsearch hit envelope; requires a context sidecar |
elastic_source | JSON shaped like an Elasticsearch _source wrapper |
With auto, jtoken picks mongo_shell when it sees ObjectId(...) or ISODate(...), elastic_hit when the object has a dict _source, mongo_extended when Extended JSON markers such as $oid or $date appear, and otherwise json.
Write the normalization context to a sidecar on encode (--context-out / NormalizationContext.to_dict()) and pass it back on decode when the output dialect is not plain JSON or Python. The sidecar records list paths, dotted keys, Elasticsearch envelope metadata, and MongoDB type markers in typed_values (object_id, datetime, long).
Mongo shell input is parsed as JSON after rewriting shell literals: ObjectId("...") and ISODate("...") become Extended JSON, NumberInt(n) becomes a plain integer, and NumberLong(n) becomes {"$numberLong": "n"}. On normalize, object_id, datetime, and long values are stored in the context so mongo_extended and mongo_shell output can restore {"$oid": ...} / ObjectId(...), {"$date": ...} / ISODate(...), and {"$numberLong": ...} / NumberLong(...). $numberInt, $numberDouble, and $numberDecimal are coerced to Python scalars and are not tracked in typed_values.
elastic_hit encodes the merged _source document (plus any fields values that are not already present in _source) and stores _index, _id, _version, _score, _type, and _routing in the context for lossless elastic_hit output.
echo '{"name": "Alice", "active": true}' | jtoken encode
echo 'name: Alice\ntrues: active' | jtoken decode
echo '{"name": "Alice", "active": true}' | jtoken stats
echo '{"name": "Alice", "active": true}' | jtoken count
Use -f/--file for file input. encode, stats, and count accept --input-format. decode accepts --output-format and --context-in when restoring non-JSON dialects. stats and count accept --model and --backend.
import jtoken
stats = jtoken.token_savings(data, model="gpt-4o", backend="tiktoken", json_indent=2)
print(stats)
# jtoken: 22 tokens | json: 36 tokens | saved: 14 (38.9%)
print(stats.jtoken_tokens, stats.json_tokens, stats.saved, stats.percent)
count_tokens and count_text_tokens are also available. Savings compare the jtoken representation against pretty JSON by default (json_indent=2).
Measured with tiktoken (cl100k_base) on 50 synthetic documents per shape, encoded one
document at a time (as they would be injected into a prompt). Every payload is verified
lossless (encode_document → decode_document → equal). Re-run it with the bundled script:
python3 benchmarks/benchmark.py
| Payload (50 docs) | JSON (pretty) | jtoken | Saved |
|---|---|---|---|
| Elasticsearch hits | 12,038 | 10,671 | 11.4% |
| MongoDB documents (extended JSON) | 9,437 | 7,630 | 19.1% |
| Nested API events | 12,773 | 11,083 | 13.2% |
| Total | 34,248 | 29,384 | 14.2% |
Savings depend on structure: short prose-heavy values cap the wins, while repetitive, nested machine-generated JSON (logs, monitoring, index payloads) compresses substantially better. Run the script on your own payloads before drawing conclusions.
jtoken.__version__jtoken.__author__encode(data: dict) -> strdecode(text: str) -> dictdumps / loadsparse_input(text, *, source="auto")normalize(data, *, source="auto", context=None) -> tuple[dict, NormalizationContext]denormalize(data, *, target="python", context)render_output(value, *, target="python") -> strencode_document(raw, *, source="auto", context=None) -> tuple[str, NormalizationContext]decode_document(text, *, target="python", context)count_tokens(data, *, model="cl100k_base", backend="auto") -> intcount_text_tokens(text, *, model="cl100k_base", backend="auto") -> inttoken_savings(data, *, model="cl100k_base", backend="auto", json_indent=2) -> TokenSavingsTokenSavingsjtoken_tokensjson_tokenssavedpercentNormalizationContextsource_formattarget_formattyped_valueslistsdotted_keyselasticto_dict() / from_dict()InputFormatOutputFormatJPackErrorJPackEncodeErrorJPackDecodeErrorNormalizationErrorDenormalizationErrorTokenCountErrorgit clone https://github.com/HermannSamimi/jtoken.git
cd jtoken
pip install -e ".[dev]"
pytest
pytest --cov=jtoken --cov-report=term-missing
Pull requests are welcome. Install the dev extra (see Development), run pytest, and open a PR against the default branch with a short description of the change.
If you discover a security issue, please report it privately via GitHub Security advisories for this repository rather than a public issue.
MIT — © 2026 Hermann Samimi
⭐ If jtoken saved you tokens (or money on your API bill), please star the repo — it helps other LLM developers find it. And if you measured savings on your payloads, an issue with the numbers (and shape of data) is very welcome.
Python
100.0%