HermannSamimi/jtoken

json token minimizer

Python

1

0 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Jtoken – lossless JSON compression for LLM prompts

2

Sep 29, 2026

README

jtoken

jtoken

PyPI version Python versions License Issues

Compress JSON for LLM prompts — same data, fewer tokens.

Author: Hermann Samimi

PyPI · Repository · Issues

jtoken strips JSON syntactic noise, collapses repeated booleans and nulls into summary lines, flattens nested dicts with dot notation, and supports normalization for Elasticsearch hits and MongoDB JSON. The package ships as a stdlib-first library with an optional tiktoken extra and a jtoken CLI.

Table of contents

Installation

[!TIP] Install jtoken[tiktoken] when you want OpenAI-compatible token counts from the real tokenizer. The core package uses only the standard library and falls back to an estimate when tiktoken is not installed.

Core

pip install jtoken

No extra runtime dependencies.

With tokenizer-accurate counting

pip install "jtoken[tiktoken]"

Use the tiktoken extra when you want OpenAI-compatible token counts instead of the built-in estimate backend.

Quick start

import jtoken

data = {
    "user": "alice",
    "age": 30,
    "premium": True,
    "verified": True,
    "is_remote": False,
    "trial": False,
    "score": 9.5,
    "referral": None,
    "last_login": None,
}

text = jtoken.encode(data)
original = jtoken.decode(text)
assert original == data

dumps / loads are json-style aliases for encode / decode.

What the format looks like

JSON

{"name": "Alice", "age": 30, "active": true, "verified": false, "ref": null}

jtoken

name: Alice
age: 30
trues: active
falses: verified
nulls: ref

Nested dicts flatten with dot notation. Booleans and nulls at any depth collapse into the same summary lines. Decode reconstructs the original nested structure.

Normalization and denormalization

[!NOTE] Decoding back to Mongo shell, Extended JSON, or an Elasticsearch hit envelope is only lossless when you keep the normalization context from encode (NormalizationContext, or --context-out on the CLI) and pass it into decode (context=, or --context-in). Plain json / python output does not need a sidecar.

Foreign document shapes can be normalized before encoding and restored after decode with a sidecar context.

flowchart LR
  raw[RawDoc] --> norm[Normalize]
  norm --> enc[Encode]
  enc --> jt[JtokenText]
  enc --> ctx[ContextSidecar]
  jt --> dec[Decode]
  ctx --> dec
  dec --> denorm[Denormalize]
  denorm --> out[Output]
import jtoken

raw_hit = {...}
normalized, context = jtoken.normalize(raw_hit, source="elastic_hit")
text = jtoken.encode(normalized)
restored = jtoken.denormalize(
    jtoken.decode(text),
    target="elastic_hit",
    context=context,
)
jtoken encode --input-format elastic_hit -f hit.json --context-out hit.ctx.json
jtoken decode --output-format mongo_shell -f hit.jtoken --context-in hit.ctx.json
Input and output formats (reference tables)

Use source= / target= in Python or --input-format / --output-format on the CLI. encode, stats, and count accept --input-format (default auto). decode accepts --output-format (default json).

Input (source / --input-format)Use when
autoLet jtoken detect the dialect from the text or object shape
jsonStandard JSON object
pythonSame JSON parser as json
mongo_extendedMongoDB Extended JSON with $oid, $date, $numberInt, $numberLong, $numberDouble, $numberDecimal
mongo_shellMongoDB shell document with ObjectId(), ISODate(), NumberInt(), NumberLong()
elastic_hitElasticsearch search hit with _source (and optional fields)
elastic_source_source payload only, or a document wrapped as {"_source": {...}}
Output (target / --output-format)Use when
pythonPython repr (Python API default)
jsonPretty-printed JSON (CLI decode default)
mongo_extendedExtended JSON; requires a context sidecar for BSON-like types
mongo_shellMongo shell document; requires a context sidecar for BSON-like types
elastic_hitFull Elasticsearch hit envelope; requires a context sidecar
elastic_sourceJSON shaped like an Elasticsearch _source wrapper

With auto, jtoken picks mongo_shell when it sees ObjectId(...) or ISODate(...), elastic_hit when the object has a dict _source, mongo_extended when Extended JSON markers such as $oid or $date appear, and otherwise json.

Write the normalization context to a sidecar on encode (--context-out / NormalizationContext.to_dict()) and pass it back on decode when the output dialect is not plain JSON or Python. The sidecar records list paths, dotted keys, Elasticsearch envelope metadata, and MongoDB type markers in typed_values (object_id, datetime, long).

MongoDB shell and Extended JSON

Mongo shell input is parsed as JSON after rewriting shell literals: ObjectId("...") and ISODate("...") become Extended JSON, NumberInt(n) becomes a plain integer, and NumberLong(n) becomes {"$numberLong": "n"}. On normalize, object_id, datetime, and long values are stored in the context so mongo_extended and mongo_shell output can restore {"$oid": ...} / ObjectId(...), {"$date": ...} / ISODate(...), and {"$numberLong": ...} / NumberLong(...). $numberInt, $numberDouble, and $numberDecimal are coerced to Python scalars and are not tracked in typed_values.

Elasticsearch hits

elastic_hit encodes the merged _source document (plus any fields values that are not already present in _source) and stores _index, _id, _version, _score, _type, and _routing in the context for lossless elastic_hit output.

CLI

echo '{"name": "Alice", "active": true}' | jtoken encode
echo 'name: Alice\ntrues: active' | jtoken decode
echo '{"name": "Alice", "active": true}' | jtoken stats
echo '{"name": "Alice", "active": true}' | jtoken count

Use -f/--file for file input. encode, stats, and count accept --input-format. decode accepts --output-format and --context-in when restoring non-JSON dialects. stats and count accept --model and --backend.

Token savings

import jtoken

stats = jtoken.token_savings(data, model="gpt-4o", backend="tiktoken", json_indent=2)
print(stats)
# jtoken: 22 tokens | json: 36 tokens | saved: 14 (38.9%)

print(stats.jtoken_tokens, stats.json_tokens, stats.saved, stats.percent)

count_tokens and count_text_tokens are also available. Savings compare the jtoken representation against pretty JSON by default (json_indent=2).

Measured token savings (reproducible benchmark)

Measured with tiktoken (cl100k_base) on 50 synthetic documents per shape, encoded one document at a time (as they would be injected into a prompt). Every payload is verified lossless (encode_document → decode_document → equal). Re-run it with the bundled script:

python3 benchmarks/benchmark.py
Payload (50 docs)JSON (pretty)jtokenSaved
Elasticsearch hits12,03810,67111.4%
MongoDB documents (extended JSON)9,4377,63019.1%
Nested API events12,77311,08313.2%
Total34,24829,38414.2%

Savings depend on structure: short prose-heavy values cap the wins, while repetitive, nested machine-generated JSON (logs, monitoring, index payloads) compresses substantially better. Run the script on your own payloads before drawing conclusions.

API reference

Expand full API list

Package metadata

  • jtoken.__version__
  • jtoken.__author__

Core codec

  • encode(data: dict) -> str
  • decode(text: str) -> dict
  • dumps / loads

Normalization

  • parse_input(text, *, source="auto")
  • normalize(data, *, source="auto", context=None) -> tuple[dict, NormalizationContext]
  • denormalize(data, *, target="python", context)
  • render_output(value, *, target="python") -> str
  • encode_document(raw, *, source="auto", context=None) -> tuple[str, NormalizationContext]
  • decode_document(text, *, target="python", context)

Token helpers

  • count_tokens(data, *, model="cl100k_base", backend="auto") -> int
  • count_text_tokens(text, *, model="cl100k_base", backend="auto") -> int
  • token_savings(data, *, model="cl100k_base", backend="auto", json_indent=2) -> TokenSavings

TokenSavings

  • jtoken_tokens
  • json_tokens
  • saved
  • percent

NormalizationContext

  • source_format
  • target_format
  • typed_values
  • lists
  • dotted_keys
  • elastic
  • to_dict() / from_dict()

Format enums

  • InputFormat
  • OutputFormat

Exceptions

  • JPackError
  • JPackEncodeError
  • JPackDecodeError
  • NormalizationError
  • DenormalizationError
  • TokenCountError

Development

git clone https://github.com/HermannSamimi/jtoken.git
cd jtoken
pip install -e ".[dev]"
pytest
pytest --cov=jtoken --cov-report=term-missing

Contributing

Pull requests are welcome. Install the dev extra (see Development), run pytest, and open a PR against the default branch with a short description of the change.

Security

If you discover a security issue, please report it privately via GitHub Security advisories for this repository rather than a public issue.

License

MIT — © 2026 Hermann Samimi


⭐ If jtoken saved you tokens (or money on your API bill), please star the repo — it helps other LLM developers find it. And if you measured savings on your payloads, an issue with the numbers (and shape of data) is very welcome.

compression
context-window
json
llm
prompt-engineering
python
rag
token-optimization

HermannSamimi/jtoken

json token minimizer

Python

1

0 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

Jtoken – lossless JSON compression for LLM prompts

2

Sep 29, 2026

README

jtoken

jtoken

PyPI version Python versions License Issues

Compress JSON for LLM prompts — same data, fewer tokens.

Author: Hermann Samimi

PyPI · Repository · Issues

jtoken strips JSON syntactic noise, collapses repeated booleans and nulls into summary lines, flattens nested dicts with dot notation, and supports normalization for Elasticsearch hits and MongoDB JSON. The package ships as a stdlib-first library with an optional tiktoken extra and a jtoken CLI.

Table of contents

Installation

[!TIP] Install jtoken[tiktoken] when you want OpenAI-compatible token counts from the real tokenizer. The core package uses only the standard library and falls back to an estimate when tiktoken is not installed.

Core

pip install jtoken

No extra runtime dependencies.

With tokenizer-accurate counting

pip install "jtoken[tiktoken]"

Use the tiktoken extra when you want OpenAI-compatible token counts instead of the built-in estimate backend.

Quick start

import jtoken

data = {
    "user": "alice",
    "age": 30,
    "premium": True,
    "verified": True,
    "is_remote": False,
    "trial": False,
    "score": 9.5,
    "referral": None,
    "last_login": None,
}

text = jtoken.encode(data)
original = jtoken.decode(text)
assert original == data

dumps / loads are json-style aliases for encode / decode.

What the format looks like

JSON

{"name": "Alice", "age": 30, "active": true, "verified": false, "ref": null}

jtoken

name: Alice
age: 30
trues: active
falses: verified
nulls: ref

Nested dicts flatten with dot notation. Booleans and nulls at any depth collapse into the same summary lines. Decode reconstructs the original nested structure.

Normalization and denormalization

[!NOTE] Decoding back to Mongo shell, Extended JSON, or an Elasticsearch hit envelope is only lossless when you keep the normalization context from encode (NormalizationContext, or --context-out on the CLI) and pass it into decode (context=, or --context-in). Plain json / python output does not need a sidecar.

Foreign document shapes can be normalized before encoding and restored after decode with a sidecar context.

flowchart LR
  raw[RawDoc] --> norm[Normalize]
  norm --> enc[Encode]
  enc --> jt[JtokenText]
  enc --> ctx[ContextSidecar]
  jt --> dec[Decode]
  ctx --> dec
  dec --> denorm[Denormalize]
  denorm --> out[Output]
import jtoken

raw_hit = {...}
normalized, context = jtoken.normalize(raw_hit, source="elastic_hit")
text = jtoken.encode(normalized)
restored = jtoken.denormalize(
    jtoken.decode(text),
    target="elastic_hit",
    context=context,
)
jtoken encode --input-format elastic_hit -f hit.json --context-out hit.ctx.json
jtoken decode --output-format mongo_shell -f hit.jtoken --context-in hit.ctx.json
Input and output formats (reference tables)

Use source= / target= in Python or --input-format / --output-format on the CLI. encode, stats, and count accept --input-format (default auto). decode accepts --output-format (default json).

Input (source / --input-format)Use when
autoLet jtoken detect the dialect from the text or object shape
jsonStandard JSON object
pythonSame JSON parser as json
mongo_extendedMongoDB Extended JSON with $oid, $date, $numberInt, $numberLong, $numberDouble, $numberDecimal
mongo_shellMongoDB shell document with ObjectId(), ISODate(), NumberInt(), NumberLong()
elastic_hitElasticsearch search hit with _source (and optional fields)
elastic_source_source payload only, or a document wrapped as {"_source": {...}}
Output (target / --output-format)Use when
pythonPython repr (Python API default)
jsonPretty-printed JSON (CLI decode default)
mongo_extendedExtended JSON; requires a context sidecar for BSON-like types
mongo_shellMongo shell document; requires a context sidecar for BSON-like types
elastic_hitFull Elasticsearch hit envelope; requires a context sidecar
elastic_sourceJSON shaped like an Elasticsearch _source wrapper

With auto, jtoken picks mongo_shell when it sees ObjectId(...) or ISODate(...), elastic_hit when the object has a dict _source, mongo_extended when Extended JSON markers such as $oid or $date appear, and otherwise json.

Write the normalization context to a sidecar on encode (--context-out / NormalizationContext.to_dict()) and pass it back on decode when the output dialect is not plain JSON or Python. The sidecar records list paths, dotted keys, Elasticsearch envelope metadata, and MongoDB type markers in typed_values (object_id, datetime, long).

MongoDB shell and Extended JSON

Mongo shell input is parsed as JSON after rewriting shell literals: ObjectId("...") and ISODate("...") become Extended JSON, NumberInt(n) becomes a plain integer, and NumberLong(n) becomes {"$numberLong": "n"}. On normalize, object_id, datetime, and long values are stored in the context so mongo_extended and mongo_shell output can restore {"$oid": ...} / ObjectId(...), {"$date": ...} / ISODate(...), and {"$numberLong": ...} / NumberLong(...). $numberInt, $numberDouble, and $numberDecimal are coerced to Python scalars and are not tracked in typed_values.

Elasticsearch hits

elastic_hit encodes the merged _source document (plus any fields values that are not already present in _source) and stores _index, _id, _version, _score, _type, and _routing in the context for lossless elastic_hit output.

CLI

echo '{"name": "Alice", "active": true}' | jtoken encode
echo 'name: Alice\ntrues: active' | jtoken decode
echo '{"name": "Alice", "active": true}' | jtoken stats
echo '{"name": "Alice", "active": true}' | jtoken count

Use -f/--file for file input. encode, stats, and count accept --input-format. decode accepts --output-format and --context-in when restoring non-JSON dialects. stats and count accept --model and --backend.

Token savings

import jtoken

stats = jtoken.token_savings(data, model="gpt-4o", backend="tiktoken", json_indent=2)
print(stats)
# jtoken: 22 tokens | json: 36 tokens | saved: 14 (38.9%)

print(stats.jtoken_tokens, stats.json_tokens, stats.saved, stats.percent)

count_tokens and count_text_tokens are also available. Savings compare the jtoken representation against pretty JSON by default (json_indent=2).

Measured token savings (reproducible benchmark)

Measured with tiktoken (cl100k_base) on 50 synthetic documents per shape, encoded one document at a time (as they would be injected into a prompt). Every payload is verified lossless (encode_document → decode_document → equal). Re-run it with the bundled script:

python3 benchmarks/benchmark.py
Payload (50 docs)JSON (pretty)jtokenSaved
Elasticsearch hits12,03810,67111.4%
MongoDB documents (extended JSON)9,4377,63019.1%
Nested API events12,77311,08313.2%
Total34,24829,38414.2%

Savings depend on structure: short prose-heavy values cap the wins, while repetitive, nested machine-generated JSON (logs, monitoring, index payloads) compresses substantially better. Run the script on your own payloads before drawing conclusions.

API reference

Expand full API list

Package metadata

  • jtoken.__version__
  • jtoken.__author__

Core codec

  • encode(data: dict) -> str
  • decode(text: str) -> dict
  • dumps / loads

Normalization

  • parse_input(text, *, source="auto")
  • normalize(data, *, source="auto", context=None) -> tuple[dict, NormalizationContext]
  • denormalize(data, *, target="python", context)
  • render_output(value, *, target="python") -> str
  • encode_document(raw, *, source="auto", context=None) -> tuple[str, NormalizationContext]
  • decode_document(text, *, target="python", context)

Token helpers

  • count_tokens(data, *, model="cl100k_base", backend="auto") -> int
  • count_text_tokens(text, *, model="cl100k_base", backend="auto") -> int
  • token_savings(data, *, model="cl100k_base", backend="auto", json_indent=2) -> TokenSavings

TokenSavings

  • jtoken_tokens
  • json_tokens
  • saved
  • percent

NormalizationContext

  • source_format
  • target_format
  • typed_values
  • lists
  • dotted_keys
  • elastic
  • to_dict() / from_dict()

Format enums

  • InputFormat
  • OutputFormat

Exceptions

  • JPackError
  • JPackEncodeError
  • JPackDecodeError
  • NormalizationError
  • DenormalizationError
  • TokenCountError

Development

git clone https://github.com/HermannSamimi/jtoken.git
cd jtoken
pip install -e ".[dev]"
pytest
pytest --cov=jtoken --cov-report=term-missing

Contributing

Pull requests are welcome. Install the dev extra (see Development), run pytest, and open a PR against the default branch with a short description of the change.

Security

If you discover a security issue, please report it privately via GitHub Security advisories for this repository rather than a public issue.

License

MIT — © 2026 Hermann Samimi


⭐ If jtoken saved you tokens (or money on your API bill), please star the repo — it helps other LLM developers find it. And if you measured savings on your payloads, an issue with the numbers (and shape of data) is very welcome.

compression
context-window
json
llm
prompt-engineering
python
rag
token-optimization

Languages

Python

100.0%