georg-jung/FastBertTokenizer

Fast and memory-efficient library for WordPiece tokenization as it is used by BERT.

C#

55

250 commits

updated Oct 6, 2026

See the code

README

FastBertTokenizer logo FastBertTokenizer

NuGet version (FastBertTokenizer) Docs .NET Build codecov

A fast and memory-efficient library for WordPiece tokenization as it is used by BERT. Tokenization correctness and speed are automatically evaluated in extensive unit tests and benchmarks. Native AOT compatible and targets net10.0, net8.0 and netstandard2.0.

Goals

  • Enabling you to run your AI workloads on .NET in production.
  • Correctness - Results that are equivalent to HuggingFace Transformers' AutoTokenizer's in all practical cases.
  • Speed - Tokenization should be as fast as reasonably possible.
  • Ease of use - The API should be easy to understand and use.

Getting Started

dotnet new console
dotnet add package FastBertTokenizer
using FastBertTokenizer;

var tok = new BertTokenizer();
await tok.LoadFromHuggingFaceAsync("bert-base-uncased");
var (inputIds, attentionMask, tokenTypeIds) = tok.Encode("Lorem ipsum dolor sit amet.");
Console.WriteLine(string.Join(", ", inputIds.ToArray()));
var decoded = tok.Decode(inputIds.Span);
Console.WriteLine(decoded);

// Output:
// 101, 19544, 2213, 12997, 17421, 2079, 10626, 4133, 2572, 3388, 1012, 102
// [CLS] lorem ipsum dolor sit amet. [SEP]

example project

Note: FastBertTokenizer currently does not support encoding two pieces of text into a single input with a separator in between and corresponding token_type_ids, as some models (e.g. cross-encoders) expect.

Speed / Benchmarks

tl;dr: FastBertTokenizer encodes ~14.5 million tokens per second on a single core, enough to tokenize a full-length novel in under 10 ms. Batched across the 4 vCPUs of a GitHub Actions runner, that grows to ~35 million tokens per second.

Market overview from a full CI run (GitHub Actions shared runner, ubuntu-24.04, 4 vCPUs): tokenizing 15,000 simple english wikipedia articles (3,657,145 tokens) with bert-base-uncased's vocabulary, truncated to 512 tokens per input. For FastBertTokenizer that is ~14.5m tokens/s single threaded and ~35.3m tokens/s multi threaded.

LibraryMeasured fromSingle threadedParallel
FastBertTokenizer.NET265 ms104 ms
tokie (Rust)Python513 ms230 ms
Microsoft.ML.Tokenizers.NET785 ms—
flash-tokenizer (C++)Python1.13 s767 ms
BlingFire (C++).NET1.22 s—
Tokenizers.DotNet (HF bindings).NET4.87 s—
Hugging Face tokenizers (Rust)Python9.30 s3.93 s

The libraries don't all do exactly the same work and cross-language numbers are only roughly comparable: e.g. Hugging Face tokenizers' single-threaded number includes per-call Python overhead, and tokie may use multiple cores even for sequential calls. See src/Benchmarks/README.md for all detailed results (incl. FastBertTokenizer's different usage patterns and runtimes), the exact environment, fairness notes, and how to run the benchmarks yourself.

Created by combining https://icons.getbootstrap.com/icons/cursor-text/ in .NET brand color with https://icons.getbootstrap.com/icons/braces/.

ai
bert
bert-embeddings
llm
machine-learning
natural-language-processing
nlp
nlp-machine-learning
tokenization
tokens
wordpiece
wordpiece-tokenization

georg-jung/FastBertTokenizer

Fast and memory-efficient library for WordPiece tokenization as it is used by BERT.

C#

55

250 commits

updated Oct 6, 2026

See the code

README

FastBertTokenizer logo FastBertTokenizer

NuGet version (FastBertTokenizer) Docs .NET Build codecov

A fast and memory-efficient library for WordPiece tokenization as it is used by BERT. Tokenization correctness and speed are automatically evaluated in extensive unit tests and benchmarks. Native AOT compatible and targets net10.0, net8.0 and netstandard2.0.

Goals

  • Enabling you to run your AI workloads on .NET in production.
  • Correctness - Results that are equivalent to HuggingFace Transformers' AutoTokenizer's in all practical cases.
  • Speed - Tokenization should be as fast as reasonably possible.
  • Ease of use - The API should be easy to understand and use.

Getting Started

dotnet new console
dotnet add package FastBertTokenizer
using FastBertTokenizer;

var tok = new BertTokenizer();
await tok.LoadFromHuggingFaceAsync("bert-base-uncased");
var (inputIds, attentionMask, tokenTypeIds) = tok.Encode("Lorem ipsum dolor sit amet.");
Console.WriteLine(string.Join(", ", inputIds.ToArray()));
var decoded = tok.Decode(inputIds.Span);
Console.WriteLine(decoded);

// Output:
// 101, 19544, 2213, 12997, 17421, 2079, 10626, 4133, 2572, 3388, 1012, 102
// [CLS] lorem ipsum dolor sit amet. [SEP]

example project

Note: FastBertTokenizer currently does not support encoding two pieces of text into a single input with a separator in between and corresponding token_type_ids, as some models (e.g. cross-encoders) expect.

Speed / Benchmarks

tl;dr: FastBertTokenizer encodes ~14.5 million tokens per second on a single core, enough to tokenize a full-length novel in under 10 ms. Batched across the 4 vCPUs of a GitHub Actions runner, that grows to ~35 million tokens per second.

Market overview from a full CI run (GitHub Actions shared runner, ubuntu-24.04, 4 vCPUs): tokenizing 15,000 simple english wikipedia articles (3,657,145 tokens) with bert-base-uncased's vocabulary, truncated to 512 tokens per input. For FastBertTokenizer that is ~14.5m tokens/s single threaded and ~35.3m tokens/s multi threaded.

LibraryMeasured fromSingle threadedParallel
FastBertTokenizer.NET265 ms104 ms
tokie (Rust)Python513 ms230 ms
Microsoft.ML.Tokenizers.NET785 ms—
flash-tokenizer (C++)Python1.13 s767 ms
BlingFire (C++).NET1.22 s—
Tokenizers.DotNet (HF bindings).NET4.87 s—
Hugging Face tokenizers (Rust)Python9.30 s3.93 s

The libraries don't all do exactly the same work and cross-language numbers are only roughly comparable: e.g. Hugging Face tokenizers' single-threaded number includes per-call Python overhead, and tokie may use multiple cores even for sequential calls. See src/Benchmarks/README.md for all detailed results (incl. FastBertTokenizer's different usage patterns and runtimes), the exact environment, fairness notes, and how to run the benchmarks yourself.

Created by combining https://icons.getbootstrap.com/icons/cursor-text/ in .NET brand color with https://icons.getbootstrap.com/icons/braces/.

ai
bert
bert-embeddings
llm
machine-learning
natural-language-processing
nlp
nlp-machine-learning
tokenization
tokens
wordpiece
wordpiece-tokenization

Significant stargazers

Ryan Moore

75 followers · starred Jan 2026