HyperLogLog with lots of sugar (Sparse, LogLog-Beta bias correction and TailCut space reduction) brought to you by Axiom
See the codeAn improved version of HyperLogLog for the count-distinct problem, approximating the number of distinct elements in a multiset. This implementation offers enhanced performance, flexibility, and simplicity while maintaining accuracy.
The initial version of this work (tagged as v0.1.0) was based on "Better with fewer bits: Improving the performance of cardinality estimation of large data streams - Qingjun Xiao, You Zhou, Shigang Chen". However, the current implementation has evolved significantly from this original basis, notably moving away from the tailcut method.
The current implementation is based on the LogLog-Beta algorithm, as described in:
"LogLog-Beta and More: A New Algorithm for Cardinality Estimation Based on LogLog Counting" by Jason Qin, Denys Kim, and Yumei Tung (2016).
Key features of the current implementation:
This implementation is now more straightforward, efficient, and flexible, while remaining backwards compatible with previous versions. It provides a balance between precision, memory usage, speed, and ease of use.
This implementation allows for creating HyperLogLog sketches with arbitrary precision between 2^4 and 2^18 registers. The memory usage scales with the number of registers:
Users can choose the precision that best fits their use case, balancing memory usage against estimation accuracy.
A big thank you to Prof. Shigang Chen and his team at the University of Florida who are actively conducting research around "Big Network Data".
Kindly check our contributing guide on how to propose bugfixes and improvements, and submit pull requests to the project.
Run the same checks used by CI before opening a pull request:
make ci
Individual targets are available for formatting (make fmt), linting
(make lint), tests (make test), race-enabled tests (make test-race), and
coverage (make coverage).
Version tags matching v*.*.* are validated by the release workflow before it
publishes a GitHub Release with generated release notes. Tags with a pre-release
suffix, such as v1.0.0-rc.1, publish as pre-releases.
© Axiom, Inc., 2026
Distributed under the MIT License (The MIT License).
See LICENSE for more information.
1,497 followers · starred Jun 2017
5,102 followers · starred Jun 2017
401 followers · starred Jan 2022
870 followers · starred Mar 2026
Go
98.8%
Makefile
1.2%
HyperLogLog with lots of sugar (Sparse, LogLog-Beta bias correction and TailCut space reduction) brought to you by Axiom
See the codeAn improved version of HyperLogLog for the count-distinct problem, approximating the number of distinct elements in a multiset. This implementation offers enhanced performance, flexibility, and simplicity while maintaining accuracy.
The initial version of this work (tagged as v0.1.0) was based on "Better with fewer bits: Improving the performance of cardinality estimation of large data streams - Qingjun Xiao, You Zhou, Shigang Chen". However, the current implementation has evolved significantly from this original basis, notably moving away from the tailcut method.
The current implementation is based on the LogLog-Beta algorithm, as described in:
"LogLog-Beta and More: A New Algorithm for Cardinality Estimation Based on LogLog Counting" by Jason Qin, Denys Kim, and Yumei Tung (2016).
Key features of the current implementation:
This implementation is now more straightforward, efficient, and flexible, while remaining backwards compatible with previous versions. It provides a balance between precision, memory usage, speed, and ease of use.
This implementation allows for creating HyperLogLog sketches with arbitrary precision between 2^4 and 2^18 registers. The memory usage scales with the number of registers:
Users can choose the precision that best fits their use case, balancing memory usage against estimation accuracy.
A big thank you to Prof. Shigang Chen and his team at the University of Florida who are actively conducting research around "Big Network Data".
Kindly check our contributing guide on how to propose bugfixes and improvements, and submit pull requests to the project.
Run the same checks used by CI before opening a pull request:
make ci
Individual targets are available for formatting (make fmt), linting
(make lint), tests (make test), race-enabled tests (make test-race), and
coverage (make coverage).
Version tags matching v*.*.* are validated by the release workflow before it
publishes a GitHub Release with generated release notes. Tags with a pre-release
suffix, such as v1.0.0-rc.1, publish as pre-releases.
© Axiom, Inc., 2026
Distributed under the MIT License (The MIT License).
See LICENSE for more information.
1,497 followers · starred Jun 2017
5,102 followers · starred Jun 2017
401 followers · starred Jan 2022
870 followers · starred Mar 2026
Go
98.8%
Makefile
1.2%