:rocket: Easy to use PDF library using Go and PDFium :rocket:
A fast, multi-threaded and easy to use PDF library for Go applications.
FPDF_SetPrintMode, FPDF_RenderPage)image.Image using either DPI or pixel sizeimage.RGBA, the default) or in grayscale (image.Gray, using
ImageFormat: requests.RenderImageFormatGrayscale), the result is in the RenderedImage response fieldThis project uses the PDFium C++ library by Google (https://pdfium.googlesource.com/pdfium/) to process the PDF documents. Therefor this project could also be called a binding.
Please be aware that PDFium comes with the Apache License 2.0 license.
Since PDFium is not a multithreaded C++ library, we can not directly make it multithreaded by calling it from Go's subroutines.
To solve this, this library has 3 different implementations that you can use to call PDFium:
Both single_threaded and multi_threaded requires PDFium to be installed on your machine for your platform during
compilation and runtime, it also requires CGO to work on the platform you're compiling to.
webassembly does not need any external dependencies and also does not require CGO to work. However, Wazero currently
only has compiler support for amd64 and arm64, meaning it will be using the interpreter on other architectures which
will be much, much slower.
All implementations use exactly the same interface, so there won't be any code changes for you to switch between them.
All implementations are thread/subroutine safe, this has been guaranteed by locking the instance that's doing your work while it's doing PDFium operations. New operations will wait until the lock becomes available again.
Be aware that PDFium could use quite some memory depending on the size of the PDF and the requests that you do, so be aware of the amount of workers that you configure.
Single-threading in CGO works by directly calling the PDFium library from the same process. Single-threaded might be preferred if the caller is managing the workers themselves and does not want the overhead of another process. Be aware that since PDFium is C++, we can't handle segfaults caused by PDFium, which may cause your process to be killed. So using this library in the multi-threaded way, with only 1 worker, can still have some benefits, since it can automatically recover from things like segfaults.
For CGO we have implemented multi-threading using HashiCorp's Go Plugin System, which allows us to launch separate PDFium worker processes, and then route the requests through the different workers. This also makes it a bit more safe to use PDFium, as it's less likely to segfaults or corrupt your main Go application. The Plugin system provides the communication between the processes using gRPC, however, when implementing this library, you won't really see anything of that. From the outside it will look like normal Go code. The inter-process communication does come with a cost as it has to serialize/deserialize input/output as it moves between the main process and the PDFium workers.
To use this Go library, you will need the actual PDFium library to run it and have it available through pkgconfig.
You can try to compile PDFium yourself, but you can also use pre-compiled binaries, for example from: https://github.com/bblanchon/pdfium-binaries/releases
If you use a pre-compiled library, make sure to extract it somewhere logical, for example /opt/pdfium.
Create/edit file /usr/lib/pkgconfig/pdfium.pc
prefix={path}
libdir={path}/lib
includedir={path}/include
Name: PDFium
Description: PDFium
Version: 8057
Requires:
Libs: -L${libdir} -lpdfium
Cflags: -I${includedir}
Replace {path} with the path you extracted/compiled pdfium in.
Make sure you extend your library path when running:
export LD_LIBRARY_PATH={path}/lib
You can do this globally or just in your editor.
this can globally be done on ubuntu by editing ~/.profile
and adding the line in this file. reloading for bash can be done by relogging or running source ~/.profile can be used
to test the change for a terminal
To get started, make sure that you create a separate package in your application that will start the worker.
The examples below can also be found in the examples folder.
For single threaded implementations we just have to initialize the library.
pdfium/renderer/renderer.go
package renderer
import (
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/single_threaded"
)
// Be sure to close pools/instances when you're done with them.
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
// Init the PDFium library and return the instance to open documents.
pool = single_threaded.Init(single_threaded.Config{})
var err error
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
This package has to be named main to make it available as a binary. The plugin system will use this to start new PDFium workers. Example:
pdfium/worker/main.go
package main
import (
"github.com/klippa-app/go-pdfium/multi_threaded/worker"
)
func main() {
worker.StartWorker(nil)
}
To actually start workers, you will have to init the PDFium library somewhere, this also allows you to dynamically start
workers when needed. The best location to add this is in the init() of a package that is going to call the PDFium
library. Example:
pdfium/renderer/renderer.go
package renderer
import (
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/multi_threaded"
)
// Be sure to close pools/instances when you're done with them.
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
// Init the PDFium library and return the instance to open documents.
// You can tweak these configs to your need. Be aware that workers can use quite some memory.
pool = multi_threaded.Init(multi_threaded.Config{
MinIdle: 1, // Makes sure that at least x workers are always available
MaxIdle: 1, // Makes sure that at most x workers are ever available
MaxTotal: 1, // Maxium amount of workers in total, allows the amount of workers to grow when needed, items between total max and idle max are automatically cleaned up, while idle workers are kept alive so they can be used directly.
Command: multi_threaded.Command{
BinPath: "go", // Only do this while developing, on production put the actual binary path in here. You should not want the Go runtime on production.
Args: []string{"run", "pdfium/worker/main.go"}, // This is a reference to the worker package, this can be left empty when using a direct binary path.
},
})
var err error
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
package renderer
import (
"os"
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the single/multi-threaded init() here.
func main() {
filePath := "example.pdf"
pageCount, err := getPageCount(filePath)
if err != nil {
log.Fatal(err)
}
log.Printf("The PDF %s has %d pages", filePath, pageCount)
}
func getPageCount(filePath string) (int, error) {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return 0, err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return 0, err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
pageCount, err := instance.FPDF_GetPageCount(&requests.FPDF_GetPageCount{
Document: doc,
})
if err != nil {
return 0, err
}
return pageCount.PageCount, nil
}
package renderer
import (
"image/png"
"log"
"os"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the single/multi-threaded init() here.
func main() {
filePath := "example.pdf"
output := "example.pdf.png"
err := renderPage(filePath, 1, output)
if err != nil {
log.Fatal(err)
}
}
func renderPage(filePath string, page int, output string) error {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
// Render the page in DPI 200.
pageRender, err := instance.RenderPageInDPI(&requests.RenderPageInDPI{
DPI: 200, // The DPI to render the page in.
Page: requests.Page{
ByIndex: &requests.PageByIndex{
Document: doc,
Index: 0,
},
}, // The page to render, 0-indexed.
})
if err != nil {
return err
}
// Write the output to a file.
f, err := os.Create(output)
if err != nil {
return err
}
defer f.Close()
err = png.Encode(f, pageRender.Image)
if err != nil {
return err
}
return nil
}
Some newer API's by PDFium are marked as experimental. We do have support for these functions, but because they are prone to change we do not compile with support for it by default. This is to keep support for most PDFium versions by default.
If you call any API methods marked as experimental, the call will result in an error
Adding support for the experimental API is quite easy. All you have to do is give the build tag pdfium_experimental
when running go build or go run, like this:
go run -tags pdfium_experimental {package}/{file.go} or go build -tags pdfium_experimental {package}/{file.go}.
We actively monitor PDFium API additions/changes/deletions and apply them in the code-base. When API methods become non-experimental, we will make them available in the default configuration.
Recently we have added support for a non-cgo implementation using WebAssembly, we use the Wazero runtime for running WebAssembly within Go. The comes with quite some advantages:
go:embed, we can embed the WebAssembly binary inside the Go binaryOf course there are also some disadvantages:
Please be aware that Wazero comes with the Apache License 2.0 license.
PDFium is compiled to 32 bit WebAssembly, which puts a ceiling on how large a single render can be. This limit does not exist in the cgo implementation, which is 64 bit.
The bitmap that a render writes into has to be allocated in one block inside the WebAssembly instance, and that allocation can not be larger than 2 GB (the maximum of a signed 32 bit integer). A bitmap is always 4 bytes per pixel, so the ceiling works out to around 536 megapixels:
| Aspect ratio | Roughly the largest render |
|---|---|
| Square | 23000 x 23000 |
| A4 (portrait) | 19400 x 27500 |
| 16:9 | 30900 x 17400 |
In practice you will hit the limit a bit earlier, because the document, the fonts and the images that PDFium loads share the same memory. The instance as a whole is also limited to 4 GB, which is the maximum that Wazero allows for a 32 bit WebAssembly module.
A render that does not fit returns an error, it never returns a partial or an empty image:
could not create a bitmap to render into, the image to render is most likely too large when the allocation failedthe image to render is too large when the size of the bitmap does not fit in a 32 bit integer at allIf you need a higher resolution than this allows, you have a few options:
Crop on the render request. The limit then applies to the
region you asked for instead of to the full page, which makes the resolution of that region effectively unbounded. It
is also much faster and uses far less memory, because the bitmap only has to hold the region.Crop and stitch them together. Regions that share an edge in points also share that
edge in pixels, so the tiles fit together without a seam.Because you can tell Wazero which folders have to be mounted in WebAssembly, you have full control over the filesystem.
By default, go-pdfium will mount the full root disk in Wazero on non-Windows environments. On Windows environments, go-pdfium will get the volume of the current working directory and mount that as the root.
You can change this behaviour by overwriting FSConfig in the pool setup.
All paths given to go-pdfium in WebAssembly mode have to be in POSIX style and have to be absolute, so for
example: /home/user/Downloads/file.pdf. If you have mounted /home/user/on the root, then the path you would have to
give is /Downloads/file.pdf, this is the same on Windows, so no backward slashes or volume names in paths.
You can set your own mounts by overwriting FSConfig in the pool setup.
The examples below can also be found in the examples folder.
To start go-pdfium workers, you will have to init the go-pdfium worker pool somewhere, this also allows you to
dynamically start
workers when needed. The best location to add this is in the init() of a package that is going to call the PDFium
library. Example:
pdfium/renderer/renderer.go
package renderer
import (
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/webassembly"
)
// Be sure to close pools/instances when you're done with them.
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
// Init the PDFium library and return the instance to open documents.
// You can tweak these configs to your need. Be aware that workers can use quite some memory.
pool, err = webassembly.Init(webassembly.Config{
MinIdle: 1, // Makes sure that at least x workers are always available
MaxIdle: 1, // Makes sure that at most x workers are ever available
MaxTotal: 1, // Maxium amount of workers in total, allows the amount of workers to grow when needed, items between total max and idle max are automatically cleaned up, while idle workers are kept alive so they can be used directly.
})
var err error
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
package renderer
import (
"os"
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the webassembly init() here.
func main() {
filePath := "example.pdf"
pageCount, err := getPageCount(filePath)
if err != nil {
log.Fatal(err)
}
log.Printf("The PDF %s has %d pages", filePath, pageCount)
}
func getPageCount(filePath string) (int, error) {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return 0, err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return 0, err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
pageCount, err := instance.FPDF_GetPageCount(&requests.FPDF_GetPageCount{
Document: doc.Document,
})
if err != nil {
return 0, err
}
return pageCount.PageCount, nil
}
package renderer
import (
"image/png"
"log"
"os"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the webassembly init() here.
func main() {
filePath := "example.pdf"
output := "example.pdf.png"
err := renderPage(filePath, 1, output)
if err != nil {
log.Fatal(err)
}
}
func renderPage(filePath string, page int, output string) error {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
// Render the page in DPI 200.
pageRender, err := instance.RenderPageInDPI(&requests.RenderPageInDPI{
DPI: 200, // The DPI to render the page in.
Page: requests.Page{
ByIndex: &requests.PageByIndex{
Document: doc.Document,
Index: 0,
},
}, // The page to render, 0-indexed.
})
if err != nil {
return err
}
// The Render* methods return a cleanup function that has to be called when
// using webassembly to make sure resources are cleaned up. Do this after
// you are done with the returned image object.
defer pageRender.Cleanup()
// Write the output to a file.
f, err := os.Create(output)
if err != nil {
return err
}
defer f.Close()
err = png.Encode(f, pageRender.Result.Image)
if err != nil {
return err
}
return nil
}
Some newer API's by PDFium are marked as experimental. The WebAssembly build has support for all of them.
We actively monitor PDFium API additions/changes/deletions and apply them in the code-base.
The WebAssembly build will always be the latest PDFium version that we added support for.
Next to wazero there is an experimental backend that runs the same embedded PDFium WebAssembly module with the
Wago runtime, a pure Go ahead-of-time compiler. It lives in the
experimental/wago package and exposes the same pool API as the webassembly package, so switching between the two
is a matter of changing the import and the Init call:
package renderer
import (
"log"
"time"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/experimental/wago"
)
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
var err error
pool, err = wago.Init(wago.Config{
MinIdle: 1,
MaxIdle: 1,
MaxTotal: 1,
})
if err != nil {
log.Fatal(err)
}
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
Wago compiles the module about ten times faster than wazero's compiler and runs it a bit faster, but it is still in beta and this backend is experimental for the following reasons:
v0.1.0-beta.9 could not compile the PDFium module on arm64 and
miscompiled parts of it on both architectures. Those bugs were fixed on wago's main branch in September 2026 and
go-pdfium pins a commit that passes the complete go-pdfium test suite on amd64 and arm64. Do not downgrade the
dependency.memory.fill clobbers a register that is still in use, which PDFium hits
in std::fill_n on a std::vector<bool>. In a sequence of 1,000 real world documents rendered on one long running
instance, 2 documents came out differently from wazero because of it (a different sequence of 5,000 documents had
no mismatch). linux/amd64 is not affected. This has been reported to wago with a 30 line reproducer.Filesystem access works through Wago's WASI plugin. Use Mounts in the config to choose which host directories PDFium
can see, by default the root of the disk is mounted read/write like in the wazero backend. All paths have to be
absolute POSIX paths inside a mount. Kill only interrupts a call that is running inside PDFium when
InterruptibleCalls is enabled in the config, which costs a little on every call.
Please be aware that Wago and its WASI plugin come with the Apache License 2.0 license.
There is also an experimental backend on the wazy runtime, a pure Go runtime
derived from wazero that keeps the same API and adds WebAssembly 3.0 and component model support. It lives in the
experimental/wazy package, takes the same configuration as the webassembly package (with wazy's RuntimeConfig
and FSConfig types), runs the complete go-pdfium test suite with the same results as wazero on amd64 and arm64, and
also has an interpreter for other platforms:
pool, err := wazy.Init(wazy.Config{
MinIdle: 1,
MaxIdle: 1,
MaxTotal: 1,
})
It is marked experimental because wazy itself is young and its API may still change. Like wazero it comes with the
Apache License 2.0 license. A comparison of the runtimes, including a 5,000 document real world corpus, is in
experimental/BENCHMARKS.md.
Both wazero and wazy compile the module with a single goroutine by default, which takes about a second on a fast
machine. Both can compile in parallel when the context passed as Context in the config carries the number of
workers: experimental.WithCompilationWorkers(ctx, n) from wazero, or api.WithCompilationWorkers(ctx, n) from wazy.
With four workers the pool starts about three times faster.
io.ReadSeeker and io.WriterDocument loading allows you to load a document with a io.ReadSeeker. Please be aware that this only works efficiently
when using the single-threaded or WebAssembly usage, as that lives in the same process. For multi-threaded usage this
will just load in
the complete file and pass the bytes through the gRPC interface.
Document/image saving allows you to save using a io.Writer. Please be aware this only works when using the
single-threaded or WebAssembly usage. It's not possible to encode the io.Writer with gRPC. Or share it between
processes for that
matter.
By default, the CGO implementation encodes JPEG images with the image/jpeg package that comes with Go to make
distribution as simple as possible. However, this package is quite slow compared to native libraries like libjpeg and
libjpeg-turbo. You can enable the usage of libjpeg-turbo with the build tag pdfium_use_turbojpeg, this requires the
package libturbojpeg-dev to be installed during build time and the libturbojpeg package during build time and
runtime.
Speed improvements that can be expected are significant, for example: on a simple PDF the full process of rendering a
page is 3x as fast compared to a build without libjpeg-turbo. On the CGO implementation, progressive JPEG output
(Progressive in RenderToFile) requires this build tag; without it the option is ignored and a baseline JPEG is
written.
The WebAssembly implementation does not need the build tag or any native library. The bundled PDFium module contains libjpeg-turbo compiled to WebAssembly, including its SIMD kernels, and exports its encoder. Rendered pages are encoded to JPEG inside the WebAssembly module directly from the rendered bitmap, so the pixels are never copied out into Go first, and progressive JPEG output is supported out of the box. When a custom module without the encoder export is used, or for image types the encoder cannot consume directly, the implementation falls back to the same encoder the CGO implementation uses.
We offer an API stability promise with semantic versioning. In other words, we promise to not break any exported function signature without incrementing the major version. New features and behaviors happen with a minor version increment, e.g. 1.0.11 to 1.1.0. We also fix bugs or change internal details with a patch version, e.g. 1.0.0 to 1.0.1. Upgrades of the supported PDFium version will cause a minor version update.
This project will support the last 2 version of Go, this means that if the last version of Go is 1.24, our go.mod
will be set to Go 1.23, and our CI tests will be run on Go 1.23 and 1.24. This is in line with Go's
Release Policy.
It won't mean that the library won't work with older versions of Go, but it will tell you what to expect of the supported Go versions. If we change the supported Go versions, we will make that a minor version upgrade. This policy allows you to not be forced to the latest Go version for a pretty long time, but it still allows us to use new language features in a pretty reasonable time-frame.
Founded in 2015, Klippa's goal is to digitize & automate administrative processes with modern technologies. We help clients enhance the effectiveness of their organization by using machine learning and OCR. Since 2015, more than a thousand happy clients have used Klippa's software solutions. Klippa currently has an international team of 50 people, with offices in Groningen, Amsterdam and Brasov.
The MIT License (MIT)
3,104 followers · starred Feb 2025
2,487 followers · starred Aug 2024
572 followers · starred Jan 2025
617 followers · starred Jul 2024
Go
99.8%
:rocket: Easy to use PDF library using Go and PDFium :rocket:
A fast, multi-threaded and easy to use PDF library for Go applications.
FPDF_SetPrintMode, FPDF_RenderPage)image.Image using either DPI or pixel sizeimage.RGBA, the default) or in grayscale (image.Gray, using
ImageFormat: requests.RenderImageFormatGrayscale), the result is in the RenderedImage response fieldThis project uses the PDFium C++ library by Google (https://pdfium.googlesource.com/pdfium/) to process the PDF documents. Therefor this project could also be called a binding.
Please be aware that PDFium comes with the Apache License 2.0 license.
Since PDFium is not a multithreaded C++ library, we can not directly make it multithreaded by calling it from Go's subroutines.
To solve this, this library has 3 different implementations that you can use to call PDFium:
Both single_threaded and multi_threaded requires PDFium to be installed on your machine for your platform during
compilation and runtime, it also requires CGO to work on the platform you're compiling to.
webassembly does not need any external dependencies and also does not require CGO to work. However, Wazero currently
only has compiler support for amd64 and arm64, meaning it will be using the interpreter on other architectures which
will be much, much slower.
All implementations use exactly the same interface, so there won't be any code changes for you to switch between them.
All implementations are thread/subroutine safe, this has been guaranteed by locking the instance that's doing your work while it's doing PDFium operations. New operations will wait until the lock becomes available again.
Be aware that PDFium could use quite some memory depending on the size of the PDF and the requests that you do, so be aware of the amount of workers that you configure.
Single-threading in CGO works by directly calling the PDFium library from the same process. Single-threaded might be preferred if the caller is managing the workers themselves and does not want the overhead of another process. Be aware that since PDFium is C++, we can't handle segfaults caused by PDFium, which may cause your process to be killed. So using this library in the multi-threaded way, with only 1 worker, can still have some benefits, since it can automatically recover from things like segfaults.
For CGO we have implemented multi-threading using HashiCorp's Go Plugin System, which allows us to launch separate PDFium worker processes, and then route the requests through the different workers. This also makes it a bit more safe to use PDFium, as it's less likely to segfaults or corrupt your main Go application. The Plugin system provides the communication between the processes using gRPC, however, when implementing this library, you won't really see anything of that. From the outside it will look like normal Go code. The inter-process communication does come with a cost as it has to serialize/deserialize input/output as it moves between the main process and the PDFium workers.
To use this Go library, you will need the actual PDFium library to run it and have it available through pkgconfig.
You can try to compile PDFium yourself, but you can also use pre-compiled binaries, for example from: https://github.com/bblanchon/pdfium-binaries/releases
If you use a pre-compiled library, make sure to extract it somewhere logical, for example /opt/pdfium.
Create/edit file /usr/lib/pkgconfig/pdfium.pc
prefix={path}
libdir={path}/lib
includedir={path}/include
Name: PDFium
Description: PDFium
Version: 8057
Requires:
Libs: -L${libdir} -lpdfium
Cflags: -I${includedir}
Replace {path} with the path you extracted/compiled pdfium in.
Make sure you extend your library path when running:
export LD_LIBRARY_PATH={path}/lib
You can do this globally or just in your editor.
this can globally be done on ubuntu by editing ~/.profile
and adding the line in this file. reloading for bash can be done by relogging or running source ~/.profile can be used
to test the change for a terminal
To get started, make sure that you create a separate package in your application that will start the worker.
The examples below can also be found in the examples folder.
For single threaded implementations we just have to initialize the library.
pdfium/renderer/renderer.go
package renderer
import (
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/single_threaded"
)
// Be sure to close pools/instances when you're done with them.
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
// Init the PDFium library and return the instance to open documents.
pool = single_threaded.Init(single_threaded.Config{})
var err error
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
This package has to be named main to make it available as a binary. The plugin system will use this to start new PDFium workers. Example:
pdfium/worker/main.go
package main
import (
"github.com/klippa-app/go-pdfium/multi_threaded/worker"
)
func main() {
worker.StartWorker(nil)
}
To actually start workers, you will have to init the PDFium library somewhere, this also allows you to dynamically start
workers when needed. The best location to add this is in the init() of a package that is going to call the PDFium
library. Example:
pdfium/renderer/renderer.go
package renderer
import (
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/multi_threaded"
)
// Be sure to close pools/instances when you're done with them.
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
// Init the PDFium library and return the instance to open documents.
// You can tweak these configs to your need. Be aware that workers can use quite some memory.
pool = multi_threaded.Init(multi_threaded.Config{
MinIdle: 1, // Makes sure that at least x workers are always available
MaxIdle: 1, // Makes sure that at most x workers are ever available
MaxTotal: 1, // Maxium amount of workers in total, allows the amount of workers to grow when needed, items between total max and idle max are automatically cleaned up, while idle workers are kept alive so they can be used directly.
Command: multi_threaded.Command{
BinPath: "go", // Only do this while developing, on production put the actual binary path in here. You should not want the Go runtime on production.
Args: []string{"run", "pdfium/worker/main.go"}, // This is a reference to the worker package, this can be left empty when using a direct binary path.
},
})
var err error
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
package renderer
import (
"os"
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the single/multi-threaded init() here.
func main() {
filePath := "example.pdf"
pageCount, err := getPageCount(filePath)
if err != nil {
log.Fatal(err)
}
log.Printf("The PDF %s has %d pages", filePath, pageCount)
}
func getPageCount(filePath string) (int, error) {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return 0, err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return 0, err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
pageCount, err := instance.FPDF_GetPageCount(&requests.FPDF_GetPageCount{
Document: doc,
})
if err != nil {
return 0, err
}
return pageCount.PageCount, nil
}
package renderer
import (
"image/png"
"log"
"os"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the single/multi-threaded init() here.
func main() {
filePath := "example.pdf"
output := "example.pdf.png"
err := renderPage(filePath, 1, output)
if err != nil {
log.Fatal(err)
}
}
func renderPage(filePath string, page int, output string) error {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
// Render the page in DPI 200.
pageRender, err := instance.RenderPageInDPI(&requests.RenderPageInDPI{
DPI: 200, // The DPI to render the page in.
Page: requests.Page{
ByIndex: &requests.PageByIndex{
Document: doc,
Index: 0,
},
}, // The page to render, 0-indexed.
})
if err != nil {
return err
}
// Write the output to a file.
f, err := os.Create(output)
if err != nil {
return err
}
defer f.Close()
err = png.Encode(f, pageRender.Image)
if err != nil {
return err
}
return nil
}
Some newer API's by PDFium are marked as experimental. We do have support for these functions, but because they are prone to change we do not compile with support for it by default. This is to keep support for most PDFium versions by default.
If you call any API methods marked as experimental, the call will result in an error
Adding support for the experimental API is quite easy. All you have to do is give the build tag pdfium_experimental
when running go build or go run, like this:
go run -tags pdfium_experimental {package}/{file.go} or go build -tags pdfium_experimental {package}/{file.go}.
We actively monitor PDFium API additions/changes/deletions and apply them in the code-base. When API methods become non-experimental, we will make them available in the default configuration.
Recently we have added support for a non-cgo implementation using WebAssembly, we use the Wazero runtime for running WebAssembly within Go. The comes with quite some advantages:
go:embed, we can embed the WebAssembly binary inside the Go binaryOf course there are also some disadvantages:
Please be aware that Wazero comes with the Apache License 2.0 license.
PDFium is compiled to 32 bit WebAssembly, which puts a ceiling on how large a single render can be. This limit does not exist in the cgo implementation, which is 64 bit.
The bitmap that a render writes into has to be allocated in one block inside the WebAssembly instance, and that allocation can not be larger than 2 GB (the maximum of a signed 32 bit integer). A bitmap is always 4 bytes per pixel, so the ceiling works out to around 536 megapixels:
| Aspect ratio | Roughly the largest render |
|---|---|
| Square | 23000 x 23000 |
| A4 (portrait) | 19400 x 27500 |
| 16:9 | 30900 x 17400 |
In practice you will hit the limit a bit earlier, because the document, the fonts and the images that PDFium loads share the same memory. The instance as a whole is also limited to 4 GB, which is the maximum that Wazero allows for a 32 bit WebAssembly module.
A render that does not fit returns an error, it never returns a partial or an empty image:
could not create a bitmap to render into, the image to render is most likely too large when the allocation failedthe image to render is too large when the size of the bitmap does not fit in a 32 bit integer at allIf you need a higher resolution than this allows, you have a few options:
Crop on the render request. The limit then applies to the
region you asked for instead of to the full page, which makes the resolution of that region effectively unbounded. It
is also much faster and uses far less memory, because the bitmap only has to hold the region.Crop and stitch them together. Regions that share an edge in points also share that
edge in pixels, so the tiles fit together without a seam.Because you can tell Wazero which folders have to be mounted in WebAssembly, you have full control over the filesystem.
By default, go-pdfium will mount the full root disk in Wazero on non-Windows environments. On Windows environments, go-pdfium will get the volume of the current working directory and mount that as the root.
You can change this behaviour by overwriting FSConfig in the pool setup.
All paths given to go-pdfium in WebAssembly mode have to be in POSIX style and have to be absolute, so for
example: /home/user/Downloads/file.pdf. If you have mounted /home/user/on the root, then the path you would have to
give is /Downloads/file.pdf, this is the same on Windows, so no backward slashes or volume names in paths.
You can set your own mounts by overwriting FSConfig in the pool setup.
The examples below can also be found in the examples folder.
To start go-pdfium workers, you will have to init the go-pdfium worker pool somewhere, this also allows you to
dynamically start
workers when needed. The best location to add this is in the init() of a package that is going to call the PDFium
library. Example:
pdfium/renderer/renderer.go
package renderer
import (
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/webassembly"
)
// Be sure to close pools/instances when you're done with them.
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
// Init the PDFium library and return the instance to open documents.
// You can tweak these configs to your need. Be aware that workers can use quite some memory.
pool, err = webassembly.Init(webassembly.Config{
MinIdle: 1, // Makes sure that at least x workers are always available
MaxIdle: 1, // Makes sure that at most x workers are ever available
MaxTotal: 1, // Maxium amount of workers in total, allows the amount of workers to grow when needed, items between total max and idle max are automatically cleaned up, while idle workers are kept alive so they can be used directly.
})
var err error
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
package renderer
import (
"os"
"log"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the webassembly init() here.
func main() {
filePath := "example.pdf"
pageCount, err := getPageCount(filePath)
if err != nil {
log.Fatal(err)
}
log.Printf("The PDF %s has %d pages", filePath, pageCount)
}
func getPageCount(filePath string) (int, error) {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return 0, err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return 0, err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
pageCount, err := instance.FPDF_GetPageCount(&requests.FPDF_GetPageCount{
Document: doc.Document,
})
if err != nil {
return 0, err
}
return pageCount.PageCount, nil
}
package renderer
import (
"image/png"
"log"
"os"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/requests"
)
// Insert the webassembly init() here.
func main() {
filePath := "example.pdf"
output := "example.pdf.png"
err := renderPage(filePath, 1, output)
if err != nil {
log.Fatal(err)
}
}
func renderPage(filePath string, page int, output string) error {
// Load the PDF file into a byte array.
pdfBytes, err := os.ReadFile(filePath)
if err != nil {
return err
}
// Open the PDF using PDFium (and claim a worker)
doc, err := instance.OpenDocument(&requests.OpenDocument{
File: &pdfBytes,
})
if err != nil {
return err
}
// Always close the document, this will release its resources.
defer instance.FPDF_CloseDocument(&requests.FPDF_CloseDocument{
Document: doc.Document,
})
// Render the page in DPI 200.
pageRender, err := instance.RenderPageInDPI(&requests.RenderPageInDPI{
DPI: 200, // The DPI to render the page in.
Page: requests.Page{
ByIndex: &requests.PageByIndex{
Document: doc.Document,
Index: 0,
},
}, // The page to render, 0-indexed.
})
if err != nil {
return err
}
// The Render* methods return a cleanup function that has to be called when
// using webassembly to make sure resources are cleaned up. Do this after
// you are done with the returned image object.
defer pageRender.Cleanup()
// Write the output to a file.
f, err := os.Create(output)
if err != nil {
return err
}
defer f.Close()
err = png.Encode(f, pageRender.Result.Image)
if err != nil {
return err
}
return nil
}
Some newer API's by PDFium are marked as experimental. The WebAssembly build has support for all of them.
We actively monitor PDFium API additions/changes/deletions and apply them in the code-base.
The WebAssembly build will always be the latest PDFium version that we added support for.
Next to wazero there is an experimental backend that runs the same embedded PDFium WebAssembly module with the
Wago runtime, a pure Go ahead-of-time compiler. It lives in the
experimental/wago package and exposes the same pool API as the webassembly package, so switching between the two
is a matter of changing the import and the Init call:
package renderer
import (
"log"
"time"
"github.com/klippa-app/go-pdfium"
"github.com/klippa-app/go-pdfium/experimental/wago"
)
var pool pdfium.Pool
var instance pdfium.Pdfium
func init() {
var err error
pool, err = wago.Init(wago.Config{
MinIdle: 1,
MaxIdle: 1,
MaxTotal: 1,
})
if err != nil {
log.Fatal(err)
}
instance, err = pool.GetInstance(time.Second * 30)
if err != nil {
log.Fatal(err)
}
}
Wago compiles the module about ten times faster than wazero's compiler and runs it a bit faster, but it is still in beta and this backend is experimental for the following reasons:
v0.1.0-beta.9 could not compile the PDFium module on arm64 and
miscompiled parts of it on both architectures. Those bugs were fixed on wago's main branch in September 2026 and
go-pdfium pins a commit that passes the complete go-pdfium test suite on amd64 and arm64. Do not downgrade the
dependency.memory.fill clobbers a register that is still in use, which PDFium hits
in std::fill_n on a std::vector<bool>. In a sequence of 1,000 real world documents rendered on one long running
instance, 2 documents came out differently from wazero because of it (a different sequence of 5,000 documents had
no mismatch). linux/amd64 is not affected. This has been reported to wago with a 30 line reproducer.Filesystem access works through Wago's WASI plugin. Use Mounts in the config to choose which host directories PDFium
can see, by default the root of the disk is mounted read/write like in the wazero backend. All paths have to be
absolute POSIX paths inside a mount. Kill only interrupts a call that is running inside PDFium when
InterruptibleCalls is enabled in the config, which costs a little on every call.
Please be aware that Wago and its WASI plugin come with the Apache License 2.0 license.
There is also an experimental backend on the wazy runtime, a pure Go runtime
derived from wazero that keeps the same API and adds WebAssembly 3.0 and component model support. It lives in the
experimental/wazy package, takes the same configuration as the webassembly package (with wazy's RuntimeConfig
and FSConfig types), runs the complete go-pdfium test suite with the same results as wazero on amd64 and arm64, and
also has an interpreter for other platforms:
pool, err := wazy.Init(wazy.Config{
MinIdle: 1,
MaxIdle: 1,
MaxTotal: 1,
})
It is marked experimental because wazy itself is young and its API may still change. Like wazero it comes with the
Apache License 2.0 license. A comparison of the runtimes, including a 5,000 document real world corpus, is in
experimental/BENCHMARKS.md.
Both wazero and wazy compile the module with a single goroutine by default, which takes about a second on a fast
machine. Both can compile in parallel when the context passed as Context in the config carries the number of
workers: experimental.WithCompilationWorkers(ctx, n) from wazero, or api.WithCompilationWorkers(ctx, n) from wazy.
With four workers the pool starts about three times faster.
io.ReadSeeker and io.WriterDocument loading allows you to load a document with a io.ReadSeeker. Please be aware that this only works efficiently
when using the single-threaded or WebAssembly usage, as that lives in the same process. For multi-threaded usage this
will just load in
the complete file and pass the bytes through the gRPC interface.
Document/image saving allows you to save using a io.Writer. Please be aware this only works when using the
single-threaded or WebAssembly usage. It's not possible to encode the io.Writer with gRPC. Or share it between
processes for that
matter.
By default, the CGO implementation encodes JPEG images with the image/jpeg package that comes with Go to make
distribution as simple as possible. However, this package is quite slow compared to native libraries like libjpeg and
libjpeg-turbo. You can enable the usage of libjpeg-turbo with the build tag pdfium_use_turbojpeg, this requires the
package libturbojpeg-dev to be installed during build time and the libturbojpeg package during build time and
runtime.
Speed improvements that can be expected are significant, for example: on a simple PDF the full process of rendering a
page is 3x as fast compared to a build without libjpeg-turbo. On the CGO implementation, progressive JPEG output
(Progressive in RenderToFile) requires this build tag; without it the option is ignored and a baseline JPEG is
written.
The WebAssembly implementation does not need the build tag or any native library. The bundled PDFium module contains libjpeg-turbo compiled to WebAssembly, including its SIMD kernels, and exports its encoder. Rendered pages are encoded to JPEG inside the WebAssembly module directly from the rendered bitmap, so the pixels are never copied out into Go first, and progressive JPEG output is supported out of the box. When a custom module without the encoder export is used, or for image types the encoder cannot consume directly, the implementation falls back to the same encoder the CGO implementation uses.
We offer an API stability promise with semantic versioning. In other words, we promise to not break any exported function signature without incrementing the major version. New features and behaviors happen with a minor version increment, e.g. 1.0.11 to 1.1.0. We also fix bugs or change internal details with a patch version, e.g. 1.0.0 to 1.0.1. Upgrades of the supported PDFium version will cause a minor version update.
This project will support the last 2 version of Go, this means that if the last version of Go is 1.24, our go.mod
will be set to Go 1.23, and our CI tests will be run on Go 1.23 and 1.24. This is in line with Go's
Release Policy.
It won't mean that the library won't work with older versions of Go, but it will tell you what to expect of the supported Go versions. If we change the supported Go versions, we will make that a minor version upgrade. This policy allows you to not be forced to the latest Go version for a pretty long time, but it still allows us to use new language features in a pretty reasonable time-frame.
Founded in 2015, Klippa's goal is to digitize & automate administrative processes with modern technologies. We help clients enhance the effectiveness of their organization by using machine learning and OCR. Since 2015, more than a thousand happy clients have used Klippa's software solutions. Klippa currently has an international team of 50 people, with offices in Groningen, Amsterdam and Brasov.
The MIT License (MIT)
3,104 followers · starred Feb 2025
2,487 followers · starred Aug 2024
572 followers · starred Jan 2025
617 followers · starred Jul 2024
Go
99.8%