u5surf/ngx_vts

Rust

2

182 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Writing nginx Modules in Rust: a book on ngx-rust (r/rust)

I've been writing nginx modules in C for about 10 years, and I'm a collaborator on [nginx-module-vts](https://github.com/vozlt/nginx-module-vts). Day to day, I kept running into two things with C: the hassle of managing memory by hand, and how hard it is to pull module logic out into unit tests. I…

13

Oct 6, 2026

README

nginx-vts-rust

CI

A Rust implementation of nginx-module-vts for virtual host traffic status monitoring, built on top of the ngx-rust framework.

Status: experimental, but the cross-worker aggregation path, the vts_zone directive, and the /status endpoint are working end-to-end with nginx 1.31.

Architecture overview

┌─────────────────────────────────────────────────────────────────────┐
│ nginx master                                                        │
│                                                                     │
│  vts_zone main 1m;  ─►  ngx_shared_memory_add  ─►  shm_zone         │
│                                                       │             │
│                                                       ▼             │
│                                       vts_init_shm_zone (Rust)      │
│                                                       │             │
│                                                       ▼             │
│           ┌──────────────────────────────────────────────┐          │
│           │ slab pool (SlabPool: Allocator)              │          │
│           │   ┌─ VtsShared                               │          │
│           │   │    ├─ RwLock< RbTreeMap<                 │          │
│           │   │    │    NgxString<SlabPool>,             │          │
│           │   │    │    ServerCounters,                  │          │
│           │   │    │    SlabPool> >       (servers)      │          │
│           │   │    └─ RwLock< RbTreeMap<                 │          │
│           │   │         NgxString<SlabPool>,             │          │
│           │   │         UpstreamCounters,                │          │
│           │   │         SlabPool> >      (upstreams)     │          │
│           │   └─ shpool->data = &VtsShared (reload-safe) │          │
│           └──────────────────────────────────────────────┘          │
└──────────────────────────────┬──────────────────────────────────────┘
                               │  fork
              ┌────────────────┴────────────────┐
              ▼                                 ▼
   ┌──────────────────────┐          ┌──────────────────────┐
   │ worker 1             │          │ worker 2             │
   │  LOG_PHASE handler   │          │  LOG_PHASE handler   │
   │   └─► record_server  │          │   └─► record_server  │
   │   └─► record_upstr.  │ ──┬────► │   └─► record_upstr.  │
   │  /status handler     │   │      │  /status handler     │
   │   └─► snapshot_*     │   │      │   └─► snapshot_*     │
   └──────────────────────┘   │      └──────────────────────┘
                              │
              ngx::sync::RwLock guards each map independently:
              - writers (record_*) take .write()
              - readers (/status snapshot_*) take .read()

The shared state is two RbTreeMaps allocated from the slab pool — one keyed by server_name, one keyed by "upstream\0server" — wrapped in ngx::sync::RwLock for concurrent worker access. Capacity scales with the configured vts_zone size rather than being capped at compile time, and /status reads no longer block concurrent writers thanks to the reader-writer lock.

Keys are derived from nginx configuration (the matched server block's first server_name, the upstream block name) — never from the raw Host header — so attacker-controlled values cannot expand the key space.

When vts_zone is not declared the FFI transparently falls back to a process-local manager — this is how the unit tests exercise the data model without nginx.

Features

  • Cross-worker aggregation — every worker writes to the same slab table; /status returns totals across the whole nginx instance.
  • vts_zone directive — declares a real shared-memory zone (ngx_shared_memory_add) whose init callback creates two RbTreeMaps inside the slab pool from Rust.
  • Server-zone metrics keyed by the matched server block's first server_name (not the raw Host header), so the table can't be blown up by adversarial Host values.
  • Metric names compatible with nginx-module-vts — the Prometheus families, label names and label values follow the original module's output, so dashboards and alerts written against it keep working. See Compatibility with nginx-module-vts for what still differs.
  • Upstream metrics per (upstream, backend) peer — status-code class counters, bytes in/out, request and upstream response times.
  • Per-attempt upstream tracking — r->upstream_states is iterated so each retry attempt (e.g. 502 from peer A followed by 200 from peer B) contributes its own sample to the upstream counters, not just the final state.
  • Upstream and server-zone request-time histograms — classic Prometheus _bucket{le=...} / _sum / _count over a shared fixed 11-bucket layout (client_golang defaults), exposed as nginx_vts_upstream_response_duration_seconds_bucket{...} and nginx_vts_server_request_duration_seconds_bucket{...}. Both feed histogram_quantile(0.99, ...) for p50/p90/p99 panels per upstream peer and per vhost.
  • Cache hit/miss metrics per cache zone (proxy_cache_path keys_zone=NAME:SIZE) — counts of HIT, MISS, BYPASS, EXPIRED, STALE, UPDATING, REVALIDATED, SCARCE aggregated across workers, exposed as nginx_vts_cache_requests_total.
  • Cache size gauges per cache zone — proxy_cache_path max_size=… and current on-disk usage (sh->size × bsize) exposed as nginx_vts_cache_usage_bytes{cache_size="max"} and {cache_size="used"}.
  • Shared zone accounting — nginx_vts_main_shm_usage_bytes with shared="max_size", "used_size", "used_node" and "free_size", so a full zone can be told from an idle one. free_size comes from the slab's own page count and used_size is the rest of the zone, rather than the sum of node sizes the original reports, because the slab spends a whole page or slot per node and that sum reads low right up to the point where inserts start failing.
  • Accurate connection counters via the global ngx_stat_* atomics when nginx is built with --with-http_stub_status_module; reading/writing/waiting match what stub_status would report. Without that build flag the module falls back to a cycle-table walk and only the active total stays meaningful.
  • Subrequest- and /status-aware counting — the LOG_PHASE handler skips internal subrequests (auth_request, mirror, addition, …) and the module's own /status scrapes, so neither double-counts the per-vhost counters.
  • Prometheus text format at /status with the text/plain; version=0.0.4 Content-Type that Prometheus 3.x requires.
  • Reload-safe — nginx -s reload reuses the existing shared table, so counters survive a config reload.

Build

Requirements

  • Rust 1.85 or later (ngx-rust 0.5 uses edition 2024).
  • nginx source tree (any 1.24+ release; CI is pinned to 1.28.0).
  • A C compiler (cc / clang).
  • pcre2 and zlib headers for the nginx build.

Build the Rust cdylib

export NGINX_SOURCE_DIR=/path/to/nginx-source     # ngx-rust looks here
cargo build --release

Output: target/release/libngx_vts_rust.{so,dylib}.

Build nginx with the module

cd /path/to/nginx-source
auto/configure --prefix=/tmp/nginx-vts-test \
               --with-compat \
               --add-dynamic-module=/path/to/ngx_vts
make

This produces:

  • objs/nginx — the nginx binary (only needed if you don't already have one built from the same source).
  • objs/ngx_http_vts_module.so — the dynamic module you load from nginx.conf via load_module.

The repository's config script picks .dylib on macOS and .so on Linux automatically.

Quick start

Minimal nginx.conf that proxies through an upstream and exposes /status:

load_module modules/ngx_http_vts_module.so;

events {
    worker_connections 64;
}

http {
    vts_zone main 1m;

    upstream backend {
        server 127.0.0.1:18091;
        server 127.0.0.1:18092;
    }

    # Two local servers acting as the upstream peers.
    server { listen 18091; location / { return 200 "peer1\n"; } }
    server { listen 18092; location / { return 200 "peer2\n"; } }

    server {
        listen 18080;
        server_name example.test;

        location /         { proxy_pass http://backend; }
        location /status   { vts_status; allow 127.0.0.1; deny all; }
    }
}

Run it:

mkdir -p /tmp/nginx-vts-test/{conf,logs,modules}
cp objs/ngx_http_vts_module.so /tmp/nginx-vts-test/modules/
cp nginx.conf                  /tmp/nginx-vts-test/conf/
objs/nginx -p /tmp/nginx-vts-test -c conf/nginx.conf

Drive traffic and read the metrics:

$ seq 1 100 | xargs -P 8 -I{} curl -sS -o /dev/null http://127.0.0.1:18080/
$ curl -sS http://127.0.0.1:18080/status

Sample output (verbatim, after 105 proxied requests across 2 workers)

Trimmed where marked ….

# Prometheus Metrics:
# HELP nginx_vts_info Nginx info
# TYPE nginx_vts_info gauge
nginx_vts_info{hostname="…",module_version="0.1.0",version="1.31.6"} 1

# HELP nginx_vts_start_time_seconds Nginx start time
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791080808
…

# HELP nginx_vts_main_connections Nginx connections
# TYPE nginx_vts_main_connections gauge
nginx_vts_main_connections{status="accepted"} 10
nginx_vts_main_connections{status="active"} 10
…

# HELP nginx_vts_main_shm_usage_bytes Shared memory zone usage
# TYPE nginx_vts_main_shm_usage_bytes gauge
nginx_vts_main_shm_usage_bytes{shared="max_size"} 1048576
nginx_vts_main_shm_usage_bytes{shared="used_size"} 98304
nginx_vts_main_shm_usage_bytes{shared="used_node"} 4
nginx_vts_main_shm_usage_bytes{shared="free_size"} 950272

# HELP nginx_vts_server_bytes_total The request/response bytes
# TYPE nginx_vts_server_bytes_total counter
nginx_vts_server_bytes_total{host="example.test",direction="in"} 8190
nginx_vts_server_bytes_total{host="example.test",direction="out"} 16065
…

# HELP nginx_vts_server_requests_total The requests counter
# TYPE nginx_vts_server_requests_total counter
nginx_vts_server_requests_total{host="example.test",code="2xx"} 105
…
nginx_vts_server_requests_total{host="_",code="2xx"} 105
…
nginx_vts_server_requests_total{host="*",code="2xx"} 210
…

# HELP nginx_vts_upstream_requests_total The upstream requests counter
# TYPE nginx_vts_upstream_requests_total counter
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18091",code="2xx"} 53
…
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18092",code="2xx"} 52
…

# HELP nginx_vts_upstream_server_up Upstream server status (1=up, 0=down)
# TYPE nginx_vts_upstream_server_up gauge
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18091"} 1
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18092"} 1

Note that peer1 (53) + peer2 (52) = 105: both workers feed the same table, so /status shows the totals regardless of which worker happened to handle the request. host="_" is the two peer servers, which have no server_name and live in the same nginx here, so the host="*" row that sums every zone counts each request twice.

Compatibility with nginx-module-vts

The Prometheus output uses the original module's metric names, label names (host, code, upstream, backend, cache_zone, …) and label values, including the host="*" row that sums every server zone. What still differs:

  • Additions — process_start_time_seconds (see Persistence) and nginx_vts_upstream_server_up.
  • Values — used_size is what the slab has spent, not the sum of node sizes. nginx_vts_start_time_seconds is when the counters started from zero, which a reload does not change. The *_request_seconds and *_response_seconds averages are cumulative (sum / count), not the original's moving average.
  • Histograms — nginx_vts_server_request_duration_seconds and nginx_vts_upstream_response_duration_seconds are always emitted, with fixed buckets, and le is written 1 rather than 1.000. The original emits them only when histogram_buckets is set.
  • Not emitted yet — nginx_vts_server_cache_total, nginx_vts_cache_bytes_total, nginx_vts_upstream_request_duration_seconds, nginx_vts_status_code_requests_total and the nginx_vts_filter_* families.
  • Upstreams without a group — a proxy_pass straight to an address is reported under its own name rather than the original's upstream="::nogroups".

Directives

DirectiveContextArgsDescription
vts_zonehttpname sizeDeclare the shared-memory zone backing all counters. Minimum size is 1 MB; without this directive the module silently falls back to process-local counters (mainly useful for tests).
vts_statuslocation—Render the Prometheus text response at this location.
vts_upstream_statshttp, server, locationon | offAccepted for backward compatibility; currently a no-op (upstream stats are always collected when vts_zone is set).

Capacity

The shared state is two RbTreeMaps — one keyed by server_name, one keyed by the (upstream, server) pair — allocated inside the slab pool that backs the vts_zone. There is no compile-time slot cap: how many distinct keys you can track is bounded only by the slab pool size you configure with vts_zone <name> <size>.

Rough sizing rule of thumb: a 1m zone comfortably holds a few thousand server-zone keys plus a few thousand upstream pairs. Each entry is on the order of ~200 bytes for the counters plus the key length plus rbtree node overhead. Bump the size if you genuinely have more virtual hosts.

When a new key cannot be allocated (the slab pool is full), it is dropped silently and existing counters keep updating. There is also a defensive upper bound on key length (VTS_MAX_KEY_BYTES = 256) to keep misconfigured server_name directives from chewing up the pool.

Keys are derived from nginx configuration (the matched server block's first server_name, the upstream block name) — never from the raw Host header — so attacker-controlled values cannot expand the key space.

Persistence

Counters live only in shared memory; nothing is written to disk.

OperationCounters
nginx -s reloadKept — nginx reuses the zone
Reload after changing the vts_zone sizeReset to zero — a new zone is allocated
Stop / start, container restartReset to zero
Binary upgrade (USR2)Reset to zero — the new master does not inherit the zone

All *_total metrics are Prometheus counters, so query them with rate() / increase(), which detect and absorb resets. What a restart loses is only the increment since the last scrape.

When a zone is configured, /status also reports when the counters last started from zero, under the original module's name and under the name other collectors look for:

# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791039391
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1791039391

It is the time the zone was built, not the time a process started, so it stays put across a reload and moves forward exactly when the counters reset. process_start_time_seconds is unprefixed because that is the name other collectors look for to tell a reset from a series they have only just started watching.

OpenTelemetry Collector — scrape /status with the prometheus receiver and let the metric_start_time processor take the start time from this metric, before any batching:

receivers:
  prometheus:
    config:
      scrape_configs:
        - job_name: nginx-vts
          metrics_path: /status
          static_configs:
            - targets: ["localhost:80"]

processors:
  metric_start_time:
    strategy: start_time_metric

service:
  pipelines:
    metrics:
      receivers: [prometheus]
      processors: [metric_start_time]
      exporters: [otlp]   # Datadog, Mackerel, or any OTLP backend

Datadog Agent — the openmetrics check reads the same metric with use_process_start_time, so counters that started after the Agent did are counted from zero on the first scrape instead of being dropped:

instances:
  - openmetrics_endpoint: http://localhost/status
    namespace: nginx_vts
    metrics: [".*"]
    use_process_start_time: true

Mackerel — send OTLP through the Collector configuration above; Mackerel stores the counters as cumulative sums and its PromQL rate() / increase() absorb resets. Scraping with mackerel-plugin-prometheus-exporter instead posts every value as is, so a restart shows up as a drop to zero on the graph.

Development

Tests

NGINX_SOURCE_DIR=/path/to/nginx-source cargo test --lib

~75 unit tests cover the shared-table data model, the upstream tracker, the Prometheus formatter (per metric family), the cache statistics helpers, the LOG_PHASE-level FFI, and the rendered /status output via the process-local VTS_MANAGER fallback.

Lints

NGINX_SOURCE_DIR=/path/to/nginx-source cargo clippy --all-targets -- -D warnings
cargo fmt --all -- --check

What's not done yet

The list below tracks known gaps relative to the original nginx-module-vts. None of them block normal traffic monitoring.

Output and control

  • JSON / HTML / JSONP output formats — only Prometheus text is emitted.
  • /control API for reset/delete.
  • vts_dump directive (periodic on-disk dump for counter recovery across restarts). Not planned while output is Prometheus-only; see Persistence.

Filtering and limits

  • Filter zones (vhost_traffic_status_filter_by_set_key, _filter_by_host, _filter_max_node) — no dynamic key-based grouping yet.
  • Traffic limiting (vhost_traffic_status_limit_traffic, _limit_traffic_by_set_key) — the module is observation-only; it cannot rate-limit responses.

Metric coverage

  • Some of the original's families are not emitted yet; see Compatibility with nginx-module-vts.
  • Upstream peer state (down, weight, max_fails, fail_timeout, backup) is not yet read from the nginx upstream configuration.
  • Per-status-code counters (vhost_traffic_status_measure_status_codes) — only the 1xx/2xx/3xx/4xx/5xx class buckets are exposed.
  • Histogram bucket layout is fixed at the Prometheus client_golang defaults (5ms..10s, 11 buckets). There is no vts_histogram_buckets-style directive to customise the bounds.
  • Average method (vhost_traffic_status_average_method AMM / WMA) — averages are plain cumulative sum / count.
  • Embedded $vts_* variables for use in log_format / if — upstream module exposes ~20; we expose none.

License

Licensed under either of

at your option.

Contribution

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.

u5surf/ngx_vts

Rust

2

182 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Writing nginx Modules in Rust: a book on ngx-rust (r/rust)

I've been writing nginx modules in C for about 10 years, and I'm a collaborator on [nginx-module-vts](https://github.com/vozlt/nginx-module-vts). Day to day, I kept running into two things with C: the hassle of managing memory by hand, and how hard it is to pull module logic out into unit tests. I…

13

Oct 6, 2026

README

nginx-vts-rust

CI

A Rust implementation of nginx-module-vts for virtual host traffic status monitoring, built on top of the ngx-rust framework.

Status: experimental, but the cross-worker aggregation path, the vts_zone directive, and the /status endpoint are working end-to-end with nginx 1.31.

Architecture overview

┌─────────────────────────────────────────────────────────────────────┐
│ nginx master                                                        │
│                                                                     │
│  vts_zone main 1m;  ─►  ngx_shared_memory_add  ─►  shm_zone         │
│                                                       │             │
│                                                       ▼             │
│                                       vts_init_shm_zone (Rust)      │
│                                                       │             │
│                                                       ▼             │
│           ┌──────────────────────────────────────────────┐          │
│           │ slab pool (SlabPool: Allocator)              │          │
│           │   ┌─ VtsShared                               │          │
│           │   │    ├─ RwLock< RbTreeMap<                 │          │
│           │   │    │    NgxString<SlabPool>,             │          │
│           │   │    │    ServerCounters,                  │          │
│           │   │    │    SlabPool> >       (servers)      │          │
│           │   │    └─ RwLock< RbTreeMap<                 │          │
│           │   │         NgxString<SlabPool>,             │          │
│           │   │         UpstreamCounters,                │          │
│           │   │         SlabPool> >      (upstreams)     │          │
│           │   └─ shpool->data = &VtsShared (reload-safe) │          │
│           └──────────────────────────────────────────────┘          │
└──────────────────────────────┬──────────────────────────────────────┘
                               │  fork
              ┌────────────────┴────────────────┐
              ▼                                 ▼
   ┌──────────────────────┐          ┌──────────────────────┐
   │ worker 1             │          │ worker 2             │
   │  LOG_PHASE handler   │          │  LOG_PHASE handler   │
   │   └─► record_server  │          │   └─► record_server  │
   │   └─► record_upstr.  │ ──┬────► │   └─► record_upstr.  │
   │  /status handler     │   │      │  /status handler     │
   │   └─► snapshot_*     │   │      │   └─► snapshot_*     │
   └──────────────────────┘   │      └──────────────────────┘
                              │
              ngx::sync::RwLock guards each map independently:
              - writers (record_*) take .write()
              - readers (/status snapshot_*) take .read()

The shared state is two RbTreeMaps allocated from the slab pool — one keyed by server_name, one keyed by "upstream\0server" — wrapped in ngx::sync::RwLock for concurrent worker access. Capacity scales with the configured vts_zone size rather than being capped at compile time, and /status reads no longer block concurrent writers thanks to the reader-writer lock.

Keys are derived from nginx configuration (the matched server block's first server_name, the upstream block name) — never from the raw Host header — so attacker-controlled values cannot expand the key space.

When vts_zone is not declared the FFI transparently falls back to a process-local manager — this is how the unit tests exercise the data model without nginx.

Features

  • Cross-worker aggregation — every worker writes to the same slab table; /status returns totals across the whole nginx instance.
  • vts_zone directive — declares a real shared-memory zone (ngx_shared_memory_add) whose init callback creates two RbTreeMaps inside the slab pool from Rust.
  • Server-zone metrics keyed by the matched server block's first server_name (not the raw Host header), so the table can't be blown up by adversarial Host values.
  • Metric names compatible with nginx-module-vts — the Prometheus families, label names and label values follow the original module's output, so dashboards and alerts written against it keep working. See Compatibility with nginx-module-vts for what still differs.
  • Upstream metrics per (upstream, backend) peer — status-code class counters, bytes in/out, request and upstream response times.
  • Per-attempt upstream tracking — r->upstream_states is iterated so each retry attempt (e.g. 502 from peer A followed by 200 from peer B) contributes its own sample to the upstream counters, not just the final state.
  • Upstream and server-zone request-time histograms — classic Prometheus _bucket{le=...} / _sum / _count over a shared fixed 11-bucket layout (client_golang defaults), exposed as nginx_vts_upstream_response_duration_seconds_bucket{...} and nginx_vts_server_request_duration_seconds_bucket{...}. Both feed histogram_quantile(0.99, ...) for p50/p90/p99 panels per upstream peer and per vhost.
  • Cache hit/miss metrics per cache zone (proxy_cache_path keys_zone=NAME:SIZE) — counts of HIT, MISS, BYPASS, EXPIRED, STALE, UPDATING, REVALIDATED, SCARCE aggregated across workers, exposed as nginx_vts_cache_requests_total.
  • Cache size gauges per cache zone — proxy_cache_path max_size=… and current on-disk usage (sh->size × bsize) exposed as nginx_vts_cache_usage_bytes{cache_size="max"} and {cache_size="used"}.
  • Shared zone accounting — nginx_vts_main_shm_usage_bytes with shared="max_size", "used_size", "used_node" and "free_size", so a full zone can be told from an idle one. free_size comes from the slab's own page count and used_size is the rest of the zone, rather than the sum of node sizes the original reports, because the slab spends a whole page or slot per node and that sum reads low right up to the point where inserts start failing.
  • Accurate connection counters via the global ngx_stat_* atomics when nginx is built with --with-http_stub_status_module; reading/writing/waiting match what stub_status would report. Without that build flag the module falls back to a cycle-table walk and only the active total stays meaningful.
  • Subrequest- and /status-aware counting — the LOG_PHASE handler skips internal subrequests (auth_request, mirror, addition, …) and the module's own /status scrapes, so neither double-counts the per-vhost counters.
  • Prometheus text format at /status with the text/plain; version=0.0.4 Content-Type that Prometheus 3.x requires.
  • Reload-safe — nginx -s reload reuses the existing shared table, so counters survive a config reload.

Build

Requirements

  • Rust 1.85 or later (ngx-rust 0.5 uses edition 2024).
  • nginx source tree (any 1.24+ release; CI is pinned to 1.28.0).
  • A C compiler (cc / clang).
  • pcre2 and zlib headers for the nginx build.

Build the Rust cdylib

export NGINX_SOURCE_DIR=/path/to/nginx-source     # ngx-rust looks here
cargo build --release

Output: target/release/libngx_vts_rust.{so,dylib}.

Build nginx with the module

cd /path/to/nginx-source
auto/configure --prefix=/tmp/nginx-vts-test \
               --with-compat \
               --add-dynamic-module=/path/to/ngx_vts
make

This produces:

  • objs/nginx — the nginx binary (only needed if you don't already have one built from the same source).
  • objs/ngx_http_vts_module.so — the dynamic module you load from nginx.conf via load_module.

The repository's config script picks .dylib on macOS and .so on Linux automatically.

Quick start

Minimal nginx.conf that proxies through an upstream and exposes /status:

load_module modules/ngx_http_vts_module.so;

events {
    worker_connections 64;
}

http {
    vts_zone main 1m;

    upstream backend {
        server 127.0.0.1:18091;
        server 127.0.0.1:18092;
    }

    # Two local servers acting as the upstream peers.
    server { listen 18091; location / { return 200 "peer1\n"; } }
    server { listen 18092; location / { return 200 "peer2\n"; } }

    server {
        listen 18080;
        server_name example.test;

        location /         { proxy_pass http://backend; }
        location /status   { vts_status; allow 127.0.0.1; deny all; }
    }
}

Run it:

mkdir -p /tmp/nginx-vts-test/{conf,logs,modules}
cp objs/ngx_http_vts_module.so /tmp/nginx-vts-test/modules/
cp nginx.conf                  /tmp/nginx-vts-test/conf/
objs/nginx -p /tmp/nginx-vts-test -c conf/nginx.conf

Drive traffic and read the metrics:

$ seq 1 100 | xargs -P 8 -I{} curl -sS -o /dev/null http://127.0.0.1:18080/
$ curl -sS http://127.0.0.1:18080/status

Sample output (verbatim, after 105 proxied requests across 2 workers)

Trimmed where marked ….

# Prometheus Metrics:
# HELP nginx_vts_info Nginx info
# TYPE nginx_vts_info gauge
nginx_vts_info{hostname="…",module_version="0.1.0",version="1.31.6"} 1

# HELP nginx_vts_start_time_seconds Nginx start time
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791080808
…

# HELP nginx_vts_main_connections Nginx connections
# TYPE nginx_vts_main_connections gauge
nginx_vts_main_connections{status="accepted"} 10
nginx_vts_main_connections{status="active"} 10
…

# HELP nginx_vts_main_shm_usage_bytes Shared memory zone usage
# TYPE nginx_vts_main_shm_usage_bytes gauge
nginx_vts_main_shm_usage_bytes{shared="max_size"} 1048576
nginx_vts_main_shm_usage_bytes{shared="used_size"} 98304
nginx_vts_main_shm_usage_bytes{shared="used_node"} 4
nginx_vts_main_shm_usage_bytes{shared="free_size"} 950272

# HELP nginx_vts_server_bytes_total The request/response bytes
# TYPE nginx_vts_server_bytes_total counter
nginx_vts_server_bytes_total{host="example.test",direction="in"} 8190
nginx_vts_server_bytes_total{host="example.test",direction="out"} 16065
…

# HELP nginx_vts_server_requests_total The requests counter
# TYPE nginx_vts_server_requests_total counter
nginx_vts_server_requests_total{host="example.test",code="2xx"} 105
…
nginx_vts_server_requests_total{host="_",code="2xx"} 105
…
nginx_vts_server_requests_total{host="*",code="2xx"} 210
…

# HELP nginx_vts_upstream_requests_total The upstream requests counter
# TYPE nginx_vts_upstream_requests_total counter
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18091",code="2xx"} 53
…
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18092",code="2xx"} 52
…

# HELP nginx_vts_upstream_server_up Upstream server status (1=up, 0=down)
# TYPE nginx_vts_upstream_server_up gauge
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18091"} 1
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18092"} 1

Note that peer1 (53) + peer2 (52) = 105: both workers feed the same table, so /status shows the totals regardless of which worker happened to handle the request. host="_" is the two peer servers, which have no server_name and live in the same nginx here, so the host="*" row that sums every zone counts each request twice.

Compatibility with nginx-module-vts

The Prometheus output uses the original module's metric names, label names (host, code, upstream, backend, cache_zone, …) and label values, including the host="*" row that sums every server zone. What still differs:

  • Additions — process_start_time_seconds (see Persistence) and nginx_vts_upstream_server_up.
  • Values — used_size is what the slab has spent, not the sum of node sizes. nginx_vts_start_time_seconds is when the counters started from zero, which a reload does not change. The *_request_seconds and *_response_seconds averages are cumulative (sum / count), not the original's moving average.
  • Histograms — nginx_vts_server_request_duration_seconds and nginx_vts_upstream_response_duration_seconds are always emitted, with fixed buckets, and le is written 1 rather than 1.000. The original emits them only when histogram_buckets is set.
  • Not emitted yet — nginx_vts_server_cache_total, nginx_vts_cache_bytes_total, nginx_vts_upstream_request_duration_seconds, nginx_vts_status_code_requests_total and the nginx_vts_filter_* families.
  • Upstreams without a group — a proxy_pass straight to an address is reported under its own name rather than the original's upstream="::nogroups".

Directives

DirectiveContextArgsDescription
vts_zonehttpname sizeDeclare the shared-memory zone backing all counters. Minimum size is 1 MB; without this directive the module silently falls back to process-local counters (mainly useful for tests).
vts_statuslocation—Render the Prometheus text response at this location.
vts_upstream_statshttp, server, locationon | offAccepted for backward compatibility; currently a no-op (upstream stats are always collected when vts_zone is set).

Capacity

The shared state is two RbTreeMaps — one keyed by server_name, one keyed by the (upstream, server) pair — allocated inside the slab pool that backs the vts_zone. There is no compile-time slot cap: how many distinct keys you can track is bounded only by the slab pool size you configure with vts_zone <name> <size>.

Rough sizing rule of thumb: a 1m zone comfortably holds a few thousand server-zone keys plus a few thousand upstream pairs. Each entry is on the order of ~200 bytes for the counters plus the key length plus rbtree node overhead. Bump the size if you genuinely have more virtual hosts.

When a new key cannot be allocated (the slab pool is full), it is dropped silently and existing counters keep updating. There is also a defensive upper bound on key length (VTS_MAX_KEY_BYTES = 256) to keep misconfigured server_name directives from chewing up the pool.

Keys are derived from nginx configuration (the matched server block's first server_name, the upstream block name) — never from the raw Host header — so attacker-controlled values cannot expand the key space.

Persistence

Counters live only in shared memory; nothing is written to disk.

OperationCounters
nginx -s reloadKept — nginx reuses the zone
Reload after changing the vts_zone sizeReset to zero — a new zone is allocated
Stop / start, container restartReset to zero
Binary upgrade (USR2)Reset to zero — the new master does not inherit the zone

All *_total metrics are Prometheus counters, so query them with rate() / increase(), which detect and absorb resets. What a restart loses is only the increment since the last scrape.

When a zone is configured, /status also reports when the counters last started from zero, under the original module's name and under the name other collectors look for:

# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791039391
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1791039391

It is the time the zone was built, not the time a process started, so it stays put across a reload and moves forward exactly when the counters reset. process_start_time_seconds is unprefixed because that is the name other collectors look for to tell a reset from a series they have only just started watching.

OpenTelemetry Collector — scrape /status with the prometheus receiver and let the metric_start_time processor take the start time from this metric, before any batching:

receivers:
  prometheus:
    config:
      scrape_configs:
        - job_name: nginx-vts
          metrics_path: /status
          static_configs:
            - targets: ["localhost:80"]

processors:
  metric_start_time:
    strategy: start_time_metric

service:
  pipelines:
    metrics:
      receivers: [prometheus]
      processors: [metric_start_time]
      exporters: [otlp]   # Datadog, Mackerel, or any OTLP backend

Datadog Agent — the openmetrics check reads the same metric with use_process_start_time, so counters that started after the Agent did are counted from zero on the first scrape instead of being dropped:

instances:
  - openmetrics_endpoint: http://localhost/status
    namespace: nginx_vts
    metrics: [".*"]
    use_process_start_time: true

Mackerel — send OTLP through the Collector configuration above; Mackerel stores the counters as cumulative sums and its PromQL rate() / increase() absorb resets. Scraping with mackerel-plugin-prometheus-exporter instead posts every value as is, so a restart shows up as a drop to zero on the graph.

Development

Tests

NGINX_SOURCE_DIR=/path/to/nginx-source cargo test --lib

~75 unit tests cover the shared-table data model, the upstream tracker, the Prometheus formatter (per metric family), the cache statistics helpers, the LOG_PHASE-level FFI, and the rendered /status output via the process-local VTS_MANAGER fallback.

Lints

NGINX_SOURCE_DIR=/path/to/nginx-source cargo clippy --all-targets -- -D warnings
cargo fmt --all -- --check

What's not done yet

The list below tracks known gaps relative to the original nginx-module-vts. None of them block normal traffic monitoring.

Output and control

  • JSON / HTML / JSONP output formats — only Prometheus text is emitted.
  • /control API for reset/delete.
  • vts_dump directive (periodic on-disk dump for counter recovery across restarts). Not planned while output is Prometheus-only; see Persistence.

Filtering and limits

  • Filter zones (vhost_traffic_status_filter_by_set_key, _filter_by_host, _filter_max_node) — no dynamic key-based grouping yet.
  • Traffic limiting (vhost_traffic_status_limit_traffic, _limit_traffic_by_set_key) — the module is observation-only; it cannot rate-limit responses.

Metric coverage

  • Some of the original's families are not emitted yet; see Compatibility with nginx-module-vts.
  • Upstream peer state (down, weight, max_fails, fail_timeout, backup) is not yet read from the nginx upstream configuration.
  • Per-status-code counters (vhost_traffic_status_measure_status_codes) — only the 1xx/2xx/3xx/4xx/5xx class buckets are exposed.
  • Histogram bucket layout is fixed at the Prometheus client_golang defaults (5ms..10s, 11 buckets). There is no vts_histogram_buckets-style directive to customise the bounds.
  • Average method (vhost_traffic_status_average_method AMM / WMA) — averages are plain cumulative sum / count.
  • Embedded $vts_* variables for use in log_format / if — upstream module exposes ~20; we expose none.

License

Licensed under either of

at your option.

Contribution

Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.