A Rust implementation of nginx-module-vts for virtual host traffic status monitoring, built on top of the ngx-rust framework.
Status: experimental, but the cross-worker aggregation path, the
vts_zone directive, and the /status endpoint are working end-to-end
with nginx 1.31.
┌─────────────────────────────────────────────────────────────────────┐
│ nginx master │
│ │
│ vts_zone main 1m; ─► ngx_shared_memory_add ─► shm_zone │
│ │ │
│ ▼ │
│ vts_init_shm_zone (Rust) │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────┐ │
│ │ slab pool (SlabPool: Allocator) │ │
│ │ ┌─ VtsShared │ │
│ │ │ ├─ RwLock< RbTreeMap< │ │
│ │ │ │ NgxString<SlabPool>, │ │
│ │ │ │ ServerCounters, │ │
│ │ │ │ SlabPool> > (servers) │ │
│ │ │ └─ RwLock< RbTreeMap< │ │
│ │ │ NgxString<SlabPool>, │ │
│ │ │ UpstreamCounters, │ │
│ │ │ SlabPool> > (upstreams) │ │
│ │ └─ shpool->data = &VtsShared (reload-safe) │ │
│ └──────────────────────────────────────────────┘ │
└──────────────────────────────┬──────────────────────────────────────┘
│ fork
┌────────────────┴────────────────┐
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ worker 1 │ │ worker 2 │
│ LOG_PHASE handler │ │ LOG_PHASE handler │
│ └─► record_server │ │ └─► record_server │
│ └─► record_upstr. │ ──┬────► │ └─► record_upstr. │
│ /status handler │ │ │ /status handler │
│ └─► snapshot_* │ │ │ └─► snapshot_* │
└──────────────────────┘ │ └──────────────────────┘
│
ngx::sync::RwLock guards each map independently:
- writers (record_*) take .write()
- readers (/status snapshot_*) take .read()
The shared state is two RbTreeMaps allocated from the slab pool — one
keyed by server_name, one keyed by "upstream\0server" — wrapped in
ngx::sync::RwLock for concurrent worker access. Capacity scales with
the configured vts_zone size rather than being capped at compile time,
and /status reads no longer block concurrent writers thanks to the
reader-writer lock.
Keys are derived from nginx configuration (the matched server block's
first server_name, the upstream block name) — never from the raw
Host header — so attacker-controlled values cannot expand the key
space.
When vts_zone is not declared the FFI transparently falls back to
a process-local manager — this is how the unit tests exercise the data
model without nginx.
/status returns totals across the whole nginx instance.vts_zone directive — declares a real shared-memory zone
(ngx_shared_memory_add) whose init callback creates two
RbTreeMaps inside the slab pool from Rust.server_name (not the raw Host header), so the table can't be
blown up by adversarial Host values.(upstream, backend) peer — status-code class
counters, bytes in/out, request and upstream response times.r->upstream_states is iterated
so each retry attempt (e.g. 502 from peer A followed by 200
from peer B) contributes its own sample to the upstream counters,
not just the final state._bucket{le=...} / _sum / _count over a shared
fixed 11-bucket layout (client_golang defaults), exposed as
nginx_vts_upstream_response_duration_seconds_bucket{...} and
nginx_vts_server_request_duration_seconds_bucket{...}. Both
feed histogram_quantile(0.99, ...) for p50/p90/p99 panels per
upstream peer and per vhost.proxy_cache_path keys_zone=NAME:SIZE) — counts of HIT, MISS, BYPASS, EXPIRED,
STALE, UPDATING, REVALIDATED, SCARCE aggregated across
workers, exposed as nginx_vts_cache_requests_total.proxy_cache_path max_size=…
and current on-disk usage (sh->size × bsize) exposed as
nginx_vts_cache_usage_bytes{cache_size="max"} and {cache_size="used"}.nginx_vts_main_shm_usage_bytes with
shared="max_size", "used_size", "used_node" and "free_size", so a
full zone can be told from an idle one. free_size comes from the slab's
own page count and used_size is the rest of the zone, rather than the sum
of node sizes the original reports, because the slab spends a whole page or
slot per node and that sum reads low right up to the point where inserts
start failing.ngx_stat_* atomics
when nginx is built with --with-http_stub_status_module;
reading/writing/waiting match what stub_status would
report. Without that build flag the module falls back to a
cycle-table walk and only the active total stays meaningful./status-aware counting — the LOG_PHASE
handler skips internal subrequests (auth_request, mirror,
addition, …) and the module's own /status scrapes, so neither
double-counts the per-vhost counters./status with the
text/plain; version=0.0.4 Content-Type that Prometheus 3.x
requires.nginx -s reload reuses the existing shared table,
so counters survive a config reload.cc / clang).export NGINX_SOURCE_DIR=/path/to/nginx-source # ngx-rust looks here
cargo build --release
Output: target/release/libngx_vts_rust.{so,dylib}.
cd /path/to/nginx-source
auto/configure --prefix=/tmp/nginx-vts-test \
--with-compat \
--add-dynamic-module=/path/to/ngx_vts
make
This produces:
objs/nginx — the nginx binary (only needed if you don't already
have one built from the same source).objs/ngx_http_vts_module.so — the dynamic module you load from
nginx.conf via load_module.The repository's config script picks .dylib on macOS and .so on
Linux automatically.
Minimal nginx.conf that proxies through an upstream and exposes
/status:
load_module modules/ngx_http_vts_module.so;
events {
worker_connections 64;
}
http {
vts_zone main 1m;
upstream backend {
server 127.0.0.1:18091;
server 127.0.0.1:18092;
}
# Two local servers acting as the upstream peers.
server { listen 18091; location / { return 200 "peer1\n"; } }
server { listen 18092; location / { return 200 "peer2\n"; } }
server {
listen 18080;
server_name example.test;
location / { proxy_pass http://backend; }
location /status { vts_status; allow 127.0.0.1; deny all; }
}
}
Run it:
mkdir -p /tmp/nginx-vts-test/{conf,logs,modules}
cp objs/ngx_http_vts_module.so /tmp/nginx-vts-test/modules/
cp nginx.conf /tmp/nginx-vts-test/conf/
objs/nginx -p /tmp/nginx-vts-test -c conf/nginx.conf
Drive traffic and read the metrics:
$ seq 1 100 | xargs -P 8 -I{} curl -sS -o /dev/null http://127.0.0.1:18080/
$ curl -sS http://127.0.0.1:18080/status
Trimmed where marked ….
# Prometheus Metrics:
# HELP nginx_vts_info Nginx info
# TYPE nginx_vts_info gauge
nginx_vts_info{hostname="…",module_version="0.1.0",version="1.31.6"} 1
# HELP nginx_vts_start_time_seconds Nginx start time
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791080808
…
# HELP nginx_vts_main_connections Nginx connections
# TYPE nginx_vts_main_connections gauge
nginx_vts_main_connections{status="accepted"} 10
nginx_vts_main_connections{status="active"} 10
…
# HELP nginx_vts_main_shm_usage_bytes Shared memory zone usage
# TYPE nginx_vts_main_shm_usage_bytes gauge
nginx_vts_main_shm_usage_bytes{shared="max_size"} 1048576
nginx_vts_main_shm_usage_bytes{shared="used_size"} 98304
nginx_vts_main_shm_usage_bytes{shared="used_node"} 4
nginx_vts_main_shm_usage_bytes{shared="free_size"} 950272
# HELP nginx_vts_server_bytes_total The request/response bytes
# TYPE nginx_vts_server_bytes_total counter
nginx_vts_server_bytes_total{host="example.test",direction="in"} 8190
nginx_vts_server_bytes_total{host="example.test",direction="out"} 16065
…
# HELP nginx_vts_server_requests_total The requests counter
# TYPE nginx_vts_server_requests_total counter
nginx_vts_server_requests_total{host="example.test",code="2xx"} 105
…
nginx_vts_server_requests_total{host="_",code="2xx"} 105
…
nginx_vts_server_requests_total{host="*",code="2xx"} 210
…
# HELP nginx_vts_upstream_requests_total The upstream requests counter
# TYPE nginx_vts_upstream_requests_total counter
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18091",code="2xx"} 53
…
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18092",code="2xx"} 52
…
# HELP nginx_vts_upstream_server_up Upstream server status (1=up, 0=down)
# TYPE nginx_vts_upstream_server_up gauge
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18091"} 1
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18092"} 1
Note that peer1 (53) + peer2 (52) = 105: both workers feed the same
table, so /status shows the totals regardless of which worker
happened to handle the request. host="_" is the two peer servers,
which have no server_name and live in the same nginx here, so the
host="*" row that sums every zone counts each request twice.
The Prometheus output uses the original module's metric names, label
names (host, code, upstream, backend, cache_zone, …) and
label values, including the host="*" row that sums every server
zone. What still differs:
process_start_time_seconds (see
Persistence) and nginx_vts_upstream_server_up.used_size is what the slab has spent, not the sum of
node sizes. nginx_vts_start_time_seconds is when the counters
started from zero, which a reload does not change. The
*_request_seconds and *_response_seconds averages are cumulative
(sum / count), not the original's moving average.nginx_vts_server_request_duration_seconds and
nginx_vts_upstream_response_duration_seconds are always emitted,
with fixed buckets, and le is written 1 rather than 1.000.
The original emits them only when histogram_buckets is set.nginx_vts_server_cache_total,
nginx_vts_cache_bytes_total,
nginx_vts_upstream_request_duration_seconds,
nginx_vts_status_code_requests_total and the nginx_vts_filter_*
families.proxy_pass straight to an
address is reported under its own name rather than the original's
upstream="::nogroups".| Directive | Context | Args | Description |
|---|---|---|---|
vts_zone | http | name size | Declare the shared-memory zone backing all counters. Minimum size is 1 MB; without this directive the module silently falls back to process-local counters (mainly useful for tests). |
vts_status | location | — | Render the Prometheus text response at this location. |
vts_upstream_stats | http, server, location | on | off | Accepted for backward compatibility; currently a no-op (upstream stats are always collected when vts_zone is set). |
The shared state is two RbTreeMaps — one keyed by server_name, one
keyed by the (upstream, server) pair — allocated inside the slab pool
that backs the vts_zone. There is no compile-time slot cap: how many
distinct keys you can track is bounded only by the slab pool size you
configure with vts_zone <name> <size>.
Rough sizing rule of thumb: a 1m zone comfortably holds a few thousand
server-zone keys plus a few thousand upstream pairs. Each entry is on the
order of ~200 bytes for the counters plus the key length plus rbtree
node overhead. Bump the size if you genuinely have more virtual hosts.
When a new key cannot be allocated (the slab pool is full), it is dropped
silently and existing counters keep updating. There is also a defensive
upper bound on key length (VTS_MAX_KEY_BYTES = 256) to keep
misconfigured server_name directives from chewing up the pool.
Keys are derived from nginx configuration (the matched server block's
first server_name, the upstream block name) — never from the raw Host
header — so attacker-controlled values cannot expand the key space.
Counters live only in shared memory; nothing is written to disk.
| Operation | Counters |
|---|---|
nginx -s reload | Kept — nginx reuses the zone |
Reload after changing the vts_zone size | Reset to zero — a new zone is allocated |
| Stop / start, container restart | Reset to zero |
Binary upgrade (USR2) | Reset to zero — the new master does not inherit the zone |
All *_total metrics are Prometheus counters, so query them with
rate() / increase(), which detect and absorb resets. What a restart
loses is only the increment since the last scrape.
When a zone is configured, /status also reports when the counters
last started from zero, under the original module's name and under the
name other collectors look for:
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791039391
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1791039391
It is the time the zone was built, not the time a process started, so
it stays put across a reload and moves forward exactly when the
counters reset. process_start_time_seconds is unprefixed because
that is the name other collectors look for to tell a reset from a
series they have only just started watching.
OpenTelemetry Collector — scrape /status with the prometheus
receiver and let the metric_start_time processor take the start time
from this metric, before any batching:
receivers:
prometheus:
config:
scrape_configs:
- job_name: nginx-vts
metrics_path: /status
static_configs:
- targets: ["localhost:80"]
processors:
metric_start_time:
strategy: start_time_metric
service:
pipelines:
metrics:
receivers: [prometheus]
processors: [metric_start_time]
exporters: [otlp] # Datadog, Mackerel, or any OTLP backend
Datadog Agent — the openmetrics check reads the same metric with
use_process_start_time, so counters that started after the Agent did
are counted from zero on the first scrape instead of being dropped:
instances:
- openmetrics_endpoint: http://localhost/status
namespace: nginx_vts
metrics: [".*"]
use_process_start_time: true
Mackerel — send OTLP through the Collector configuration above;
Mackerel stores the counters as cumulative sums and its PromQL
rate() / increase() absorb resets. Scraping with
mackerel-plugin-prometheus-exporter instead posts every value as is,
so a restart shows up as a drop to zero on the graph.
NGINX_SOURCE_DIR=/path/to/nginx-source cargo test --lib
~75 unit tests cover the shared-table data model, the upstream
tracker, the Prometheus formatter (per metric family), the cache
statistics helpers, the LOG_PHASE-level FFI, and the rendered
/status output via the process-local VTS_MANAGER fallback.
NGINX_SOURCE_DIR=/path/to/nginx-source cargo clippy --all-targets -- -D warnings
cargo fmt --all -- --check
The list below tracks known gaps relative to the original
nginx-module-vts. None of them block normal traffic monitoring.
/control API for reset/delete.vts_dump directive (periodic on-disk dump for counter recovery
across restarts). Not planned while output is Prometheus-only; see
Persistence.vhost_traffic_status_filter_by_set_key,
_filter_by_host, _filter_max_node) — no dynamic key-based
grouping yet.vhost_traffic_status_limit_traffic,
_limit_traffic_by_set_key) — the module is observation-only; it
cannot rate-limit responses.down, weight, max_fails,
fail_timeout, backup) is not yet read from the nginx upstream
configuration.vhost_traffic_status_measure_status_codes) — only the
1xx/2xx/3xx/4xx/5xx class buckets are exposed.client_golang defaults (5ms..10s, 11 buckets). There is no
vts_histogram_buckets-style directive to customise the bounds.vhost_traffic_status_average_method AMM / WMA) —
averages are plain cumulative sum / count.$vts_* variables for use in log_format / if —
upstream module exposes ~20; we expose none.Licensed under either of
at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.
A Rust implementation of nginx-module-vts for virtual host traffic status monitoring, built on top of the ngx-rust framework.
Status: experimental, but the cross-worker aggregation path, the
vts_zone directive, and the /status endpoint are working end-to-end
with nginx 1.31.
┌─────────────────────────────────────────────────────────────────────┐
│ nginx master │
│ │
│ vts_zone main 1m; ─► ngx_shared_memory_add ─► shm_zone │
│ │ │
│ ▼ │
│ vts_init_shm_zone (Rust) │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────┐ │
│ │ slab pool (SlabPool: Allocator) │ │
│ │ ┌─ VtsShared │ │
│ │ │ ├─ RwLock< RbTreeMap< │ │
│ │ │ │ NgxString<SlabPool>, │ │
│ │ │ │ ServerCounters, │ │
│ │ │ │ SlabPool> > (servers) │ │
│ │ │ └─ RwLock< RbTreeMap< │ │
│ │ │ NgxString<SlabPool>, │ │
│ │ │ UpstreamCounters, │ │
│ │ │ SlabPool> > (upstreams) │ │
│ │ └─ shpool->data = &VtsShared (reload-safe) │ │
│ └──────────────────────────────────────────────┘ │
└──────────────────────────────┬──────────────────────────────────────┘
│ fork
┌────────────────┴────────────────┐
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ worker 1 │ │ worker 2 │
│ LOG_PHASE handler │ │ LOG_PHASE handler │
│ └─► record_server │ │ └─► record_server │
│ └─► record_upstr. │ ──┬────► │ └─► record_upstr. │
│ /status handler │ │ │ /status handler │
│ └─► snapshot_* │ │ │ └─► snapshot_* │
└──────────────────────┘ │ └──────────────────────┘
│
ngx::sync::RwLock guards each map independently:
- writers (record_*) take .write()
- readers (/status snapshot_*) take .read()
The shared state is two RbTreeMaps allocated from the slab pool — one
keyed by server_name, one keyed by "upstream\0server" — wrapped in
ngx::sync::RwLock for concurrent worker access. Capacity scales with
the configured vts_zone size rather than being capped at compile time,
and /status reads no longer block concurrent writers thanks to the
reader-writer lock.
Keys are derived from nginx configuration (the matched server block's
first server_name, the upstream block name) — never from the raw
Host header — so attacker-controlled values cannot expand the key
space.
When vts_zone is not declared the FFI transparently falls back to
a process-local manager — this is how the unit tests exercise the data
model without nginx.
/status returns totals across the whole nginx instance.vts_zone directive — declares a real shared-memory zone
(ngx_shared_memory_add) whose init callback creates two
RbTreeMaps inside the slab pool from Rust.server_name (not the raw Host header), so the table can't be
blown up by adversarial Host values.(upstream, backend) peer — status-code class
counters, bytes in/out, request and upstream response times.r->upstream_states is iterated
so each retry attempt (e.g. 502 from peer A followed by 200
from peer B) contributes its own sample to the upstream counters,
not just the final state._bucket{le=...} / _sum / _count over a shared
fixed 11-bucket layout (client_golang defaults), exposed as
nginx_vts_upstream_response_duration_seconds_bucket{...} and
nginx_vts_server_request_duration_seconds_bucket{...}. Both
feed histogram_quantile(0.99, ...) for p50/p90/p99 panels per
upstream peer and per vhost.proxy_cache_path keys_zone=NAME:SIZE) — counts of HIT, MISS, BYPASS, EXPIRED,
STALE, UPDATING, REVALIDATED, SCARCE aggregated across
workers, exposed as nginx_vts_cache_requests_total.proxy_cache_path max_size=…
and current on-disk usage (sh->size × bsize) exposed as
nginx_vts_cache_usage_bytes{cache_size="max"} and {cache_size="used"}.nginx_vts_main_shm_usage_bytes with
shared="max_size", "used_size", "used_node" and "free_size", so a
full zone can be told from an idle one. free_size comes from the slab's
own page count and used_size is the rest of the zone, rather than the sum
of node sizes the original reports, because the slab spends a whole page or
slot per node and that sum reads low right up to the point where inserts
start failing.ngx_stat_* atomics
when nginx is built with --with-http_stub_status_module;
reading/writing/waiting match what stub_status would
report. Without that build flag the module falls back to a
cycle-table walk and only the active total stays meaningful./status-aware counting — the LOG_PHASE
handler skips internal subrequests (auth_request, mirror,
addition, …) and the module's own /status scrapes, so neither
double-counts the per-vhost counters./status with the
text/plain; version=0.0.4 Content-Type that Prometheus 3.x
requires.nginx -s reload reuses the existing shared table,
so counters survive a config reload.cc / clang).export NGINX_SOURCE_DIR=/path/to/nginx-source # ngx-rust looks here
cargo build --release
Output: target/release/libngx_vts_rust.{so,dylib}.
cd /path/to/nginx-source
auto/configure --prefix=/tmp/nginx-vts-test \
--with-compat \
--add-dynamic-module=/path/to/ngx_vts
make
This produces:
objs/nginx — the nginx binary (only needed if you don't already
have one built from the same source).objs/ngx_http_vts_module.so — the dynamic module you load from
nginx.conf via load_module.The repository's config script picks .dylib on macOS and .so on
Linux automatically.
Minimal nginx.conf that proxies through an upstream and exposes
/status:
load_module modules/ngx_http_vts_module.so;
events {
worker_connections 64;
}
http {
vts_zone main 1m;
upstream backend {
server 127.0.0.1:18091;
server 127.0.0.1:18092;
}
# Two local servers acting as the upstream peers.
server { listen 18091; location / { return 200 "peer1\n"; } }
server { listen 18092; location / { return 200 "peer2\n"; } }
server {
listen 18080;
server_name example.test;
location / { proxy_pass http://backend; }
location /status { vts_status; allow 127.0.0.1; deny all; }
}
}
Run it:
mkdir -p /tmp/nginx-vts-test/{conf,logs,modules}
cp objs/ngx_http_vts_module.so /tmp/nginx-vts-test/modules/
cp nginx.conf /tmp/nginx-vts-test/conf/
objs/nginx -p /tmp/nginx-vts-test -c conf/nginx.conf
Drive traffic and read the metrics:
$ seq 1 100 | xargs -P 8 -I{} curl -sS -o /dev/null http://127.0.0.1:18080/
$ curl -sS http://127.0.0.1:18080/status
Trimmed where marked ….
# Prometheus Metrics:
# HELP nginx_vts_info Nginx info
# TYPE nginx_vts_info gauge
nginx_vts_info{hostname="…",module_version="0.1.0",version="1.31.6"} 1
# HELP nginx_vts_start_time_seconds Nginx start time
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791080808
…
# HELP nginx_vts_main_connections Nginx connections
# TYPE nginx_vts_main_connections gauge
nginx_vts_main_connections{status="accepted"} 10
nginx_vts_main_connections{status="active"} 10
…
# HELP nginx_vts_main_shm_usage_bytes Shared memory zone usage
# TYPE nginx_vts_main_shm_usage_bytes gauge
nginx_vts_main_shm_usage_bytes{shared="max_size"} 1048576
nginx_vts_main_shm_usage_bytes{shared="used_size"} 98304
nginx_vts_main_shm_usage_bytes{shared="used_node"} 4
nginx_vts_main_shm_usage_bytes{shared="free_size"} 950272
# HELP nginx_vts_server_bytes_total The request/response bytes
# TYPE nginx_vts_server_bytes_total counter
nginx_vts_server_bytes_total{host="example.test",direction="in"} 8190
nginx_vts_server_bytes_total{host="example.test",direction="out"} 16065
…
# HELP nginx_vts_server_requests_total The requests counter
# TYPE nginx_vts_server_requests_total counter
nginx_vts_server_requests_total{host="example.test",code="2xx"} 105
…
nginx_vts_server_requests_total{host="_",code="2xx"} 105
…
nginx_vts_server_requests_total{host="*",code="2xx"} 210
…
# HELP nginx_vts_upstream_requests_total The upstream requests counter
# TYPE nginx_vts_upstream_requests_total counter
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18091",code="2xx"} 53
…
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18092",code="2xx"} 52
…
# HELP nginx_vts_upstream_server_up Upstream server status (1=up, 0=down)
# TYPE nginx_vts_upstream_server_up gauge
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18091"} 1
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18092"} 1
Note that peer1 (53) + peer2 (52) = 105: both workers feed the same
table, so /status shows the totals regardless of which worker
happened to handle the request. host="_" is the two peer servers,
which have no server_name and live in the same nginx here, so the
host="*" row that sums every zone counts each request twice.
The Prometheus output uses the original module's metric names, label
names (host, code, upstream, backend, cache_zone, …) and
label values, including the host="*" row that sums every server
zone. What still differs:
process_start_time_seconds (see
Persistence) and nginx_vts_upstream_server_up.used_size is what the slab has spent, not the sum of
node sizes. nginx_vts_start_time_seconds is when the counters
started from zero, which a reload does not change. The
*_request_seconds and *_response_seconds averages are cumulative
(sum / count), not the original's moving average.nginx_vts_server_request_duration_seconds and
nginx_vts_upstream_response_duration_seconds are always emitted,
with fixed buckets, and le is written 1 rather than 1.000.
The original emits them only when histogram_buckets is set.nginx_vts_server_cache_total,
nginx_vts_cache_bytes_total,
nginx_vts_upstream_request_duration_seconds,
nginx_vts_status_code_requests_total and the nginx_vts_filter_*
families.proxy_pass straight to an
address is reported under its own name rather than the original's
upstream="::nogroups".| Directive | Context | Args | Description |
|---|---|---|---|
vts_zone | http | name size | Declare the shared-memory zone backing all counters. Minimum size is 1 MB; without this directive the module silently falls back to process-local counters (mainly useful for tests). |
vts_status | location | — | Render the Prometheus text response at this location. |
vts_upstream_stats | http, server, location | on | off | Accepted for backward compatibility; currently a no-op (upstream stats are always collected when vts_zone is set). |
The shared state is two RbTreeMaps — one keyed by server_name, one
keyed by the (upstream, server) pair — allocated inside the slab pool
that backs the vts_zone. There is no compile-time slot cap: how many
distinct keys you can track is bounded only by the slab pool size you
configure with vts_zone <name> <size>.
Rough sizing rule of thumb: a 1m zone comfortably holds a few thousand
server-zone keys plus a few thousand upstream pairs. Each entry is on the
order of ~200 bytes for the counters plus the key length plus rbtree
node overhead. Bump the size if you genuinely have more virtual hosts.
When a new key cannot be allocated (the slab pool is full), it is dropped
silently and existing counters keep updating. There is also a defensive
upper bound on key length (VTS_MAX_KEY_BYTES = 256) to keep
misconfigured server_name directives from chewing up the pool.
Keys are derived from nginx configuration (the matched server block's
first server_name, the upstream block name) — never from the raw Host
header — so attacker-controlled values cannot expand the key space.
Counters live only in shared memory; nothing is written to disk.
| Operation | Counters |
|---|---|
nginx -s reload | Kept — nginx reuses the zone |
Reload after changing the vts_zone size | Reset to zero — a new zone is allocated |
| Stop / start, container restart | Reset to zero |
Binary upgrade (USR2) | Reset to zero — the new master does not inherit the zone |
All *_total metrics are Prometheus counters, so query them with
rate() / increase(), which detect and absorb resets. What a restart
loses is only the increment since the last scrape.
When a zone is configured, /status also reports when the counters
last started from zero, under the original module's name and under the
name other collectors look for:
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791039391
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1791039391
It is the time the zone was built, not the time a process started, so
it stays put across a reload and moves forward exactly when the
counters reset. process_start_time_seconds is unprefixed because
that is the name other collectors look for to tell a reset from a
series they have only just started watching.
OpenTelemetry Collector — scrape /status with the prometheus
receiver and let the metric_start_time processor take the start time
from this metric, before any batching:
receivers:
prometheus:
config:
scrape_configs:
- job_name: nginx-vts
metrics_path: /status
static_configs:
- targets: ["localhost:80"]
processors:
metric_start_time:
strategy: start_time_metric
service:
pipelines:
metrics:
receivers: [prometheus]
processors: [metric_start_time]
exporters: [otlp] # Datadog, Mackerel, or any OTLP backend
Datadog Agent — the openmetrics check reads the same metric with
use_process_start_time, so counters that started after the Agent did
are counted from zero on the first scrape instead of being dropped:
instances:
- openmetrics_endpoint: http://localhost/status
namespace: nginx_vts
metrics: [".*"]
use_process_start_time: true
Mackerel — send OTLP through the Collector configuration above;
Mackerel stores the counters as cumulative sums and its PromQL
rate() / increase() absorb resets. Scraping with
mackerel-plugin-prometheus-exporter instead posts every value as is,
so a restart shows up as a drop to zero on the graph.
NGINX_SOURCE_DIR=/path/to/nginx-source cargo test --lib
~75 unit tests cover the shared-table data model, the upstream
tracker, the Prometheus formatter (per metric family), the cache
statistics helpers, the LOG_PHASE-level FFI, and the rendered
/status output via the process-local VTS_MANAGER fallback.
NGINX_SOURCE_DIR=/path/to/nginx-source cargo clippy --all-targets -- -D warnings
cargo fmt --all -- --check
The list below tracks known gaps relative to the original
nginx-module-vts. None of them block normal traffic monitoring.
/control API for reset/delete.vts_dump directive (periodic on-disk dump for counter recovery
across restarts). Not planned while output is Prometheus-only; see
Persistence.vhost_traffic_status_filter_by_set_key,
_filter_by_host, _filter_max_node) — no dynamic key-based
grouping yet.vhost_traffic_status_limit_traffic,
_limit_traffic_by_set_key) — the module is observation-only; it
cannot rate-limit responses.down, weight, max_fails,
fail_timeout, backup) is not yet read from the nginx upstream
configuration.vhost_traffic_status_measure_status_codes) — only the
1xx/2xx/3xx/4xx/5xx class buckets are exposed.client_golang defaults (5ms..10s, 11 buckets). There is no
vts_histogram_buckets-style directive to customise the bounds.vhost_traffic_status_average_method AMM / WMA) —
averages are plain cumulative sum / count.$vts_* variables for use in log_format / if —
upstream module exposes ~20; we expose none.Licensed under either of
at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.