A native Rust sampling memory profiler designed for fast insights in production
Rust
36
46 commits
updated Sep 10, 2026
Ying is a native Rust sampling memory profiler which tracks retained memory and allocations. It is designed for production usage in asynchronous Rust programs. Ying(鷹) is Chinese for eagle. 🦅🦅🦅
I started this project because existing solutions I looked at were either consuming too many resources, or wrote profiling files that were too large or too cumbersome to consume, or were very bad at producing useful stack traces especially for Rust async programs, or did not support tracking retained memory in its profiling.
Features:
::poll:: frames in expanded stack traces for clarityProfilerRunner -- utility to spin up thread to dump out reports and optionally flamegraphs every N minutes when total memory usage changes significantlyTo see an example of a long-running toy service that generates the periodic reports and flamegraphs:
cargo run --profile bench --example ying_example
The example runs a cache workload plus a simulated leak for ~12 minutes — long enough for the
reporting thread to write a ying.<timestamp>.<MB>MB.report and a .svg flamegraph under
ying-profiles/ at every check. For a quicker demo, shorten the interval and runtime:
YING_EXAMPLE_INTERVAL_SECS=10 YING_EXAMPLE_RUNTIME_SECS=45 cargo run --profile bench --example ying_example
Every report and flamegraph pair is named with the ISO8601 timestamp of when it was written and
the total retained memory at that moment, so a plain ls of the output directory is already a
memory timeline — you can see when memory grew and by how much before opening anything:
$ ls -l ying-profiles/
total 440
-rw-r--r-- 1 evan staff 17583 Sep 9 16:45 ying.2026-09-09T16:45:05-04:00.15MB.report
-rw-r--r-- 1 evan staff 50492 Sep 9 16:45 ying.2026-09-09T16:45:05-04:00.15MB.svg
-rw-r--r-- 1 evan staff 16254 Sep 9 16:45 ying.2026-09-09T16:45:15-04:00.23MB.report
-rw-r--r-- 1 evan staff 49262 Sep 9 16:45 ying.2026-09-09T16:45:15-04:00.23MB.svg
-rw-r--r-- 1 evan staff 26162 Sep 9 16:48 ying.2026-09-09T16:48:51-04:00.28MB.report
-rw-r--r-- 1 evan staff 50765 Sep 9 16:48 ying.2026-09-09T16:48:51-04:00.28MB.svg
Here memory went from 15MB to 23MB in ten seconds, then to 28MB — pick the report just before the jump and diff it against the one after to see exactly which stacks grew.
Set Ying as the global allocator, then turn on reporting with one line:
use ying_profiler::YingProfiler;
#[global_allocator]
static YING_ALLOC: YingProfiler = YingProfiler::default();
fn main() {
YING_ALLOC.start_profiling();
// ... the rest of your program
}
There are three layers here, and it is worth knowing which one you are using.
The #[global_allocator] declaration is what makes Ying sample allocations at all. On its own it is complete and valid: Ying accumulates stats in memory and defers to the System allocator, but writes nothing anywhere. Nothing is scheduled and no thread is spawned.
ProfilerRunner: the primitive that decides how stats get reportedProfilerRunner is the piece to reach for whenever the defaults do not fit. It owns every decision about turning collected stats into output — how often to look, how much of a change is worth reporting, what to measure, where output goes, and in what form:
| Setting | Meaning |
|---|---|
check_interval_secs | how often the background thread wakes up to compare memory use |
report_pct_change_trigger | how much retained memory must move before a report is written |
reporting_path | directory for reports and flamegraphs; created if missing |
measure_allocated_not_retained | rank stacks by total allocated bytes instead of retained |
gen_flamegraphs | also write an SVG flamegraph alongside each text report |
expand_frames | expand inlined symbols within each stack frame in reports |
Build one with ProfilerRunnerBuilder, which defaults anything you leave out, then hand it the allocator static to start its thread:
use ying_profiler::{utils::ProfilerRunnerBuilder, YingProfiler};
#[global_allocator]
static YING_ALLOC: YingProfiler = YingProfiler::default();
fn main() {
let runner = ProfilerRunnerBuilder::default()
.check_interval_secs(60usize)
.report_pct_change_trigger(5usize)
.reporting_path("/var/log/ying")
.gen_flamegraphs(true)
.build()
.unwrap();
runner.spawn(&YING_ALLOC);
}
start_profiling is not a separate mechanism: it builds exactly one of these with a specific set of values (5 minute interval, 10% trigger, flamegraphs on, retained memory, output to ying-profiles) and spawns it. Anything it can do, a ProfilerRunner you build yourself can do too.
You do not need a runner at all if you would rather decide when to look. YingProfiler exposes the stats as data — top_k_stacks_by_retained, top_k_stacks_by_allocated, total_retained_bytes — and utils::gen_flamegraph writes a one-off flamegraph on demand. This is the route to take when reports should be triggered by your own signals, such as an HTTP endpoint or a health check noticing memory growth.
Ying is the global allocator, so everything Ying allocates re-enters Ying. Four rules keep that from turning into a deadlock, and they all matter if you plan to change the code:
alloc,
dealloc and realloc alike. The profiler's maps allocate and free internally (when a shard
resizes, for instance), and those calls land back in the allocator while a shard lock is held.
Touching the map first would try to take a lock the same thread already holds.&mut handed out from &self. Two threads whose
ids collide then share one guard, which is aliasing UB and, worse, loses increments: a non-atomic
bump can be dropped, the guard falls to zero while a thread is still inside its critical section,
and rule 1 silently stops holding. It is now a thread_local! with a const {} initializer and
no Drop, which is what keeps TLS access from allocating - see the comment on THREAD_STATE.DashMap,
whose RwLock is built on parking_lot_core; parking_lot_core allocates its global parking
table while holding its internal bucket locks, and that allocation comes back through Ying into
DashMap, whose lock slow path then re-enters parking_lot_core and deadlocks. No choice of
allocator for the hash table fixes this. Ying therefore uses a small sharded map built on
std::sync::RwLock, which never allocates.Vec
growing inside a map scan, calling back into realloc, and blocking on a shard lock the same
thread already held. ShardedMap::map_to_vec therefore reserves with no lock held and only
fills within that reservation, so a future hole in the guard degrades into a retry rather than a
process-wide hang.Every bug this crate has had in the allocator path was a deadlock, and deadlocks are probabilistic: one green test run means very little. The guard bug above showed up once in roughly a hundred suite runs, and a 300 run soak of the stock configuration missed it entirely, so reproducing these needs help.
YING_TEST_MAP_CAPACITY overrides the initial capacity of every profiler map. Setting it to 1
starts each map at zero capacity so shards resize constantly, and resizing is the path where a map
allocates and frees while holding its own lock. With that set, a re-entrancy bug that otherwise
appears once in tens of runs wedges the process immediately:
YING_TEST_MAP_CAPACITY=1 cargo nextest run --release --profile stress
The general technique that found the guard bug is worth repeating for anything in this area: make the rare condition the normal one. Shrinking the thread-id slot table from 1024 entries to 64 took the failure rate from 0 in 300 runs to 5 in 300, which was enough to catch it in minutes.
Timeouts must come from outside the process. Once the allocator is stuck, so is panicking and
printing, so a test cannot report its own hang. cargo nextest runs each test in a separate
process and kills it on timeout (see .config/nextest.toml), and CI puts a timeout-minutes on
every test step, because a bad enough bug wedges the binary during startup before nextest can apply
a per-test timeout at all.
When a test does hang, with_watchdog writes to stderr without allocating and then calls abort,
so SIGABRT makes the OS capture a backtrace for every thread. That is the only practical way to
see which lock the hung threads are parked on; on macOS the report lands in
~/Library/Logs/DiagnosticReports, and on Linux you need core dumps enabled.
profile_spans - records tracing-span information in stacks. NOTE: this feature is experimental and incomplete.All profilers are useful and represent amazing work by their authors. I do have a writeup of different memory profilers here.
Rust as an ecosystem is lacking in good memory profiling tools. Bytehound is quite good but has a large CPU impact and writes out huge profiling files as it measures every allocation. Jemalloc/Jeprof does sampling, so it's great for production use, but its output is difficult to interpret, and it does not track retained memory, which is quite critical for debugging memory issues in production. Both of the above tools are written with generic C/C++ malloc/preload ABI in mind, so the backtraces that one gets from their use are really limited, especially for profiling Rust binaries built for an optimized/release target, and especially async code. The output also often has trouble with mangled symbols.
If we use the backtrace crate and analyze release/bench backtraces in detail, we can see why that is.
The Rust compiler does a good job of inlining function calls - even ones across async/await boundaries -
in release code. Thus, for a single instruction pointer (IP) in the stack trace, it might correspond to
many different places in the code. This is from examples/ying_example.rs:
Some(ying_example::insert_one::{{closure}}::h7eddb5f8ebb3289b)
> Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::h7a53098577c44da0)
> Some(ying_example::cache_update_loop::{{closure}}::h38556c7e7ae06bfa)
> Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hd319a0f603a1d426)
> Some(ying_example::main::{{closure}}::h33aa63760e836e2f)
> Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hb2fd3cb904946c24)
A generic tool which just examines the IP and tries to figure out a single symbol would miss out on all of the inlined symbols. Some tools can expand on symbols, but the results still aren't very good.
Future:
222 followers · starred Sep 2026
332 followers · starred Mar 2025
507 followers · starred Sep 2026
102 followers · starred Sep 2026
Rust
100.0%
A native Rust sampling memory profiler designed for fast insights in production
Rust
36
46 commits
updated Sep 10, 2026
Ying is a native Rust sampling memory profiler which tracks retained memory and allocations. It is designed for production usage in asynchronous Rust programs. Ying(鷹) is Chinese for eagle. 🦅🦅🦅
I started this project because existing solutions I looked at were either consuming too many resources, or wrote profiling files that were too large or too cumbersome to consume, or were very bad at producing useful stack traces especially for Rust async programs, or did not support tracking retained memory in its profiling.
Features:
::poll:: frames in expanded stack traces for clarityProfilerRunner -- utility to spin up thread to dump out reports and optionally flamegraphs every N minutes when total memory usage changes significantlyTo see an example of a long-running toy service that generates the periodic reports and flamegraphs:
cargo run --profile bench --example ying_example
The example runs a cache workload plus a simulated leak for ~12 minutes — long enough for the
reporting thread to write a ying.<timestamp>.<MB>MB.report and a .svg flamegraph under
ying-profiles/ at every check. For a quicker demo, shorten the interval and runtime:
YING_EXAMPLE_INTERVAL_SECS=10 YING_EXAMPLE_RUNTIME_SECS=45 cargo run --profile bench --example ying_example
Every report and flamegraph pair is named with the ISO8601 timestamp of when it was written and
the total retained memory at that moment, so a plain ls of the output directory is already a
memory timeline — you can see when memory grew and by how much before opening anything:
$ ls -l ying-profiles/
total 440
-rw-r--r-- 1 evan staff 17583 Sep 9 16:45 ying.2026-09-09T16:45:05-04:00.15MB.report
-rw-r--r-- 1 evan staff 50492 Sep 9 16:45 ying.2026-09-09T16:45:05-04:00.15MB.svg
-rw-r--r-- 1 evan staff 16254 Sep 9 16:45 ying.2026-09-09T16:45:15-04:00.23MB.report
-rw-r--r-- 1 evan staff 49262 Sep 9 16:45 ying.2026-09-09T16:45:15-04:00.23MB.svg
-rw-r--r-- 1 evan staff 26162 Sep 9 16:48 ying.2026-09-09T16:48:51-04:00.28MB.report
-rw-r--r-- 1 evan staff 50765 Sep 9 16:48 ying.2026-09-09T16:48:51-04:00.28MB.svg
Here memory went from 15MB to 23MB in ten seconds, then to 28MB — pick the report just before the jump and diff it against the one after to see exactly which stacks grew.
Set Ying as the global allocator, then turn on reporting with one line:
use ying_profiler::YingProfiler;
#[global_allocator]
static YING_ALLOC: YingProfiler = YingProfiler::default();
fn main() {
YING_ALLOC.start_profiling();
// ... the rest of your program
}
There are three layers here, and it is worth knowing which one you are using.
The #[global_allocator] declaration is what makes Ying sample allocations at all. On its own it is complete and valid: Ying accumulates stats in memory and defers to the System allocator, but writes nothing anywhere. Nothing is scheduled and no thread is spawned.
ProfilerRunner: the primitive that decides how stats get reportedProfilerRunner is the piece to reach for whenever the defaults do not fit. It owns every decision about turning collected stats into output — how often to look, how much of a change is worth reporting, what to measure, where output goes, and in what form:
| Setting | Meaning |
|---|---|
check_interval_secs | how often the background thread wakes up to compare memory use |
report_pct_change_trigger | how much retained memory must move before a report is written |
reporting_path | directory for reports and flamegraphs; created if missing |
measure_allocated_not_retained | rank stacks by total allocated bytes instead of retained |
gen_flamegraphs | also write an SVG flamegraph alongside each text report |
expand_frames | expand inlined symbols within each stack frame in reports |
Build one with ProfilerRunnerBuilder, which defaults anything you leave out, then hand it the allocator static to start its thread:
use ying_profiler::{utils::ProfilerRunnerBuilder, YingProfiler};
#[global_allocator]
static YING_ALLOC: YingProfiler = YingProfiler::default();
fn main() {
let runner = ProfilerRunnerBuilder::default()
.check_interval_secs(60usize)
.report_pct_change_trigger(5usize)
.reporting_path("/var/log/ying")
.gen_flamegraphs(true)
.build()
.unwrap();
runner.spawn(&YING_ALLOC);
}
start_profiling is not a separate mechanism: it builds exactly one of these with a specific set of values (5 minute interval, 10% trigger, flamegraphs on, retained memory, output to ying-profiles) and spawns it. Anything it can do, a ProfilerRunner you build yourself can do too.
You do not need a runner at all if you would rather decide when to look. YingProfiler exposes the stats as data — top_k_stacks_by_retained, top_k_stacks_by_allocated, total_retained_bytes — and utils::gen_flamegraph writes a one-off flamegraph on demand. This is the route to take when reports should be triggered by your own signals, such as an HTTP endpoint or a health check noticing memory growth.
Ying is the global allocator, so everything Ying allocates re-enters Ying. Four rules keep that from turning into a deadlock, and they all matter if you plan to change the code:
alloc,
dealloc and realloc alike. The profiler's maps allocate and free internally (when a shard
resizes, for instance), and those calls land back in the allocator while a shard lock is held.
Touching the map first would try to take a lock the same thread already holds.&mut handed out from &self. Two threads whose
ids collide then share one guard, which is aliasing UB and, worse, loses increments: a non-atomic
bump can be dropped, the guard falls to zero while a thread is still inside its critical section,
and rule 1 silently stops holding. It is now a thread_local! with a const {} initializer and
no Drop, which is what keeps TLS access from allocating - see the comment on THREAD_STATE.DashMap,
whose RwLock is built on parking_lot_core; parking_lot_core allocates its global parking
table while holding its internal bucket locks, and that allocation comes back through Ying into
DashMap, whose lock slow path then re-enters parking_lot_core and deadlocks. No choice of
allocator for the hash table fixes this. Ying therefore uses a small sharded map built on
std::sync::RwLock, which never allocates.Vec
growing inside a map scan, calling back into realloc, and blocking on a shard lock the same
thread already held. ShardedMap::map_to_vec therefore reserves with no lock held and only
fills within that reservation, so a future hole in the guard degrades into a retry rather than a
process-wide hang.Every bug this crate has had in the allocator path was a deadlock, and deadlocks are probabilistic: one green test run means very little. The guard bug above showed up once in roughly a hundred suite runs, and a 300 run soak of the stock configuration missed it entirely, so reproducing these needs help.
YING_TEST_MAP_CAPACITY overrides the initial capacity of every profiler map. Setting it to 1
starts each map at zero capacity so shards resize constantly, and resizing is the path where a map
allocates and frees while holding its own lock. With that set, a re-entrancy bug that otherwise
appears once in tens of runs wedges the process immediately:
YING_TEST_MAP_CAPACITY=1 cargo nextest run --release --profile stress
The general technique that found the guard bug is worth repeating for anything in this area: make the rare condition the normal one. Shrinking the thread-id slot table from 1024 entries to 64 took the failure rate from 0 in 300 runs to 5 in 300, which was enough to catch it in minutes.
Timeouts must come from outside the process. Once the allocator is stuck, so is panicking and
printing, so a test cannot report its own hang. cargo nextest runs each test in a separate
process and kills it on timeout (see .config/nextest.toml), and CI puts a timeout-minutes on
every test step, because a bad enough bug wedges the binary during startup before nextest can apply
a per-test timeout at all.
When a test does hang, with_watchdog writes to stderr without allocating and then calls abort,
so SIGABRT makes the OS capture a backtrace for every thread. That is the only practical way to
see which lock the hung threads are parked on; on macOS the report lands in
~/Library/Logs/DiagnosticReports, and on Linux you need core dumps enabled.
profile_spans - records tracing-span information in stacks. NOTE: this feature is experimental and incomplete.All profilers are useful and represent amazing work by their authors. I do have a writeup of different memory profilers here.
Rust as an ecosystem is lacking in good memory profiling tools. Bytehound is quite good but has a large CPU impact and writes out huge profiling files as it measures every allocation. Jemalloc/Jeprof does sampling, so it's great for production use, but its output is difficult to interpret, and it does not track retained memory, which is quite critical for debugging memory issues in production. Both of the above tools are written with generic C/C++ malloc/preload ABI in mind, so the backtraces that one gets from their use are really limited, especially for profiling Rust binaries built for an optimized/release target, and especially async code. The output also often has trouble with mangled symbols.
If we use the backtrace crate and analyze release/bench backtraces in detail, we can see why that is.
The Rust compiler does a good job of inlining function calls - even ones across async/await boundaries -
in release code. Thus, for a single instruction pointer (IP) in the stack trace, it might correspond to
many different places in the code. This is from examples/ying_example.rs:
Some(ying_example::insert_one::{{closure}}::h7eddb5f8ebb3289b)
> Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::h7a53098577c44da0)
> Some(ying_example::cache_update_loop::{{closure}}::h38556c7e7ae06bfa)
> Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hd319a0f603a1d426)
> Some(ying_example::main::{{closure}}::h33aa63760e836e2f)
> Some(<core::future::from_generator::GenFuture<T> as core::future::future::Future>::poll::hb2fd3cb904946c24)
A generic tool which just examines the IP and tries to figure out a single symbol would miss out on all of the inlined symbols. Some tools can expand on symbols, but the results still aren't very good.
Future:
222 followers · starred Sep 2026
332 followers · starred Mar 2025
507 followers · starred Sep 2026
102 followers · starred Sep 2026
Rust
100.0%