hmijail/ObsSyncBugHunt

A semantic fuzzer to find bugs in Obsidian Sync as a distributed system

TypeScript

0

90 commits

updated Sep 30, 2026

See the code

See what people are saying

README

Obsidian Sync Bug Hunter

This is a semantic fuzzer for a simple distributed system (Obsidian Sync and its clients). In other words, a test harness that hunts for data loss in Obsidian Sync when the same note(s) are edited alternatively on multiple Obsidian instances.

Inspired by Jepsen, which would be overkill for something like Obsidian Sync.

I wrote this README personally. Everything else, including the docs/ directory, are Claude artifacts.

Lose data on your own Obsidian Sync here.
See the timeline of that example here, as reported by the fuzzer.

Some background: data loss in Obsidian Sync!?

Obsidian is a nice note-taking app. It's closed-source but free. It has a sync service, Obsidian Sync, which is subscription-based. This service has data-losing bugs. A thread in the Obsidian forums has been running for 2 years now gathering complaints, but the devs seem unable to find the problem. They proposed workarounds, but they fail too.

I lost data to Obsidian Sync too, and found that thread. I proposed using e.g. Jepsen to find bugs in a systematic way. There was no response.

I wanted some test project to use Claude Code on, so I told it I wanted to apply Jepsen to Obsidian. Claude jumped to make things happen; unfortunately those were pretty silly things. It quickly became clear that Claude needs its tasks to have much, much tighter scope.

Long story short, I started guiding the design of a semantic fuzzer, following the adversarial/paranoiac themes from my own DARUM, which (in its own way) also plays with randomness and repetitions to force an uncollaborative black box to reveal a bit of its inner workings. Plus containers and network control.

So 100% of the design is mine (and this README), but the code is 100% Claude's. In fact, I never used TypeScript; I chose it because it's a language used in the Obsidian ecosystem... and to force myself to stay hands-off and trust Claude. (I won't be repeating that)

Results

One result is that the fuzzer works: the harness finds different sequences of operations that trigger sync bugs in Obsidian, measures their repeatability and even helps understand how the sequence failed. Yay!

The other result is that Claude is a surprisingly, increasingly incompetent assistant. The project ballooned into 60h, half of it fighting the bug treadmill. Full experience report here, plus some posible remediations.

The summary is that keeping Claude in a leash tight enough to stop it from doing silly stuff is hard work, which eventually turns into a bug treadmill anyway. Claude is like an intern that knows far too much for their own good, uses that knowledge to make bad choices... plus periodically forgets instructions... but rarely lets go of pointless minutiae.
Also, you're responsible for what it remembers, even though you only have blunt tools to control that.
Also, those tools keep changing, and even Anthropic's instructions aren't very consistent, which confuses even Claude.

So that's a blurry mess. OK, but what did I learn from this project? Only things about Claude itself, the stuff that keeps changing. But nothing about the matter at hand. In fact, it's the opposite: I had to teach Claude how to build this.

If Claude was an intern, I would expect that they learnt something, and maybe even that they'd take over and keep the project moving forward. But Claude doesn't learn.

In a nutshell: this is an insta-legacy project, that ties you to LLMs, and required experience, but didn't create it.

Let's lose some data

As a motivating example, this is one sequence of operations found by the fuzzer that causes data loss ~100% of the time in Obsidian 1.12 and 1.13.7 (latest as of this writing). Reported 2 months ago, still not acknowledged nor fixed as of this writing.

Let’s assume you use Obsidian with Sync in your phone and your laptop. Both should set their Sync settings to “Conflict file” mode, which is the devs’ recommendation hoping to minimize data loss.

Note that this particular bug needs you to finish all the steps within 60 seconds! Later we’ll see why.

  1. Set your iPhone on airplane mode; ensure Wifi is also disconnected.
  2. In your laptop, created the note “buggy” (or whatever you want)
  3. In that note, type “laptop”
  4. Wait until Obsidian syncs up and shows the green sync icon (a couple of seconds)
  5. On your phone, create the same note “buggy”
  6. In that note, type “phone”
  7. Disable airplane mode on the phone and wait for sync.

After sync finishes, only the line “laptop” remains in the “buggy” note. That’s to be expected, since there was a conflict; the problem is that a Conflict File should have been created to save the “phone” line, but it didn’t. So that line is gone everywhere. And you didn’t dream it: looking at the Sync version history, you’ll see that the line did indeed reach the server.

The fuzzer shows how the bug happens with an ASCII timeline.

Fuzzer features

  • Sets up multiple Obsidian instances running in containers, prepared to use Obsidian Sync (requires an Obsidian Sync subscription)
  • Represents sequences of user actions as a compact string (aka history)
  • Can generate histories randomly, with configurabale complexity (number of nodes, number of notes, user's patience to wait for sync status, network failures)
  • Runs each history for a configurable number of repetitions
  • Samples whole-system behavior during the execution of histories, with configurable granularity / invasiveness
  • Gathers logs of everything for later analysis, and plots timelines summarizing them
  • Analyzes results, ranking histories by % of data-losing repetitions, plus flags suspicious behavior for further analysis
  • Generates shell scripts implementing a history, for verifiable reproducibility of bugs with minimal machinery
  • Semi-automatic upgradeability to new Obsidian versions, with self-checks to confirm whether the harness can still deal with the new version or requires modifications.

Prerequisites

  • Obsidian Sync subscription (if you want one just to test this project, know that it seems refundable during the first week)
  • Docker (in macOS, 2 vCPUs / 4GB RAM in the VM is enough for 2 Obsidian containers). Podman also works but seems to make sampling slower; prefer Apple Hypervisor to krunkit.
  • Node >= 22
  • Python 3
  • A VNC client to make the containerized Obsidian log in to your Sync account (and optionally to watch how Obsidian runs histories)
  • Curl
  • Optional:
    • Coreutils (for gtimeout)
    • Fnm (to pin down Node versions)
    • A local (non-containerized) Obsidian instance, mainly useful to find and test bugs in the macOS Obsidian version. (Needs its CLI activated)

On macOS you can install all the prerequisites with brew, and you can use the standard Screen Sharing.app for the VNC connection.

No LLM is used in the harness. Bugs found can't be hallucinations.

Quick start

The test harness will be creating and editing lots of notes on your Sync vault. It will try to keep the vault safe, by only ever acting on notes inside a folder ("bughunt") in your vault. (In any case you should backup your vault; personally, until bugs are fixed I moved my vault out of Sync and into iCloud Drive)

make is the easy entry point to the project, which maps to other tools as needed. make help lists every available command.

Common flow:

make install && make check        # install (npm ci) + typecheck + tests

# Create node and prepare it for Obsidian Sync's login:
make login
# Connect through VNC to the container (localhost:5901). A pristine Obsidian is waiting. Configure it to Sync to a vault and set it to "Create conflict file". Enable the Obsidian CLI.
make capture-login                # extracts the settings and login credentials into ./secrets
make containers-up                # launch n1 + n2 fresh with the captured credentials
make unpause-sync                 # let the nodes start syncing
make check-assumptions            # does everything look as expected? (run after updates, etc)
make clean-data                   # OPTIONAL clean slate: empty the vault + wipe runs/

make run HISTORY=N1AaWN2Aa REPEAT=3      # run one specific history
make soak HISTORY=N1AaWN2Aa              # soak one history until Ctrl-C
make soak                                # generate histories and run them until Ctrl-C
make analyze                             # aggregate runs/ into a report in runs/analysis.md

make repro HISTORY=N1DAaWN2AaC           # create a shell script to run that history (bug reproduction with minimal machinery)

How it all works

In a nutshell: we will set up a couple (or more!) Obsidian clients in containers, make them Sync, and then we'll create and edit notes in them, while checking that no data is lost. When it does, we'll record how repeatable is that case, in a way amenable to be reported to the Obsidian devs.

A set of Obsidian clients, ready to Sync

When you run make build-image, a container image will be prepared to run Obsidian. Then, make login will run it for you to connect through VNC. You will log in to your Obsidian Sync account, connect to a Sync vault, enable creation of conflict files, and enable the Obsidian CLI.

Next, make capture-login will extract from that container your Sync credentials and copy them into the secrets directory. This is so that multiple containers can reuse the same credentials, without manually setting up each one of them individually. This information never leaves your computer. Don't publish that directory; it's already in .gitignore.

And then, make containers-up will start your nodes and get them syncing. If there's notes in the vault, they might take some time to finish their initial sync.

If you want to see or interact with them, you can always connect to each node through VNC, in the port 5900+(node number).

Generating sequences of Obsidian edits with a tiny DSL

We will test sequences of actions, AKA histories. A history is a string of user actions replayed against multiple Obsidian nodes. Commands are uppercase, parameters lowercase/digits.

Commandmeaning
N<d>set the active node (N1, N2)
Lset the active node to local Obsidian instance
A<x>make the active node append a uniquely-tagged line to note x (creates it if necessary)
D / Cdisconnect / connect the active node from the network (applies to containers, not the Local node)
W[n]wait until the active node is synced (with optional timeout of n seconds)
P[n]pause ~n seconds (default 10)

Example: N1AaWN2Aa= node 1 appends to note a, waits for sync; node 2 appends to the same note.

At the end of the history, the harness reconnects all nodes to the network, waits for them all to report synced, and still waits for a settling window to ensure that no further changes happen (e.g. generation of conflict files, which later get synced, which could still generate conflicts in other clients, etc etc). An end result is judged only when things remain stable for FINAL_SETTLE_SEC seconds.

Note that the harness models a single user using Obsidian across multiple devices, so there's a single thread of control doing everything. This means that e.g. a Pause command applies across all nodes at once: the control thread does nothing for n seconds, while the Obsidian nodes will keep doing their thing (e.g. try to sync). Similarly, W waits for the current node to be synced, which doesn't stop the rest of nodes from working.

Histories can be auto-generated randomly (each op's probability is adjustable, see Parameters) or typed manually. They are normalized so that histories that would be similar in practice also look similar as a string (see the History Normalization section).

Timings are unavoidably variable across repetitions of a history, since we don't have control of the Sync server, internet traffic, etc. This might cause Sync results to change every time you repeat the history. Therefore histories are run for REPEAT times to sample the distribution of end results. If one wants to minimize variability, command W waits until Obsidian the node is synced.

Edits to notes are append-only (for now?), since that is a case supported by the CLI. Edits optionally (OPEN_NOTES) can also open the note in the nodes' GUI, so you can watch the history unfold through a VNC connection if you want to.

One key difficulty: W formalizes (somewhat) the synced status

Syncing is hard. It can happen that the user is editing stuff while changes are coming in or out of Obsidian. It can even happen that something is wrong with the network and syncing gets delayed, accumulating conflicting changes at different points of the process, while you, the user, might not know or even care that this is happening.

The user can (in theory) count on Obsidian to deal with this. But for the fuzzer to do its job better we need to try to formalize the different underlying scenarios. Hence, W doesn't just trust the Sync status reported by Obsidian (which corresponds to the Sync icon in the GUI), but also checks metadata from the Sync server and content of modified notes; it will only let the history continue when the synced status does seem correct (and will otherwise record a data loss bug if it looks wrong after a configurable grace period).

On the other hand, Obsidian Sync bugs might also get triggered by an impatient user who tries editing notes while the Sync status is not settled. This case is covered by W<n>. For example, W10 simulates a user who waits for a maximum of 10 seconds for the GUI Sync icon to get green; after that, they just go ahead and edit a note anyway. (Which makes sense, because sometimes the icon changes are hard to see).

Practical example

The motivating example at the beginning of this README was found by the fuzzer as this history: N2DN1AaWN2AaCW

(Found in Obsidian 1.12.7, still there in 1.13.7)

  • N2: selects N2 as the current node
  • D : disconnects the current node
  • N1: selects N1 as the current node
  • Aa: appends a token to note "a" in node 1 (this is a "logical name"; see section on naming of notes)
  • W : wait until the current note in the current node is reported as synced by Obsidian
  • N2: selects N2
  • Aa: appends a new token to note "a" in node 2 (same "logical name" as before)
  • C : connect the current node
  • W : wait for sync

Timelines

Running a history creates long logs. For easier inspection, these are summarized into an ASCII timeline that shows what happened, where and when.

For example, this is the timeline generated by the motivating example:

make run HISTORY=N2DN1AaWN2AaCW SAMPLING=everything

seconds    0    1      2  3 4   5    7            60
n1 ops     |   a|  W   |  | |   ||   |            |   |   ||    |
n1 sync    |    |x .   |x | |.  ||.  |            |.  |.  ||x   |.
n1 vers a  | -  |x     |1 | |2  ||.  |            |.  |x  ||.   |.
n1 file a  |    |x .   |M | |M  ||m  |            |m  |m  ||m   |m
n2 ops     |D   |    aC|  |W|   ||   |    ...     |   |   ||    |
n2 sync    |    | x    | x|s| ..|| ..|            | x.| x.|| x. | x
n2 vers a  |  - | x -  | x| | 2 || . |            | . | . || .  | .
n2 file a  |    | x    | M| | m || m |            | m | m || m L| m

How to read that?

TImelines consist of lanes (rows), with each lane dedicated to one type of data we sample. In each lane, various letters represent events. Samples that are taken at the same time are plotted in the same column. For ease of reference, seconds are separated by columns of |.

Letter meanings:

  • general: x the call was cut off · ? reply not recognised · not sampled · | a wallclock second
  • ops:a/b/… a note was written to · D/C disconnect/connect · W a wait started · P a pause · h a W<n> handed off early
  • sync: . synced · s syncing · p paused · e error · o offline · h stopped
  • vers: a digit when the server version counter moved, to that value · . sampled but unchanged · - no server history yet
  • file (per node per note) — . file is complete and unchanged · u file changed here, is complete · m missing a token · M changed here, missing a token · c changed here and a conflict file appeared · C same, but some token is missing · - no file · L loss declared · ! readings that do not add up

A column where nothing was sampled will be empty (||). In general, dots mean "we sampled here but wasn't interesting".

Data loss, play by play

Can you see now what happened in that example history?

  • The operations actually finished on second 3: Node 2 waited for sync to happen, and at that moment, Obsidian reported it was in the process of syncing. But actually, we can see that in second 2, Node 2's note a had already been modified, and it was Missing a token. This shows that there's some time skew between the moment that the note is written to the moment that Obsidian reports sync is actually finished.
  • To track the note history across nodes, we take advantage of the note version counter in the Sync server. Version 1 of the note was reported by Node 1 on second 2. Then, on second 4, version 2 appears in both nodes, just after we allowed Node 2 to sync. Interestingly, we never catch a glimpse of version 1 in Node 2.
  • When the history operations finish on second 3 and sync status and version counter stabilize on second 4, everything looks finished. Only, the notes in both nodes are missing a token. The fuzzer then waits a grace period of 60 seconds, before judging data Loss on second 63.

See how timings change Sync behavior and hide the bug

I mentioned that this history needs all steps happening within 60 seconds. But why? Let's compare what happens if you wait e.g. 70 seconds just before disabling Airplane mode on the phone at step 7:

make run HISTORY=N2DN1AaWN2AaP70CW SAMPLING=everything

seconds    0      1      2   4             60           65         70
n1 ops     |   a  |W     ||  |             |  |  ||  |  | | |  |  ||  | |   ||
n1 sync    |    x |.   s ||. |             |. | .|| .| .| |.| .| .|| .| |.  ||.
n1 vers a  | -  1 |    . ||. |             |. | .|| .| .| |.| .| .|| .| |x  ||.
n1 file a  |    u |.   m ||m |             |m | m|| m| m| |m| m| m|| m| |m  ||.
n2 ops     |D     |  aP  ||  |     ...     |  |  ||  |  | | |  |  ||  |C|  W||
n2 sync    |     x|     x|| x|             | x|e ||e |e |e| |e |e ||e | | xe|| ss
n2 vers a  |  -  -| -   x|| x|             | x|? ||? |? |?| |? |? ||? | | ? || 1
n2 file a  |     -|     M|| m|             | m|m ||m |m |m| |m |m ||m | | m || c
  • The newly introduced Pause starts on second 1. This pushes the Connection of Node 2 to second 71.
  • The interesting part is what happens at second 61. Until 60, Obsidian was taking too long to respond to our sync status and version counter probes, so we were timing them out early (x) to avoid blocking . But at second 61, suddenly Obsidian started reporting sync error! That's when it noticed that the network is down.
  • When the network Connected again at 71, Obsidian eventually reconnected to Sync, exercising some error recovery path that didn't trigger this particular bug: the conflict file exists now, see the c at second 74; plus now version 1 is reported in Node 2!
  • Since the error happened at second 61, we can infer that from that moment on this particular bug won't be triggered. Hence the 60 second limit.

Pacing between cross-node edits to the same note

A history can edit the same note in different nodes: e.g., N1AaWN2Aa. This can be done conservatively (i.e., waiting for Obsidian to report it is synced, like in this example) or aggressively (as if the user typed into a note at the desktop and immediately afterwards typed into that same note on the phone). This is controlled via FORCED_TURNS, which contains ops that will be introduced in generated histories whenever a note is edited across different nodes. Some examples:

  • FORCED_TURNS=W (default) : before switching nodes, there is a Wait for confirmed upload of changes.
  • FORCED_TURNS=P60 : a Pause command (of 60s in this case) is inserted before switching nodes.
  • FORCED_TURNS= (empty) : cross-node edits can happen immediately. Unrealistic, but maybe useful as a stress test (...once Obsidian Sync can deal with easier scenarios).

Exercising sync recovery after disconnections

The main expected source of bugs is synchronization across nodes, particularly when the nodes get disconnected and reconnected to the network while the notes might keep changing. Just as if you edited a note on your phone on the go, while connectivity comes and goes.

CD_PROB defines the probability of Disconnect/Connect appearing in a history, causing a node going offline / online again.

The exact way in which nodes go offline is selected via ISOLATOR:

  • network: Default. Detach/attach the container from/to the container network, while keeping its IP.
  • sync: Obsidian-cli sync off / sync on commands. Note that this is unrealistically benevolent to Obsidian!

Using a Local node

Containers virtualize Obsidian clients that run on Linux. But what if we want to introduce a Mac client? Bugs might be different, so confirming reproducibility would be nice.

A possible future improvement could be to use a Mac VM (or even an iPhone simulator). But for now, if you are in a Mac, then you don't need to virtualize it! You can have the harness talk to your local Obsidian: just open the Sync vault, configure it to create Conflict files, and enable the CLI. Then you can use L in your histories, or use e.g. make soak NODES=n1,l LOCAL_VAULT=MyNotes so that generated histories include the local node.

Of course this means that while the histories are being run, your Obsidian client will be doing stuff to this vault. You should not disturb it (e.g. by changing the note in focus), so it's best to do this when you will not be using Obsidian yourself, e.g. during the night.

Outcomes

Automatic analysis

make analyze aggregates all the runs' results into tables in a file runs/analysis.md, to ease eyeballing of failure patterns across many histories and repetitions.

It also surfaces timing distributions, which might end up hinting at the reason why some history reps were OK while others lost data.

One can also comb manually through the logs. Read on for the gory details.

Logs, naming conventions and directories

The main idea is that eyeballing the runs/ directory should quickly allow you to see what histories were run, how many of their repetitions failed and why, and then zero in to the interesting cases:

  • At the top level there's directories named after each history: runs/<timestamps>-<history>[-<RESULT>]/
  • Inside of each history, there's the log for each of its repetitions: <timestamp>[-RESULT].jsonl.

Each log contains all the information needed to reconstruct the scenario.

Timestamps are formatted as DDTHHMMSS for ease of eyeballing and of cross-referencing. This will be helpful when you have dozens of files and directories and need to relate an Obsidian note against a particular history and repetition.

The notes that are created in Obsidian are named bughunt/<repTs>-<letter>-<history>, e.g. bughunt/26T181530-a-N1AaN2WAa.md. <repTs> is the repetition's timestamp, the trailing -<history> is the DSL string, and -<letter> is the DSL note letter the concrete note maps to. So e.g. a multi-note history (NOTES>1, HISTORY=AaAb) generates notes named …-a-…, …-b-….

If any of the history repetitions ended up in a non-OK state, its log's filename gets a suffix according to the failure:

rep suffixmeaning
(none)PASS
-LOSTa token was writen but disappeared. Data loss!
-DUPLa token is duplicated
-NOUPLOADa token was writen in a node but never reached the server
-OBSFAILobsidian-cli reports something but the filesystem disagrees
-UNKNOWNsome situation couldn't be recognised
-ENVFAILa container took too long to reconnect
-ABORTEDinterrupted by ^C

OBSFAIL, UNKNOWN and ENVFAIL mean that something is seriously wrong and needs special handling, so they are additionally logged to files in runs/{OBSFAIL, UNKNOWN, ENVFAIL}.log.

If a rep was non-OK, then the containing directory also gets a suffix -BAD<pct> indicating the % of repetitions that ended badly.

Judging whether there was data loss: token survival

Each command Ax in a history appends a unique token (<node>-<seq>-<note>) to the note x. At the end of the history, the oracle (src/oracle.ts) checks that those tokens are still there. It can detect 3 types of problems:

  • loss : a token was introduced but at the end of the history it's been lost;
  • duplication : a token is repeated;
  • divergence : nodes disagree on final content or conflict-file set.

Nodes run with Obsidian Sync in "create conflict file" mode, and the oracle checks for tokens either in the notes explicitly created during that history, or in any corresponding "Conflicted copy" created by Obsidian.

Cleaning up

The harness creates its notes in the bughunt/ folder of the Obsidian vault, and make clean-notes only ever deletes in that folder. So even if pointed at a real, in-use vault, the harness should keep your own notes safe. You should have backups, though.

The contents of the runs/ directory can be deleted at will. make clean-runs will do so.

make clean-data cleans both notes and logs.

Future ideas (?)

A reflection: Claude Code allows you to build ideas out very quickly. But many ideas should be discarded instead of built. Friction of idea implementation against reality used to help filter the craziest stuff out; if Claude Code removes that friction... what happens?

So here's is a dump of ideas that may, or may not, be interesting or cool to work on.

  • Obsidian is driven through its CLI, hoping that it behaves just like it would when driven through the GUI. There's an Obsidian headless option, currently in beta, that could also be interesting to try. Maybe it'll surface bugs differently to either the Linux or Mac GUI versions.
  • Obsidian Sync's auto-merge mode is not tested yet. Conflict file mode is the official recommendation in the Obsidian forums' thread about data loss, so I thought I'd start here.
  • Outcome judgment is very lenient towards Obsidian: as long as the input tokens are stored somewhere (actual note or conflict file), the result is considered OK. However, a real user surely wouldn't be happy if their inputs keep getting moved into conflict files randomly. So judgment should probably be made more... judgmental.
  • Both auto-merge and stricter judgment of conflict files would probably require keeping an internal model of acceptable results according to Obsidian Sync docs. That would probably be a big can of worms, given the closed-source nature of the beast and how little is pinned down in the docs.
  • I started this project inspired by Jepsen. Even if it's overkill for Obsidian Sync, there could be much to learn from it; plus there's a lot of other research on fuzzing a black box with semantics, surely also including internal models of legal outputs.
  • Relatedly, it'd be interesting to change the history generator so that it takes into account the failure rate of past histories to generate new ones, à la genetic algorithms. Just like AFL does.
  • It would be interesting to force network failures (packet loss) or slowness, once Obsidian Sync is solid enough over a well-behaved network.
  • In fact, the way in which Sync is blocked from working (network dis/connection vs obsidian-cli commands) changes the bugs found. This hints at Obsidian behaving specially on those commands. So, what if we added some new interruption mechanism, like suddenly killing Obsidian? (to model e.g. iOS quitting Obsidian because of memory pressure)
  • Interposing a MITM proxy on the Sync protocol might allow to have a reliable oracle of sync status, instead of just recording what Obsidian reports. That would allow to characterize client state independently of timers.
  • The code checking Obsidian Sync status could probably be made to work on other sync backends. Would e.g. Obsidian-on-iCloud lose more or less data? What about Syncthing, etc?
  • In fact, the very Obsidian driver could be made generic to work on other programs, like Logseq. That'd be kinda funny, given that I left Logseq because of how data-lossy it was.
  • The local node's purpose is to allow a Mac Obsidian client into the otherwise Linux mix. But since the local node works directly on the host's own Obsidian instance, this limits what can be done with it: e.g., no network faults (because it would also kill the containers' network). So it could be interesting to remove that local corner case and instead use tart to have a macOS VM, just as another ~container.
  • Another alternative would be to use macOS' pfctl to selectively block Obsidian Sync connections. But that gets into another can of worms with sudo, etc.
  • Conflict files are only supposed to appear in concrete Sync scenarios. The bugs found until now are pretty clearly about conflict files failing to be created by the Obsidian client. Tuning the pause lengths is an easy way to bias towards which client should create a conflict file. Therefore, could the pause time be enough to point to different bugs in the code?
  • Looks like there's some correlation between container CPU availability and some bugs' reproducibility. Could this reduce to pause length again?
  • Relatedly, given that Obsidian is closed-source, could the exact failure mode be reconstructed / reverse-engineered with DTrace / eBPF? or maybe something Electron-specific?

Reference

History Normalization

Histories are normalized to ensure they make sense, and so that those histories that would be similar in practice also look similar as a string. Example: history N1CPN2CAaAa would be reduced to N2PAa:

  • Histories start with all nodes connected, and redundant Dis/Connects are removed. (N1CPN2CAaAa → N1PN2AaAa)
  • A Pause not adjacent to an action (D/C/A/W) floats forward to the next action (N1PN2AaAa → N1N2PAaAa)
  • Redundant node selections vanish (N1N2PAaAa → N2PAaAa)
  • Contiguous Appends to the same note collapse into a single Append. (N2PAaAa → N2PAa)

Upgrading Obsidian

A new Obsidian version will eventually be released and you'll want to check if the bugs you found are still there.

make obsidian-latest                     # check latest GitHub .tar.gz release of Obsidian
make obsidian-upgrade                    # update the Obsidian version number that will be used
make containers-up                       # rebuild the image + relaunch the nodes
make check-assumptions                   # does everything look right?

The captured login information should keep working all the same.

Checking the assumptions this harness rests on

For our experiments to make sense, we depend on quite a few things that could change after a software update: the container engine and its networking behavior, the output format of the obsidian-cli command, the timings of it all, etc. So there's a Makefile target to check that everything still looks as expected.

make containers-up
make unpause-sync
make check-assumptions     # when coming back to the project after a long break, a software update, etc

Other auxiliary tools

make timeline-rep REP=... plots the rep's timeline from the data in the given log.

make probe-propagation runs histories with the goal of measuring the timings imposed by Obsidian: when is an edited note synced to the server? Are writes batched? How long until the other clients download it? As of 1.13.7, syncs happen immediately on first write, but subsequent ones are spaced to happen once every 10s per note.

make bench-cli measures the speed of running various Obsidian sampling commands in a container, sequentially or in parallel, batched or not. It helps ensure that the sampling mechanisms being used are still the fastest available.

Parameters to make and npm

There are many ways to fine-tune how things run, though the defaults are sane. (In fact, having so many available parameters feels like something I wouldn't do :P)

The table below shows the parameters available both at the make level (to be used as VAR=value: make soak FORCED_TURNS=P) and at the npm flag level (npm run start -- --forced-turns P).

Histories are generated by drawing ops randomly, one at a time. There's parameters to adjust the probability of each op, relative to A's 1.

make varCLI flagdefaultmeaning
HISTORY--history(generate)run a specific DSL string instead of generating
STEPS--steps—with HISTORY: run only its first N ops
REPEAT--repeat10repeats per history
HISTORIES--histories1number of histories to run (≤0 = until killed)
DURATION_MIN--duration-min—run for (at least) N minutes instead of a count
OPS--ops6-12edit-count range — counts A only; collapse may leave fewer. A single number (9) fixes the count (same as 9-9)
NOTES--notes1max number of potential notes being edited per generated history
FORCED_TURNS--forced-turnsWops inserted between edits across nodes (see Pacing section)
PAUSE_PROB--pause-prob0.3draw probability for a P
PAUSE_SEC--pause-sec10length of an ordinary pause
LONG_PAUSE_PROB--long-pause-prob0.25chance that an emitted pause is a long one
LONG_PAUSE_SEC--long-pause-sec100length of a long pause
WAIT_PROB--wait-prob0.2draw probability for a standalone W
CD_PROB--cd-prob0.4draw probability of a D/C
ISOLATOR--isolatornetworknetwork (partition) or sync (cooperative baseline)
NODES / NET / OBSIDIAN_BIN--nodes / --net / --binn1,n2 / obsidian-net / /opt/…container plumbing. NODES is only consulted when HISTORY is not set.
LOCAL_BIN--local-binobsidianpath to a local obsidian CLI binary, if used
LOCAL_NODE_ID--local-node-idOS's hostnamethe local instance's own Sync-reported device name, used to attribute its conflict files correctly
SKIP_HOST_CHECK--skip-host-checkoffdisable the checks ensuring that the host is online (at preflight and while waiting for sync settling)
POLL_SEC--poll-sec1how often (s) to re-read every node's state while waiting
MIN_FLOOR_SEC--min-floor-sec3observe at least this long before declaring done, to catch syncs slow to start after a Connect
CAP_SEC--cap-sec120how long to wait, once not-yet-settled, before also checking whether the host itself is offline
FINAL_SETTLE_SEC--final-settle-sec15end-of-history settle window; needs to cover a potential round-trip sync
PROBE_SEC--probe-sec5per-call cap on the settle's sync:status probe, in case it blocks
RUNS_DIR--runs-dircurrent pathparent dir for the whole runs/ tree
SKIP_SNAPSHOT--skip-snapshotoffskip the whole pause-snapshot mechanism (no extra CLI calls during a P), in case it's suspected of perturbing timings/results
CONTAINER_ENGINE-autodocker if it's on PATH, otherwise podman
RECONNECT_BUDGET_MS--reconnect-budget-ms1000if a container takes longer than this to reconnect, abort with ENVFAIL
PREFIX--prefix-fixed starting ops prefixed to histories
SAMPLING--samplingstrategicstrategic samples status and timings only where it's estimated to be necessary, hoping to avoid disturbing Obsidian. everything samples at every node, even if they aren't active in the history. everything-no-sleep samples everywhere, as frequently as possible.
DISPLAY--displayendend prints the timeline block after each rep, bar keeps it pinned as a live status bar above the scrolling log, off suppresses it
OPEN_NOTES--open-notesoffmake Obsidian open the note being edited in the GUI, allowing the history to be watched as it happens.
LOSS_GRACE_SEC--loss-grace-sec60when W detects that Sync is finished but a token is missing, it waits for this long before recording a case of data loss
LOCAL_VAULT--local-vault(none)vault the local Obsidian must have focused; required when the run includes L

hmijail/ObsSyncBugHunt

A semantic fuzzer to find bugs in Obsidian Sync as a distributed system

TypeScript

0

90 commits

updated Sep 30, 2026

See the code

See what people are saying

README

Obsidian Sync Bug Hunter

This is a semantic fuzzer for a simple distributed system (Obsidian Sync and its clients). In other words, a test harness that hunts for data loss in Obsidian Sync when the same note(s) are edited alternatively on multiple Obsidian instances.

Inspired by Jepsen, which would be overkill for something like Obsidian Sync.

I wrote this README personally. Everything else, including the docs/ directory, are Claude artifacts.

Lose data on your own Obsidian Sync here.
See the timeline of that example here, as reported by the fuzzer.

Some background: data loss in Obsidian Sync!?

Obsidian is a nice note-taking app. It's closed-source but free. It has a sync service, Obsidian Sync, which is subscription-based. This service has data-losing bugs. A thread in the Obsidian forums has been running for 2 years now gathering complaints, but the devs seem unable to find the problem. They proposed workarounds, but they fail too.

I lost data to Obsidian Sync too, and found that thread. I proposed using e.g. Jepsen to find bugs in a systematic way. There was no response.

I wanted some test project to use Claude Code on, so I told it I wanted to apply Jepsen to Obsidian. Claude jumped to make things happen; unfortunately those were pretty silly things. It quickly became clear that Claude needs its tasks to have much, much tighter scope.

Long story short, I started guiding the design of a semantic fuzzer, following the adversarial/paranoiac themes from my own DARUM, which (in its own way) also plays with randomness and repetitions to force an uncollaborative black box to reveal a bit of its inner workings. Plus containers and network control.

So 100% of the design is mine (and this README), but the code is 100% Claude's. In fact, I never used TypeScript; I chose it because it's a language used in the Obsidian ecosystem... and to force myself to stay hands-off and trust Claude. (I won't be repeating that)

Results

One result is that the fuzzer works: the harness finds different sequences of operations that trigger sync bugs in Obsidian, measures their repeatability and even helps understand how the sequence failed. Yay!

The other result is that Claude is a surprisingly, increasingly incompetent assistant. The project ballooned into 60h, half of it fighting the bug treadmill. Full experience report here, plus some posible remediations.

The summary is that keeping Claude in a leash tight enough to stop it from doing silly stuff is hard work, which eventually turns into a bug treadmill anyway. Claude is like an intern that knows far too much for their own good, uses that knowledge to make bad choices... plus periodically forgets instructions... but rarely lets go of pointless minutiae.
Also, you're responsible for what it remembers, even though you only have blunt tools to control that.
Also, those tools keep changing, and even Anthropic's instructions aren't very consistent, which confuses even Claude.

So that's a blurry mess. OK, but what did I learn from this project? Only things about Claude itself, the stuff that keeps changing. But nothing about the matter at hand. In fact, it's the opposite: I had to teach Claude how to build this.

If Claude was an intern, I would expect that they learnt something, and maybe even that they'd take over and keep the project moving forward. But Claude doesn't learn.

In a nutshell: this is an insta-legacy project, that ties you to LLMs, and required experience, but didn't create it.

Let's lose some data

As a motivating example, this is one sequence of operations found by the fuzzer that causes data loss ~100% of the time in Obsidian 1.12 and 1.13.7 (latest as of this writing). Reported 2 months ago, still not acknowledged nor fixed as of this writing.

Let’s assume you use Obsidian with Sync in your phone and your laptop. Both should set their Sync settings to “Conflict file” mode, which is the devs’ recommendation hoping to minimize data loss.

Note that this particular bug needs you to finish all the steps within 60 seconds! Later we’ll see why.

  1. Set your iPhone on airplane mode; ensure Wifi is also disconnected.
  2. In your laptop, created the note “buggy” (or whatever you want)
  3. In that note, type “laptop”
  4. Wait until Obsidian syncs up and shows the green sync icon (a couple of seconds)
  5. On your phone, create the same note “buggy”
  6. In that note, type “phone”
  7. Disable airplane mode on the phone and wait for sync.

After sync finishes, only the line “laptop” remains in the “buggy” note. That’s to be expected, since there was a conflict; the problem is that a Conflict File should have been created to save the “phone” line, but it didn’t. So that line is gone everywhere. And you didn’t dream it: looking at the Sync version history, you’ll see that the line did indeed reach the server.

The fuzzer shows how the bug happens with an ASCII timeline.

Fuzzer features

  • Sets up multiple Obsidian instances running in containers, prepared to use Obsidian Sync (requires an Obsidian Sync subscription)
  • Represents sequences of user actions as a compact string (aka history)
  • Can generate histories randomly, with configurabale complexity (number of nodes, number of notes, user's patience to wait for sync status, network failures)
  • Runs each history for a configurable number of repetitions
  • Samples whole-system behavior during the execution of histories, with configurable granularity / invasiveness
  • Gathers logs of everything for later analysis, and plots timelines summarizing them
  • Analyzes results, ranking histories by % of data-losing repetitions, plus flags suspicious behavior for further analysis
  • Generates shell scripts implementing a history, for verifiable reproducibility of bugs with minimal machinery
  • Semi-automatic upgradeability to new Obsidian versions, with self-checks to confirm whether the harness can still deal with the new version or requires modifications.

Prerequisites

  • Obsidian Sync subscription (if you want one just to test this project, know that it seems refundable during the first week)
  • Docker (in macOS, 2 vCPUs / 4GB RAM in the VM is enough for 2 Obsidian containers). Podman also works but seems to make sampling slower; prefer Apple Hypervisor to krunkit.
  • Node >= 22
  • Python 3
  • A VNC client to make the containerized Obsidian log in to your Sync account (and optionally to watch how Obsidian runs histories)
  • Curl
  • Optional:
    • Coreutils (for gtimeout)
    • Fnm (to pin down Node versions)
    • A local (non-containerized) Obsidian instance, mainly useful to find and test bugs in the macOS Obsidian version. (Needs its CLI activated)

On macOS you can install all the prerequisites with brew, and you can use the standard Screen Sharing.app for the VNC connection.

No LLM is used in the harness. Bugs found can't be hallucinations.

Quick start

The test harness will be creating and editing lots of notes on your Sync vault. It will try to keep the vault safe, by only ever acting on notes inside a folder ("bughunt") in your vault. (In any case you should backup your vault; personally, until bugs are fixed I moved my vault out of Sync and into iCloud Drive)

make is the easy entry point to the project, which maps to other tools as needed. make help lists every available command.

Common flow:

make install && make check        # install (npm ci) + typecheck + tests

# Create node and prepare it for Obsidian Sync's login:
make login
# Connect through VNC to the container (localhost:5901). A pristine Obsidian is waiting. Configure it to Sync to a vault and set it to "Create conflict file". Enable the Obsidian CLI.
make capture-login                # extracts the settings and login credentials into ./secrets
make containers-up                # launch n1 + n2 fresh with the captured credentials
make unpause-sync                 # let the nodes start syncing
make check-assumptions            # does everything look as expected? (run after updates, etc)
make clean-data                   # OPTIONAL clean slate: empty the vault + wipe runs/

make run HISTORY=N1AaWN2Aa REPEAT=3      # run one specific history
make soak HISTORY=N1AaWN2Aa              # soak one history until Ctrl-C
make soak                                # generate histories and run them until Ctrl-C
make analyze                             # aggregate runs/ into a report in runs/analysis.md

make repro HISTORY=N1DAaWN2AaC           # create a shell script to run that history (bug reproduction with minimal machinery)

How it all works

In a nutshell: we will set up a couple (or more!) Obsidian clients in containers, make them Sync, and then we'll create and edit notes in them, while checking that no data is lost. When it does, we'll record how repeatable is that case, in a way amenable to be reported to the Obsidian devs.

A set of Obsidian clients, ready to Sync

When you run make build-image, a container image will be prepared to run Obsidian. Then, make login will run it for you to connect through VNC. You will log in to your Obsidian Sync account, connect to a Sync vault, enable creation of conflict files, and enable the Obsidian CLI.

Next, make capture-login will extract from that container your Sync credentials and copy them into the secrets directory. This is so that multiple containers can reuse the same credentials, without manually setting up each one of them individually. This information never leaves your computer. Don't publish that directory; it's already in .gitignore.

And then, make containers-up will start your nodes and get them syncing. If there's notes in the vault, they might take some time to finish their initial sync.

If you want to see or interact with them, you can always connect to each node through VNC, in the port 5900+(node number).

Generating sequences of Obsidian edits with a tiny DSL

We will test sequences of actions, AKA histories. A history is a string of user actions replayed against multiple Obsidian nodes. Commands are uppercase, parameters lowercase/digits.

Commandmeaning
N<d>set the active node (N1, N2)
Lset the active node to local Obsidian instance
A<x>make the active node append a uniquely-tagged line to note x (creates it if necessary)
D / Cdisconnect / connect the active node from the network (applies to containers, not the Local node)
W[n]wait until the active node is synced (with optional timeout of n seconds)
P[n]pause ~n seconds (default 10)

Example: N1AaWN2Aa= node 1 appends to note a, waits for sync; node 2 appends to the same note.

At the end of the history, the harness reconnects all nodes to the network, waits for them all to report synced, and still waits for a settling window to ensure that no further changes happen (e.g. generation of conflict files, which later get synced, which could still generate conflicts in other clients, etc etc). An end result is judged only when things remain stable for FINAL_SETTLE_SEC seconds.

Note that the harness models a single user using Obsidian across multiple devices, so there's a single thread of control doing everything. This means that e.g. a Pause command applies across all nodes at once: the control thread does nothing for n seconds, while the Obsidian nodes will keep doing their thing (e.g. try to sync). Similarly, W waits for the current node to be synced, which doesn't stop the rest of nodes from working.

Histories can be auto-generated randomly (each op's probability is adjustable, see Parameters) or typed manually. They are normalized so that histories that would be similar in practice also look similar as a string (see the History Normalization section).

Timings are unavoidably variable across repetitions of a history, since we don't have control of the Sync server, internet traffic, etc. This might cause Sync results to change every time you repeat the history. Therefore histories are run for REPEAT times to sample the distribution of end results. If one wants to minimize variability, command W waits until Obsidian the node is synced.

Edits to notes are append-only (for now?), since that is a case supported by the CLI. Edits optionally (OPEN_NOTES) can also open the note in the nodes' GUI, so you can watch the history unfold through a VNC connection if you want to.

One key difficulty: W formalizes (somewhat) the synced status

Syncing is hard. It can happen that the user is editing stuff while changes are coming in or out of Obsidian. It can even happen that something is wrong with the network and syncing gets delayed, accumulating conflicting changes at different points of the process, while you, the user, might not know or even care that this is happening.

The user can (in theory) count on Obsidian to deal with this. But for the fuzzer to do its job better we need to try to formalize the different underlying scenarios. Hence, W doesn't just trust the Sync status reported by Obsidian (which corresponds to the Sync icon in the GUI), but also checks metadata from the Sync server and content of modified notes; it will only let the history continue when the synced status does seem correct (and will otherwise record a data loss bug if it looks wrong after a configurable grace period).

On the other hand, Obsidian Sync bugs might also get triggered by an impatient user who tries editing notes while the Sync status is not settled. This case is covered by W<n>. For example, W10 simulates a user who waits for a maximum of 10 seconds for the GUI Sync icon to get green; after that, they just go ahead and edit a note anyway. (Which makes sense, because sometimes the icon changes are hard to see).

Practical example

The motivating example at the beginning of this README was found by the fuzzer as this history: N2DN1AaWN2AaCW

(Found in Obsidian 1.12.7, still there in 1.13.7)

  • N2: selects N2 as the current node
  • D : disconnects the current node
  • N1: selects N1 as the current node
  • Aa: appends a token to note "a" in node 1 (this is a "logical name"; see section on naming of notes)
  • W : wait until the current note in the current node is reported as synced by Obsidian
  • N2: selects N2
  • Aa: appends a new token to note "a" in node 2 (same "logical name" as before)
  • C : connect the current node
  • W : wait for sync

Timelines

Running a history creates long logs. For easier inspection, these are summarized into an ASCII timeline that shows what happened, where and when.

For example, this is the timeline generated by the motivating example:

make run HISTORY=N2DN1AaWN2AaCW SAMPLING=everything

seconds    0    1      2  3 4   5    7            60
n1 ops     |   a|  W   |  | |   ||   |            |   |   ||    |
n1 sync    |    |x .   |x | |.  ||.  |            |.  |.  ||x   |.
n1 vers a  | -  |x     |1 | |2  ||.  |            |.  |x  ||.   |.
n1 file a  |    |x .   |M | |M  ||m  |            |m  |m  ||m   |m
n2 ops     |D   |    aC|  |W|   ||   |    ...     |   |   ||    |
n2 sync    |    | x    | x|s| ..|| ..|            | x.| x.|| x. | x
n2 vers a  |  - | x -  | x| | 2 || . |            | . | . || .  | .
n2 file a  |    | x    | M| | m || m |            | m | m || m L| m

How to read that?

TImelines consist of lanes (rows), with each lane dedicated to one type of data we sample. In each lane, various letters represent events. Samples that are taken at the same time are plotted in the same column. For ease of reference, seconds are separated by columns of |.

Letter meanings:

  • general: x the call was cut off · ? reply not recognised · not sampled · | a wallclock second
  • ops:a/b/… a note was written to · D/C disconnect/connect · W a wait started · P a pause · h a W<n> handed off early
  • sync: . synced · s syncing · p paused · e error · o offline · h stopped
  • vers: a digit when the server version counter moved, to that value · . sampled but unchanged · - no server history yet
  • file (per node per note) — . file is complete and unchanged · u file changed here, is complete · m missing a token · M changed here, missing a token · c changed here and a conflict file appeared · C same, but some token is missing · - no file · L loss declared · ! readings that do not add up

A column where nothing was sampled will be empty (||). In general, dots mean "we sampled here but wasn't interesting".

Data loss, play by play

Can you see now what happened in that example history?

  • The operations actually finished on second 3: Node 2 waited for sync to happen, and at that moment, Obsidian reported it was in the process of syncing. But actually, we can see that in second 2, Node 2's note a had already been modified, and it was Missing a token. This shows that there's some time skew between the moment that the note is written to the moment that Obsidian reports sync is actually finished.
  • To track the note history across nodes, we take advantage of the note version counter in the Sync server. Version 1 of the note was reported by Node 1 on second 2. Then, on second 4, version 2 appears in both nodes, just after we allowed Node 2 to sync. Interestingly, we never catch a glimpse of version 1 in Node 2.
  • When the history operations finish on second 3 and sync status and version counter stabilize on second 4, everything looks finished. Only, the notes in both nodes are missing a token. The fuzzer then waits a grace period of 60 seconds, before judging data Loss on second 63.

See how timings change Sync behavior and hide the bug

I mentioned that this history needs all steps happening within 60 seconds. But why? Let's compare what happens if you wait e.g. 70 seconds just before disabling Airplane mode on the phone at step 7:

make run HISTORY=N2DN1AaWN2AaP70CW SAMPLING=everything

seconds    0      1      2   4             60           65         70
n1 ops     |   a  |W     ||  |             |  |  ||  |  | | |  |  ||  | |   ||
n1 sync    |    x |.   s ||. |             |. | .|| .| .| |.| .| .|| .| |.  ||.
n1 vers a  | -  1 |    . ||. |             |. | .|| .| .| |.| .| .|| .| |x  ||.
n1 file a  |    u |.   m ||m |             |m | m|| m| m| |m| m| m|| m| |m  ||.
n2 ops     |D     |  aP  ||  |     ...     |  |  ||  |  | | |  |  ||  |C|  W||
n2 sync    |     x|     x|| x|             | x|e ||e |e |e| |e |e ||e | | xe|| ss
n2 vers a  |  -  -| -   x|| x|             | x|? ||? |? |?| |? |? ||? | | ? || 1
n2 file a  |     -|     M|| m|             | m|m ||m |m |m| |m |m ||m | | m || c
  • The newly introduced Pause starts on second 1. This pushes the Connection of Node 2 to second 71.
  • The interesting part is what happens at second 61. Until 60, Obsidian was taking too long to respond to our sync status and version counter probes, so we were timing them out early (x) to avoid blocking . But at second 61, suddenly Obsidian started reporting sync error! That's when it noticed that the network is down.
  • When the network Connected again at 71, Obsidian eventually reconnected to Sync, exercising some error recovery path that didn't trigger this particular bug: the conflict file exists now, see the c at second 74; plus now version 1 is reported in Node 2!
  • Since the error happened at second 61, we can infer that from that moment on this particular bug won't be triggered. Hence the 60 second limit.

Pacing between cross-node edits to the same note

A history can edit the same note in different nodes: e.g., N1AaWN2Aa. This can be done conservatively (i.e., waiting for Obsidian to report it is synced, like in this example) or aggressively (as if the user typed into a note at the desktop and immediately afterwards typed into that same note on the phone). This is controlled via FORCED_TURNS, which contains ops that will be introduced in generated histories whenever a note is edited across different nodes. Some examples:

  • FORCED_TURNS=W (default) : before switching nodes, there is a Wait for confirmed upload of changes.
  • FORCED_TURNS=P60 : a Pause command (of 60s in this case) is inserted before switching nodes.
  • FORCED_TURNS= (empty) : cross-node edits can happen immediately. Unrealistic, but maybe useful as a stress test (...once Obsidian Sync can deal with easier scenarios).

Exercising sync recovery after disconnections

The main expected source of bugs is synchronization across nodes, particularly when the nodes get disconnected and reconnected to the network while the notes might keep changing. Just as if you edited a note on your phone on the go, while connectivity comes and goes.

CD_PROB defines the probability of Disconnect/Connect appearing in a history, causing a node going offline / online again.

The exact way in which nodes go offline is selected via ISOLATOR:

  • network: Default. Detach/attach the container from/to the container network, while keeping its IP.
  • sync: Obsidian-cli sync off / sync on commands. Note that this is unrealistically benevolent to Obsidian!

Using a Local node

Containers virtualize Obsidian clients that run on Linux. But what if we want to introduce a Mac client? Bugs might be different, so confirming reproducibility would be nice.

A possible future improvement could be to use a Mac VM (or even an iPhone simulator). But for now, if you are in a Mac, then you don't need to virtualize it! You can have the harness talk to your local Obsidian: just open the Sync vault, configure it to create Conflict files, and enable the CLI. Then you can use L in your histories, or use e.g. make soak NODES=n1,l LOCAL_VAULT=MyNotes so that generated histories include the local node.

Of course this means that while the histories are being run, your Obsidian client will be doing stuff to this vault. You should not disturb it (e.g. by changing the note in focus), so it's best to do this when you will not be using Obsidian yourself, e.g. during the night.

Outcomes

Automatic analysis

make analyze aggregates all the runs' results into tables in a file runs/analysis.md, to ease eyeballing of failure patterns across many histories and repetitions.

It also surfaces timing distributions, which might end up hinting at the reason why some history reps were OK while others lost data.

One can also comb manually through the logs. Read on for the gory details.

Logs, naming conventions and directories

The main idea is that eyeballing the runs/ directory should quickly allow you to see what histories were run, how many of their repetitions failed and why, and then zero in to the interesting cases:

  • At the top level there's directories named after each history: runs/<timestamps>-<history>[-<RESULT>]/
  • Inside of each history, there's the log for each of its repetitions: <timestamp>[-RESULT].jsonl.

Each log contains all the information needed to reconstruct the scenario.

Timestamps are formatted as DDTHHMMSS for ease of eyeballing and of cross-referencing. This will be helpful when you have dozens of files and directories and need to relate an Obsidian note against a particular history and repetition.

The notes that are created in Obsidian are named bughunt/<repTs>-<letter>-<history>, e.g. bughunt/26T181530-a-N1AaN2WAa.md. <repTs> is the repetition's timestamp, the trailing -<history> is the DSL string, and -<letter> is the DSL note letter the concrete note maps to. So e.g. a multi-note history (NOTES>1, HISTORY=AaAb) generates notes named …-a-…, …-b-….

If any of the history repetitions ended up in a non-OK state, its log's filename gets a suffix according to the failure:

rep suffixmeaning
(none)PASS
-LOSTa token was writen but disappeared. Data loss!
-DUPLa token is duplicated
-NOUPLOADa token was writen in a node but never reached the server
-OBSFAILobsidian-cli reports something but the filesystem disagrees
-UNKNOWNsome situation couldn't be recognised
-ENVFAILa container took too long to reconnect
-ABORTEDinterrupted by ^C

OBSFAIL, UNKNOWN and ENVFAIL mean that something is seriously wrong and needs special handling, so they are additionally logged to files in runs/{OBSFAIL, UNKNOWN, ENVFAIL}.log.

If a rep was non-OK, then the containing directory also gets a suffix -BAD<pct> indicating the % of repetitions that ended badly.

Judging whether there was data loss: token survival

Each command Ax in a history appends a unique token (<node>-<seq>-<note>) to the note x. At the end of the history, the oracle (src/oracle.ts) checks that those tokens are still there. It can detect 3 types of problems:

  • loss : a token was introduced but at the end of the history it's been lost;
  • duplication : a token is repeated;
  • divergence : nodes disagree on final content or conflict-file set.

Nodes run with Obsidian Sync in "create conflict file" mode, and the oracle checks for tokens either in the notes explicitly created during that history, or in any corresponding "Conflicted copy" created by Obsidian.

Cleaning up

The harness creates its notes in the bughunt/ folder of the Obsidian vault, and make clean-notes only ever deletes in that folder. So even if pointed at a real, in-use vault, the harness should keep your own notes safe. You should have backups, though.

The contents of the runs/ directory can be deleted at will. make clean-runs will do so.

make clean-data cleans both notes and logs.

Future ideas (?)

A reflection: Claude Code allows you to build ideas out very quickly. But many ideas should be discarded instead of built. Friction of idea implementation against reality used to help filter the craziest stuff out; if Claude Code removes that friction... what happens?

So here's is a dump of ideas that may, or may not, be interesting or cool to work on.

  • Obsidian is driven through its CLI, hoping that it behaves just like it would when driven through the GUI. There's an Obsidian headless option, currently in beta, that could also be interesting to try. Maybe it'll surface bugs differently to either the Linux or Mac GUI versions.
  • Obsidian Sync's auto-merge mode is not tested yet. Conflict file mode is the official recommendation in the Obsidian forums' thread about data loss, so I thought I'd start here.
  • Outcome judgment is very lenient towards Obsidian: as long as the input tokens are stored somewhere (actual note or conflict file), the result is considered OK. However, a real user surely wouldn't be happy if their inputs keep getting moved into conflict files randomly. So judgment should probably be made more... judgmental.
  • Both auto-merge and stricter judgment of conflict files would probably require keeping an internal model of acceptable results according to Obsidian Sync docs. That would probably be a big can of worms, given the closed-source nature of the beast and how little is pinned down in the docs.
  • I started this project inspired by Jepsen. Even if it's overkill for Obsidian Sync, there could be much to learn from it; plus there's a lot of other research on fuzzing a black box with semantics, surely also including internal models of legal outputs.
  • Relatedly, it'd be interesting to change the history generator so that it takes into account the failure rate of past histories to generate new ones, à la genetic algorithms. Just like AFL does.
  • It would be interesting to force network failures (packet loss) or slowness, once Obsidian Sync is solid enough over a well-behaved network.
  • In fact, the way in which Sync is blocked from working (network dis/connection vs obsidian-cli commands) changes the bugs found. This hints at Obsidian behaving specially on those commands. So, what if we added some new interruption mechanism, like suddenly killing Obsidian? (to model e.g. iOS quitting Obsidian because of memory pressure)
  • Interposing a MITM proxy on the Sync protocol might allow to have a reliable oracle of sync status, instead of just recording what Obsidian reports. That would allow to characterize client state independently of timers.
  • The code checking Obsidian Sync status could probably be made to work on other sync backends. Would e.g. Obsidian-on-iCloud lose more or less data? What about Syncthing, etc?
  • In fact, the very Obsidian driver could be made generic to work on other programs, like Logseq. That'd be kinda funny, given that I left Logseq because of how data-lossy it was.
  • The local node's purpose is to allow a Mac Obsidian client into the otherwise Linux mix. But since the local node works directly on the host's own Obsidian instance, this limits what can be done with it: e.g., no network faults (because it would also kill the containers' network). So it could be interesting to remove that local corner case and instead use tart to have a macOS VM, just as another ~container.
  • Another alternative would be to use macOS' pfctl to selectively block Obsidian Sync connections. But that gets into another can of worms with sudo, etc.
  • Conflict files are only supposed to appear in concrete Sync scenarios. The bugs found until now are pretty clearly about conflict files failing to be created by the Obsidian client. Tuning the pause lengths is an easy way to bias towards which client should create a conflict file. Therefore, could the pause time be enough to point to different bugs in the code?
  • Looks like there's some correlation between container CPU availability and some bugs' reproducibility. Could this reduce to pause length again?
  • Relatedly, given that Obsidian is closed-source, could the exact failure mode be reconstructed / reverse-engineered with DTrace / eBPF? or maybe something Electron-specific?

Reference

History Normalization

Histories are normalized to ensure they make sense, and so that those histories that would be similar in practice also look similar as a string. Example: history N1CPN2CAaAa would be reduced to N2PAa:

  • Histories start with all nodes connected, and redundant Dis/Connects are removed. (N1CPN2CAaAa → N1PN2AaAa)
  • A Pause not adjacent to an action (D/C/A/W) floats forward to the next action (N1PN2AaAa → N1N2PAaAa)
  • Redundant node selections vanish (N1N2PAaAa → N2PAaAa)
  • Contiguous Appends to the same note collapse into a single Append. (N2PAaAa → N2PAa)

Upgrading Obsidian

A new Obsidian version will eventually be released and you'll want to check if the bugs you found are still there.

make obsidian-latest                     # check latest GitHub .tar.gz release of Obsidian
make obsidian-upgrade                    # update the Obsidian version number that will be used
make containers-up                       # rebuild the image + relaunch the nodes
make check-assumptions                   # does everything look right?

The captured login information should keep working all the same.

Checking the assumptions this harness rests on

For our experiments to make sense, we depend on quite a few things that could change after a software update: the container engine and its networking behavior, the output format of the obsidian-cli command, the timings of it all, etc. So there's a Makefile target to check that everything still looks as expected.

make containers-up
make unpause-sync
make check-assumptions     # when coming back to the project after a long break, a software update, etc

Other auxiliary tools

make timeline-rep REP=... plots the rep's timeline from the data in the given log.

make probe-propagation runs histories with the goal of measuring the timings imposed by Obsidian: when is an edited note synced to the server? Are writes batched? How long until the other clients download it? As of 1.13.7, syncs happen immediately on first write, but subsequent ones are spaced to happen once every 10s per note.

make bench-cli measures the speed of running various Obsidian sampling commands in a container, sequentially or in parallel, batched or not. It helps ensure that the sampling mechanisms being used are still the fastest available.

Parameters to make and npm

There are many ways to fine-tune how things run, though the defaults are sane. (In fact, having so many available parameters feels like something I wouldn't do :P)

The table below shows the parameters available both at the make level (to be used as VAR=value: make soak FORCED_TURNS=P) and at the npm flag level (npm run start -- --forced-turns P).

Histories are generated by drawing ops randomly, one at a time. There's parameters to adjust the probability of each op, relative to A's 1.

make varCLI flagdefaultmeaning
HISTORY--history(generate)run a specific DSL string instead of generating
STEPS--steps—with HISTORY: run only its first N ops
REPEAT--repeat10repeats per history
HISTORIES--histories1number of histories to run (≤0 = until killed)
DURATION_MIN--duration-min—run for (at least) N minutes instead of a count
OPS--ops6-12edit-count range — counts A only; collapse may leave fewer. A single number (9) fixes the count (same as 9-9)
NOTES--notes1max number of potential notes being edited per generated history
FORCED_TURNS--forced-turnsWops inserted between edits across nodes (see Pacing section)
PAUSE_PROB--pause-prob0.3draw probability for a P
PAUSE_SEC--pause-sec10length of an ordinary pause
LONG_PAUSE_PROB--long-pause-prob0.25chance that an emitted pause is a long one
LONG_PAUSE_SEC--long-pause-sec100length of a long pause
WAIT_PROB--wait-prob0.2draw probability for a standalone W
CD_PROB--cd-prob0.4draw probability of a D/C
ISOLATOR--isolatornetworknetwork (partition) or sync (cooperative baseline)
NODES / NET / OBSIDIAN_BIN--nodes / --net / --binn1,n2 / obsidian-net / /opt/…container plumbing. NODES is only consulted when HISTORY is not set.
LOCAL_BIN--local-binobsidianpath to a local obsidian CLI binary, if used
LOCAL_NODE_ID--local-node-idOS's hostnamethe local instance's own Sync-reported device name, used to attribute its conflict files correctly
SKIP_HOST_CHECK--skip-host-checkoffdisable the checks ensuring that the host is online (at preflight and while waiting for sync settling)
POLL_SEC--poll-sec1how often (s) to re-read every node's state while waiting
MIN_FLOOR_SEC--min-floor-sec3observe at least this long before declaring done, to catch syncs slow to start after a Connect
CAP_SEC--cap-sec120how long to wait, once not-yet-settled, before also checking whether the host itself is offline
FINAL_SETTLE_SEC--final-settle-sec15end-of-history settle window; needs to cover a potential round-trip sync
PROBE_SEC--probe-sec5per-call cap on the settle's sync:status probe, in case it blocks
RUNS_DIR--runs-dircurrent pathparent dir for the whole runs/ tree
SKIP_SNAPSHOT--skip-snapshotoffskip the whole pause-snapshot mechanism (no extra CLI calls during a P), in case it's suspected of perturbing timings/results
CONTAINER_ENGINE-autodocker if it's on PATH, otherwise podman
RECONNECT_BUDGET_MS--reconnect-budget-ms1000if a container takes longer than this to reconnect, abort with ENVFAIL
PREFIX--prefix-fixed starting ops prefixed to histories
SAMPLING--samplingstrategicstrategic samples status and timings only where it's estimated to be necessary, hoping to avoid disturbing Obsidian. everything samples at every node, even if they aren't active in the history. everything-no-sleep samples everywhere, as frequently as possible.
DISPLAY--displayendend prints the timeline block after each rep, bar keeps it pinned as a live status bar above the scrolling log, off suppresses it
OPEN_NOTES--open-notesoffmake Obsidian open the note being edited in the GUI, allowing the history to be watched as it happens.
LOSS_GRACE_SEC--loss-grace-sec60when W detects that Sync is finished but a token is missing, it waits for this long before recording a case of data loss
LOCAL_VAULT--local-vault(none)vault the local Obsidian must have focused; required when the run includes L

Languages

TypeScript

79.1%

Shell

14.7%

Makefile

5.2%