ArcadeData/ldbc_graphalytics_platforms_arcadedb

LDBC Graphalytics benchmark platform driver for ArcadeDB + Most Popular DBMS

Python

3

136 commits

updated Oct 6, 2026

See the code

README

LDBC Graphalytics ArcadeDB Platform Driver

Platform driver implementation for the LDBC Graphalytics benchmark using ArcadeDB.

Uses ArcadeDB in embedded mode with the Graph Analytical View (GAV) engine, which builds a CSR (Compressed Sparse Row) adjacency index for high-performance graph algorithm execution with zero GC pressure.

This repository contains three benchmark modes:

  1. Official LDBC Graphalytics — standardized framework with per-algorithm isolation, validation, and reporting
  2. Native multi-vendor comparison — load once, run all algorithms, compare ArcadeDB vs Kuzu vs DuckPGQ vs Memgraph vs Neo4j vs FalkorDB vs HugeGraph
  3. LSQB (Labelled Subgraph Query Benchmark) — 9 subgraph pattern matching queries on the LDBC SNB social network, comparing ArcadeDB (Cypher) vs DuckDB (SQL) vs FalkorDB (Cypher) and others

Supported Algorithms

AlgorithmImplementationComplexity
BFS (Breadth-First Search)Parallel frontier expansion with bitmap visited set and push/pull direction optimizationO(V + E)
PR (PageRank)Pull-based parallel iteration via backward CSRO(iterations * E)
WCC (Weakly Connected Components)Synchronous parallel min-label propagationO(diameter * E)
CDLP (Community Detection Label Propagation)Synchronous parallel label propagation with sort-based mode findingO(iterations * E * log(d))
LCC (Local Clustering Coefficient)Parallel sorted-merge triangle countingO(E * sqrt(E))
SSSP (Single Source Shortest Paths)Dijkstra with binary min-heap on CSR + columnar weightsO((V + E) * log(V))

Prerequisites

  • Java 21 or later (required for jdk.incubator.vector SIMD support). The published benchmark numbers use Eclipse Temurin 25 with -XX:+UseCompactObjectHeaders (GraalVM is not used)
  • Maven 3.x
  • Python 3.10+ (for Mode 2 and Mode 3 multi-vendor comparisons; see pyproject.toml for per-vendor install extras)

Build

mvn package -DskipTests

The build produces a self-contained distribution in graphalytics-1.3.0-arcadedb-0.1-SNAPSHOT/.

Dataset

Use the built-in dataset manager to browse and download datasets from the LDBC data repository:

# See all available datasets (40+ Graphalytics + 9 LSQB scale factors)
python3 datasets.py available

# Download the standard Graphalytics benchmark dataset (633K vertices, 34M edges, ~155 MB)
python3 datasets.py download datagen-7_5-fb

# Download the LSQB social network dataset (SF1, ~3.9M vertices, ~17.9M edges)
python3 datasets.py download lsqb-sf1

# Show downloaded datasets with size and vertex/edge counts
python3 datasets.py

Datasets are downloaded into the datasets/ directory (git-ignored). After downloading datagen-7_5-fb:

datasets/
  datagen-7_5-fb/
    datagen-7_5-fb.v              # vertex file (one ID per line)
    datagen-7_5-fb.e              # edge file (src dst weight, space-separated)
    datagen-7_5-fb.properties     # graph metadata
    datagen-7_5-fb-BFS/           # validation data per algorithm
    datagen-7_5-fb-WCC/
    ...

Weekly unattended run

python3 weekend.py --jvm-flags "-XX:+UseCompactObjectHeaders"

Runs every suite and vendor in sequence (ArcadeDB embedded Java benchmarks, the Graphalytics and LSQB multi-vendor suites, optionally Mode 1 with --mode1-dist) and writes a Markdown/JSON report to weekly-results/<timestamp>/. Each vendor runs in its own process group with a total-time and a no-output limit; a hung vendor gets SIGTERM, then SIGKILL, and the run continues. Loaded databases are kept under ~/.cache/ldbc-graph-bench and reused, so only the first run (or --reset, a new image, a changed dataset) pays for the loads. See CLAUDE.md for the details and python3 -m unittest discover -s tests for the tests.

Benchmarks and results

Every number is a warm number (first call discarded, median of 3 timed runs), measured one system at a time on AC power and validated against the official LDBC reference outputs (or the official expected counts for LSQB). A value marked ✗ failed that check and is not ranked; N/A = no implementation; timeout = the 5-minute limit per operation. Footnote symbols (§ ‡ ¶ *) are explained on the detailed pages. Machine: MacBook Pro 16" (2026), Apple M5 Pro, 48 GB RAM; ArcadeDB on Eclipse Temurin 25 with -XX:+UseCompactObjectHeaders, 12 GB heap for every JVM system, Docker Desktop with 32 GB.

BenchmarkWhat it measuresDetailed page
Graphalytics, official framework (Mode 1)ArcadeDB only, official LDBC harness, one graph load per algorithmdocs/benchmark-graphalytics-official.md
Graphalytics, multi-vendor (Mode 2), datagen-7_5-fb6 algorithms, 10 systems, 633K vertices / 34M edgesdocs/benchmark-graphalytics-multivendor.md
Graphalytics, multi-vendor, graph500-225 algorithms, 10 systems, 2.4M vertices / 64M edgesdocs/benchmark-graphalytics-graph500-22.md
LSQB SF1 (Mode 3)9 subgraph pattern matching queries, 7+ systemsdocs/benchmark-lsqb.md

How ArcadeDB itself changes from release to release (official framework, Mode 2 and LSQB): ArcadeDB-release-progress.md.

Graphalytics, datagen-7_5-fb (633,432 vertices, 34,185,747 edges)

Seconds, last column peak memory in GiB. Bold marks the fastest valid result per column. Systems, versions, how we measure, per-system notes and how to run each vendor: detailed page.

SystemLoadPageRankWCCBFSLCCSSSPCDLPPeak memory (GiB)
ArcadeDB embedded71.90.0850.0030.0202.050.750.965.6
ArcadeDB Docker43.90.160.020.062.611.391.4512.3
Neo4j6576.98§0.1110.480‡15.4N/AN/A13.3
Kuzu28.81.160.4340.328N/AN/AN/A0.87
LadybugDB5.16N/AN/A7.89N/AN/AN/A1.0
DuckPGQ0.851.51✗2.00timeout¶49.3N/AN/A8.3
Memgraph4375.491893.85N/A75.9timeout25.1
ArangoDB *72693.040.938.4N/A173254✗23.2
FalkorDB1162.67✗2.500.057N/AN/A10.6✗8.4
HugeGraph34.72.410.2930.195110N/A22.3✗3.2

Graphalytics, graph500-22 (2,396,657 vertices, 64,155,735 edges)

Seconds; no SSSP (not defined for this dataset). ArcadeDB is 26.11.1-SNAPSHOT; bold marks the fastest valid result per column. Per-system notes, the harness problems found and how to reproduce: detailed page.

SystemLoadPageRankWCCBFSLCCCDLPPeak memory (GiB)
ArcadeDB embedded114.50.2680.0130.07646.91.781.6**
ArcadeDB Docker101†0.590.090.2460.14.3313.5
Neo4j2017†12.50.201.05timeoutN/A13.3
Kuzu53.8†3.131.260.94N/AN/A5.9
LadybugDB10.1†N/AN/A16.7N/AN/A2.0
DuckPGQ0.57†9.27✗4.94timeout‡timeoutN/A15.7
Memgraph932§31.0OOM12.5N/AOOM25.0
ArangoDB 3.11.141700†261.9106.4OOM¶N/Atimeout30.6
FalkorDB348.57.97✗7.300.151N/A39.8✗15.9
HugeGraph34.2†6.030.670.42timeout44.7✗8.8

OOM = out of memory (Docker Desktop's 32 GB). † remembered original load, ‡ DuckPGQ BFS does not finish (ended after 20 min), § Memgraph load includes the reverse-edge step, ¶ ArangoDB container ran out of memory, ** live heap after GC (process peak not measured): see the detailed page.

LSQB SF1 (3,947,829 vertices, 17,882,623 edges)

Seconds; ArcadeDB embedded is shown with the Graph Analytical View (OLAP) and without it (OLTP); bold marks the fastest result per query. All counts match the official expected output. Queries, run commands and analysis: detailed page.

SystemLoadQ1Q2Q3Q4Q5Q6Q7Q8Q9Peak GiB
ArcadeDB OLAP119.50.080.130.050.040.170.070.040.110.304.3
ArcadeDB OLTP159.22.454.403.171.0611.3911.001.076.910.8712.0
ArcadeDB Docker99.40.170.170.060.050.250.130.060.140.3312.9
DuckDB0.460.110.010.040.060.041.840.070.076.030.9
Kuzu2.444.610.152.30N/AN/A1.38N/AN/A6.395.1
LadybugDB2.950.120.1010.440.170.180.660.410.320.078.8
Neo4j252.04.991.6310.766.285.6628.498.0412.39254.2913.3
PostgreSQL15.19.170.651.606.352.1715.5011.163.3950.620.5
Memgraph222.258.06timeouttimeout4.303.54121.154.923.02timeout3.1
FalkorDB659.336.4353.784.593.533.9340.9347.964.90225.222.4

More

License

Apache License, Version 2.0

algorithms
benchmark
graph
graph-database
graphdb
performance-testing

ArcadeData/ldbc_graphalytics_platforms_arcadedb

LDBC Graphalytics benchmark platform driver for ArcadeDB + Most Popular DBMS

Python

3

136 commits

updated Oct 6, 2026

See the code

README

LDBC Graphalytics ArcadeDB Platform Driver

Platform driver implementation for the LDBC Graphalytics benchmark using ArcadeDB.

Uses ArcadeDB in embedded mode with the Graph Analytical View (GAV) engine, which builds a CSR (Compressed Sparse Row) adjacency index for high-performance graph algorithm execution with zero GC pressure.

This repository contains three benchmark modes:

  1. Official LDBC Graphalytics — standardized framework with per-algorithm isolation, validation, and reporting
  2. Native multi-vendor comparison — load once, run all algorithms, compare ArcadeDB vs Kuzu vs DuckPGQ vs Memgraph vs Neo4j vs FalkorDB vs HugeGraph
  3. LSQB (Labelled Subgraph Query Benchmark) — 9 subgraph pattern matching queries on the LDBC SNB social network, comparing ArcadeDB (Cypher) vs DuckDB (SQL) vs FalkorDB (Cypher) and others

Supported Algorithms

AlgorithmImplementationComplexity
BFS (Breadth-First Search)Parallel frontier expansion with bitmap visited set and push/pull direction optimizationO(V + E)
PR (PageRank)Pull-based parallel iteration via backward CSRO(iterations * E)
WCC (Weakly Connected Components)Synchronous parallel min-label propagationO(diameter * E)
CDLP (Community Detection Label Propagation)Synchronous parallel label propagation with sort-based mode findingO(iterations * E * log(d))
LCC (Local Clustering Coefficient)Parallel sorted-merge triangle countingO(E * sqrt(E))
SSSP (Single Source Shortest Paths)Dijkstra with binary min-heap on CSR + columnar weightsO((V + E) * log(V))

Prerequisites

  • Java 21 or later (required for jdk.incubator.vector SIMD support). The published benchmark numbers use Eclipse Temurin 25 with -XX:+UseCompactObjectHeaders (GraalVM is not used)
  • Maven 3.x
  • Python 3.10+ (for Mode 2 and Mode 3 multi-vendor comparisons; see pyproject.toml for per-vendor install extras)

Build

mvn package -DskipTests

The build produces a self-contained distribution in graphalytics-1.3.0-arcadedb-0.1-SNAPSHOT/.

Dataset

Use the built-in dataset manager to browse and download datasets from the LDBC data repository:

# See all available datasets (40+ Graphalytics + 9 LSQB scale factors)
python3 datasets.py available

# Download the standard Graphalytics benchmark dataset (633K vertices, 34M edges, ~155 MB)
python3 datasets.py download datagen-7_5-fb

# Download the LSQB social network dataset (SF1, ~3.9M vertices, ~17.9M edges)
python3 datasets.py download lsqb-sf1

# Show downloaded datasets with size and vertex/edge counts
python3 datasets.py

Datasets are downloaded into the datasets/ directory (git-ignored). After downloading datagen-7_5-fb:

datasets/
  datagen-7_5-fb/
    datagen-7_5-fb.v              # vertex file (one ID per line)
    datagen-7_5-fb.e              # edge file (src dst weight, space-separated)
    datagen-7_5-fb.properties     # graph metadata
    datagen-7_5-fb-BFS/           # validation data per algorithm
    datagen-7_5-fb-WCC/
    ...

Weekly unattended run

python3 weekend.py --jvm-flags "-XX:+UseCompactObjectHeaders"

Runs every suite and vendor in sequence (ArcadeDB embedded Java benchmarks, the Graphalytics and LSQB multi-vendor suites, optionally Mode 1 with --mode1-dist) and writes a Markdown/JSON report to weekly-results/<timestamp>/. Each vendor runs in its own process group with a total-time and a no-output limit; a hung vendor gets SIGTERM, then SIGKILL, and the run continues. Loaded databases are kept under ~/.cache/ldbc-graph-bench and reused, so only the first run (or --reset, a new image, a changed dataset) pays for the loads. See CLAUDE.md for the details and python3 -m unittest discover -s tests for the tests.

Benchmarks and results

Every number is a warm number (first call discarded, median of 3 timed runs), measured one system at a time on AC power and validated against the official LDBC reference outputs (or the official expected counts for LSQB). A value marked ✗ failed that check and is not ranked; N/A = no implementation; timeout = the 5-minute limit per operation. Footnote symbols (§ ‡ ¶ *) are explained on the detailed pages. Machine: MacBook Pro 16" (2026), Apple M5 Pro, 48 GB RAM; ArcadeDB on Eclipse Temurin 25 with -XX:+UseCompactObjectHeaders, 12 GB heap for every JVM system, Docker Desktop with 32 GB.

BenchmarkWhat it measuresDetailed page
Graphalytics, official framework (Mode 1)ArcadeDB only, official LDBC harness, one graph load per algorithmdocs/benchmark-graphalytics-official.md
Graphalytics, multi-vendor (Mode 2), datagen-7_5-fb6 algorithms, 10 systems, 633K vertices / 34M edgesdocs/benchmark-graphalytics-multivendor.md
Graphalytics, multi-vendor, graph500-225 algorithms, 10 systems, 2.4M vertices / 64M edgesdocs/benchmark-graphalytics-graph500-22.md
LSQB SF1 (Mode 3)9 subgraph pattern matching queries, 7+ systemsdocs/benchmark-lsqb.md

How ArcadeDB itself changes from release to release (official framework, Mode 2 and LSQB): ArcadeDB-release-progress.md.

Graphalytics, datagen-7_5-fb (633,432 vertices, 34,185,747 edges)

Seconds, last column peak memory in GiB. Bold marks the fastest valid result per column. Systems, versions, how we measure, per-system notes and how to run each vendor: detailed page.

SystemLoadPageRankWCCBFSLCCSSSPCDLPPeak memory (GiB)
ArcadeDB embedded71.90.0850.0030.0202.050.750.965.6
ArcadeDB Docker43.90.160.020.062.611.391.4512.3
Neo4j6576.98§0.1110.480‡15.4N/AN/A13.3
Kuzu28.81.160.4340.328N/AN/AN/A0.87
LadybugDB5.16N/AN/A7.89N/AN/AN/A1.0
DuckPGQ0.851.51✗2.00timeout¶49.3N/AN/A8.3
Memgraph4375.491893.85N/A75.9timeout25.1
ArangoDB *72693.040.938.4N/A173254✗23.2
FalkorDB1162.67✗2.500.057N/AN/A10.6✗8.4
HugeGraph34.72.410.2930.195110N/A22.3✗3.2

Graphalytics, graph500-22 (2,396,657 vertices, 64,155,735 edges)

Seconds; no SSSP (not defined for this dataset). ArcadeDB is 26.11.1-SNAPSHOT; bold marks the fastest valid result per column. Per-system notes, the harness problems found and how to reproduce: detailed page.

SystemLoadPageRankWCCBFSLCCCDLPPeak memory (GiB)
ArcadeDB embedded114.50.2680.0130.07646.91.781.6**
ArcadeDB Docker101†0.590.090.2460.14.3313.5
Neo4j2017†12.50.201.05timeoutN/A13.3
Kuzu53.8†3.131.260.94N/AN/A5.9
LadybugDB10.1†N/AN/A16.7N/AN/A2.0
DuckPGQ0.57†9.27✗4.94timeout‡timeoutN/A15.7
Memgraph932§31.0OOM12.5N/AOOM25.0
ArangoDB 3.11.141700†261.9106.4OOM¶N/Atimeout30.6
FalkorDB348.57.97✗7.300.151N/A39.8✗15.9
HugeGraph34.2†6.030.670.42timeout44.7✗8.8

OOM = out of memory (Docker Desktop's 32 GB). † remembered original load, ‡ DuckPGQ BFS does not finish (ended after 20 min), § Memgraph load includes the reverse-edge step, ¶ ArangoDB container ran out of memory, ** live heap after GC (process peak not measured): see the detailed page.

LSQB SF1 (3,947,829 vertices, 17,882,623 edges)

Seconds; ArcadeDB embedded is shown with the Graph Analytical View (OLAP) and without it (OLTP); bold marks the fastest result per query. All counts match the official expected output. Queries, run commands and analysis: detailed page.

SystemLoadQ1Q2Q3Q4Q5Q6Q7Q8Q9Peak GiB
ArcadeDB OLAP119.50.080.130.050.040.170.070.040.110.304.3
ArcadeDB OLTP159.22.454.403.171.0611.3911.001.076.910.8712.0
ArcadeDB Docker99.40.170.170.060.050.250.130.060.140.3312.9
DuckDB0.460.110.010.040.060.041.840.070.076.030.9
Kuzu2.444.610.152.30N/AN/A1.38N/AN/A6.395.1
LadybugDB2.950.120.1010.440.170.180.660.410.320.078.8
Neo4j252.04.991.6310.766.285.6628.498.0412.39254.2913.3
PostgreSQL15.19.170.651.606.352.1715.5011.163.3950.620.5
Memgraph222.258.06timeouttimeout4.303.54121.154.923.02timeout3.1
FalkorDB659.336.4353.784.593.533.9340.9347.964.90225.222.4

More

License

Apache License, Version 2.0

algorithms
benchmark
graph
graph-database
graphdb
performance-testing