Big Data SQL & Processing Engines

20 repos

Distributed SQL engines and data processing frameworks built on JVM languages, designed for large-scale analytics and data warehousing. These systems emphasize query optimization, data lake integration, and polyglot support—enabling SQL interfaces over distributed compute clusters. The cluster includes foundational engines like Apache Spark alongside specialized tools for metadata management, query federation, and interactive analytics, with particular strength in Scala and Java implementations.

Java · 8
Scala · 8
TypeScript · 2
C++ · 1
Rust · 1
scala ·111,095
java ·110,961
big-data ·79,432
sql ·76,217
python ·70,308
spark ·58,746
jdbc ·49,756
kafka ·44,309
r ·43,981
streaming ·33,675

neo4j/neo4j

Graphs for Everyone

Java

17,190

70,747 commits

clyphy/Human-AI

No description

TypeScript

0

1 commits

apache/flink

Apache Flink

Java

26,327

30,378 commits