20 repos
Distributed SQL engines and data processing frameworks built on JVM languages, designed for large-scale analytics and data warehousing. These systems emphasize query optimization, data lake integration, and polyglot support—enabling SQL interfaces over distributed compute clusters. The cluster includes foundational engines like Apache Spark alongside specialized tools for metadata management, query federation, and interactive analytics, with particular strength in Scala and Java implementations.