Scala data processing and streaming

49 repos across 3 sub-areas

Distributed computing frameworks and data pipeline tools built primarily in Scala and Java. The cluster centers on Apache Spark for large-scale data processing, Akka for actor-based concurrency and reactive systems, and Akka HTTP for building reactive web services. These projects enable building scalable data workflows, streaming applications, and real-time processing systems, with supporting libraries like Alpakka for reactive integrations and Linkis for data engine orchestration.