All Categories → Data Processing → big-data

Top 369 big-data open source projects

ClickHouse® is a free analytics DBMS for big data

✭ 21,089

C++python assembly shell CMake javascript hacktoberfest sql analytics big-data clickhouse distributed-database olap dbms mpp

Koalas

Koalas: pandas API on Apache Spark

✭ 3,044

python shell data-science spark pandas big-data dataframe pydata mlflow

Vue Virtual Scroll List

⚡️A vue component support big amount data list with high render performance and efficient.

✭ 3,201

javascript big-data infinite-scroll virtual-list

Cboard

An easy to use, self-service open BI reporting and BI dashboard platform.

✭ 2,795

javascript HTML java CSS PLSQL PHP dashboard data-visualization big-data business-intelligence olap superset metabase cboard

Data Accelerator

Data Accelerator for Apache Spark simplifies onboarding to Streaming of Big Data. It offers a rich, easy to use experience to help with creation, editing and management of Spark jobs on Azure HDInsights or Databricks while enabling the full power of the Spark engine.

✭ 247

react nodejs docker iot kafka azure spark streaming big-data apache-spark kafka-streams spark-streaming streaming-data

Hyperspace

An open source indexing subsystem that brings index-based query acceleration to Apache Spark™ and big data workloads.

✭ 246

scala spark analytics big-data databases indexing

Aws Etl Orchestrator

A serverless architecture for orchestrating ETL jobs in arbitrarily-complex workflows using AWS Step Functions and AWS Lambda.

✭ 245

python aws serverless lambda big-data state-machine bigdata etl orchestration extract amazon-web-services transform

Trafodion

Apache Trafodion

✭ 242

cplusplus big-data

Kafka Ui

Open-Source Web GUI for Apache Kafka Management

✭ 230

typescript ui open-source gui kafka opensource big-data streams apache-kafka kafka-streams kafka-connect kafka-client kafka-producer pub-sub

Selinon

An advanced distributed task flow management on top of Celery

✭ 237

python kubernetes big-data openshift task distributed-computing celery

Eland

Python Client and Toolkit for DataFrames, Big Data, Machine Learning and ETL in Elasticsearch

✭ 235

python machine-learning elasticsearch data-analysis pandas big-data scikit-learn etl dataframe lightgbm

Books

整理一些书籍 ,包含 C&C++ 、git 、Java、Keras 、Linux 、NLP 、Python 、Scala 、TensorFlow 、大数据、推荐系统、数据库、数据挖掘、机器学习、深度学习、算法等。

✭ 222

python java c scala cpp tensorflow database nlp git keras big-data ml

Lite Virtual List

Virtual list component library supporting waterfall flow based on vue

✭ 223

vue big-data dom

Nakedtensor

Bare bone examples of machine learning in TensorFlow

✭ 2,443

python tensorflow big-data simple distributed-computing tensorflow-tutorials tensorflow-examples linear-regression tensorflow-exercises

Usql

U-SQL Examples and Issue Tracking

✭ 221

azure big-data

Gimel

Big Data Processing Framework - Unified Data API or SQL on Any Storage

✭ 216

python scala elasticsearch kafka spark big-data cassandra jdbc hbase pyspark paypal spark-streaming

Awkward 0.x

Manipulate arrays of complex data structures as easily as Numpy.

✭ 216

python python3 numpy big-data analysis parquet root arrow

Sparkrdma

RDMA accelerated, high-performance, scalable and efficient ShuffleManager plugin for Apache Spark

✭ 215

java scala spark big-data hadoop bigdata apache-spark

Helicalinsight

Helical Insight software is world’s first Open Source Business Intelligence framework which helps you to make sense out of your data and make well informed decisions.

✭ 214

java mysql mongodb postgresql dashboard data-visualization data-analysis big-data nosql neo4j reporting graph-database hive druid business-intelligence rdbms oracle-database

Calcite

Apache Calcite

✭ 2,816

java kotlin HTML SCSS FreeMarker shell sql big-data geospatial hadoop calcite

Attic Predictionio Sdk Python

PredictionIO Python SDK

✭ 196

python scala big-data

Couchdb Docker

Semi-official Apache CouchDB Docker images

✭ 194

javascript erlang cplusplus database http dockerfile cloud big-data apache couchdb

Data Science Live Book

An open source book to learn data science, data analysis and machine learning, suitable for all ages!

✭ 193

machine-learning tex data-science visualization learning statistics analytics data-analysis big-data predictive-modeling

Attic Predictionio Sdk Ruby

PredictionIO Ruby SDK

✭ 192

ruby scala big-data

Gun

An open source cybersecurity protocol for syncing decentralized graph data.

Presto Go Client

A Presto client for the Go programming language.

✭ 183

go golang sql big-data presto

Flume

Mirror of Apache Flume

✭ 2,200

java Rich Text Format shell powershell python Thrift big-data flume

Bigdata Playground

A complete example of a big data application using : Kubernetes (kops/aws), Apache Spark SQL/Streaming/MLib, Apache Flink, Scala, Python, Apache Kafka, Apache Hbase, Apache Parquet, Apache Avro, Apache Storm, Twitter Api, MongoDB, NodeJS, Angular, GraphQL

✭ 177

python typescript scala nodejs machine-learning docker angular graphql mongodb kafka big-data hadoop apache-spark twitter-api hbase avro parquet spark-streaming

Dvid

Distributed, Versioned, Image-oriented Dataservice

✭ 174

go big-data key-value neuroscience

Keyvi

Keyvi - a key value index that powers Cliqz search engine. It is an in-memory FST-based data structure highly optimized for size and lookup performance.

✭ 171

python cpp search data-structures big-data

Attic Predictionio

PredictionIO, a machine learning server for developers and ML engineers.

✭ 12,522

scala shell python HTML Dockerfile java Smarty big-data predictionio

Geopyspark

GeoTrellis for PySpark

✭ 167

python spark big-data geospatial

Keyvi

Keyvi - the key value index. It is an in-memory FST-based data structure highly optimized for size and lookup performance.