TajoMirror of Apache Tajo
Stars: ✭ 128 (-34.69%)
MetamodelMirror of Apache Metamodel
Stars: ✭ 143 (-27.04%)
Pythondatarepo for code published on pythondata.com
Stars: ✭ 113 (-42.35%)
Spark With PythonFundamentals of Spark with Python (using PySpark), code examples
Stars: ✭ 150 (-23.47%)
RichdemHigh-performance Terrain and Hydrology Analysis
Stars: ✭ 127 (-35.2%)
GeopysparkGeoTrellis for PySpark
Stars: ✭ 167 (-14.8%)
CmakCMAK is a tool for managing Apache Kafka clusters
Stars: ✭ 10,544 (+5279.59%)
Eel SdkBig Data Toolkit for the JVM
Stars: ✭ 140 (-28.57%)
Spark R Notebooks R on Apache Spark (SparkR) tutorials for Big Data analysis and Machine Learning as IPython / Jupyter notebooks
Stars: ✭ 109 (-44.39%)
DatasciencevmTools and Docs on the Azure Data Science Virtual Machine (http://aka.ms/dsvm)
Stars: ✭ 153 (-21.94%)
GafferA large-scale entity and relation database supporting aggregation of properties
Stars: ✭ 1,642 (+737.76%)
KeyviKeyvi - a key value index that powers Cliqz search engine. It is an in-memory FST-based data structure highly optimized for size and lookup performance.
Stars: ✭ 171 (-12.76%)
FeastFeature Store for Machine Learning
Stars: ✭ 2,576 (+1214.29%)
100daysofmlcodeMy journey to learn and grow in the domain of Machine Learning and Artificial Intelligence by performing the #100DaysofMLCode Challenge.
Stars: ✭ 146 (-25.51%)
Presto Go ClientA Presto client for the Go programming language.
Stars: ✭ 183 (-6.63%)
Hdfs ShellHDFS Shell is a HDFS manipulation tool to work with functions integrated in Hadoop DFS
Stars: ✭ 117 (-40.31%)
AsakusafwAsakusa Framework
Stars: ✭ 114 (-41.84%)
FluoApache Fluo
Stars: ✭ 159 (-18.88%)
GenieDistributed Big Data Orchestration Service
Stars: ✭ 1,544 (+687.76%)
Sparkling GraphSparklingGraph provides easy to use set of features that will give you ability to proces large scala graphs using Spark and GraphX.
Stars: ✭ 139 (-29.08%)
AcceleratorThe Accelerator is a tool for fast and reproducible processing of large amounts of data.
Stars: ✭ 137 (-30.1%)
Spark.jlJulia binding for Apache Spark
Stars: ✭ 153 (-21.94%)
DvidDistributed, Versioned, Image-oriented Dataservice
Stars: ✭ 174 (-11.22%)
HamaMirror of Apache Hama
Stars: ✭ 129 (-34.18%)
FiliEasily make RESTful web services for time series reporting with Big Data analytics engines like Druid and SQL Databases.
Stars: ✭ 151 (-22.96%)
GunAn open source cybersecurity protocol for syncing decentralized graph data.
Stars: ✭ 15,172 (+7640.82%)
AzuredatalakeSamples and Docs for Azure Data Lake Store and Analytics
Stars: ✭ 128 (-34.69%)
ParquetviewerSimple windows desktop application for viewing & querying Apache Parquet files
Stars: ✭ 145 (-26.02%)
Griffon VmGriffon Data Science Virtual Machine
Stars: ✭ 128 (-34.69%)
Attic PredictionioPredictionIO, a machine learning server for developers and ML engineers.
Stars: ✭ 12,522 (+6288.78%)
Mobydq🐳 Tool to automate data quality checks on data pipelines
Stars: ✭ 123 (-37.24%)
HydrographA visual ETL development and debugging tool for big data
Stars: ✭ 144 (-26.53%)
Report自动化配置报表平台。演示地址http://58.87.112.247/report 账号 visitor密码123456
Stars: ✭ 123 (-37.24%)
Data Science Live BookAn open source book to learn data science, data analysis and machine learning, suitable for all ages!
Stars: ✭ 193 (-1.53%)
SigmfThe Signal Metadata Format Specification
Stars: ✭ 120 (-38.78%)
DrillApache Drill is a distributed MPP query layer for self describing data
Stars: ✭ 1,619 (+726.02%)
KeyviKeyvi - the key value index. It is an in-memory FST-based data structure highly optimized for size and lookup performance.
Stars: ✭ 161 (-17.86%)
Amazon S3 Find And ForgetAmazon S3 Find and Forget is a solution to handle data erasure requests from data lakes stored on Amazon S3, for example, pursuant to the European General Data Protection Regulation (GDPR)
Stars: ✭ 115 (-41.33%)
Belajarpython.comOpen Source Indonesian Python Programming Tutorial Site
Stars: ✭ 141 (-28.06%)
Just Dashboard📊 📋 Dashboards using YAML or JSON files
Stars: ✭ 1,511 (+670.92%)
FlumeMirror of Apache Flume
Stars: ✭ 2,200 (+1022.45%)
AmbariMirror of Apache Ambari
Stars: ✭ 1,576 (+704.08%)
BigdataclassTwo-day workshop that covers how to use R to interact databases and Spark
Stars: ✭ 110 (-43.88%)
PrestoThe official home of the Presto distributed SQL query engine for big data
Stars: ✭ 12,957 (+6510.71%)
PoseidonA search engine which can hold 100 trillion lines of log data.
Stars: ✭ 1,793 (+814.8%)
Couchdb DockerSemi-official Apache CouchDB Docker images
Stars: ✭ 194 (-1.02%)
Bigdata PlaygroundA complete example of a big data application using : Kubernetes (kops/aws), Apache Spark SQL/Streaming/MLib, Apache Flink, Scala, Python, Apache Kafka, Apache Hbase, Apache Parquet, Apache Avro, Apache Storm, Twitter Api, MongoDB, NodeJS, Angular, GraphQL
Stars: ✭ 177 (-9.69%)
GeniA Clojure dataframe library that runs on Spark
Stars: ✭ 152 (-22.45%)