predictionioPredictionIO, a machine learning server for developers and ML engineers.
Stars: ✭ 12,510 (-0.1%)
RichdemHigh-performance Terrain and Hydrology Analysis
Stars: ✭ 127 (-98.99%)
Eel SdkBig Data Toolkit for the JVM
Stars: ✭ 140 (-98.88%)
Spark With PythonFundamentals of Spark with Python (using PySpark), code examples
Stars: ✭ 150 (-98.8%)
Sparkling GraphSparklingGraph provides easy to use set of features that will give you ability to proces large scala graphs using Spark and GraphX.
Stars: ✭ 139 (-98.89%)
Hdfs ShellHDFS Shell is a HDFS manipulation tool to work with functions integrated in Hadoop DFS
Stars: ✭ 117 (-99.07%)
CmakCMAK is a tool for managing Apache Kafka clusters
Stars: ✭ 10,544 (-15.8%)
AsakusafwAsakusa Framework
Stars: ✭ 114 (-99.09%)
Pythondatarepo for code published on pythondata.com
Stars: ✭ 113 (-99.1%)
GeniA Clojure dataframe library that runs on Spark
Stars: ✭ 152 (-98.79%)
100daysofmlcodeMy journey to learn and grow in the domain of Machine Learning and Artificial Intelligence by performing the #100DaysofMLCode Challenge.
Stars: ✭ 146 (-98.83%)
GenieDistributed Big Data Orchestration Service
Stars: ✭ 1,544 (-87.67%)
Spark R Notebooks R on Apache Spark (SparkR) tutorials for Big Data analysis and Machine Learning as IPython / Jupyter notebooks
Stars: ✭ 109 (-99.13%)
Griffon VmGriffon Data Science Virtual Machine
Stars: ✭ 128 (-98.98%)
Belajarpython.comOpen Source Indonesian Python Programming Tutorial Site
Stars: ✭ 141 (-98.87%)
Mobydq🐳 Tool to automate data quality checks on data pipelines
Stars: ✭ 123 (-99.02%)
FiliEasily make RESTful web services for time series reporting with Big Data analytics engines like Druid and SQL Databases.
Stars: ✭ 151 (-98.79%)
Report自动化配置报表平台。演示地址http://58.87.112.247/report 账号 visitor密码123456
Stars: ✭ 123 (-99.02%)
SigmfThe Signal Metadata Format Specification
Stars: ✭ 120 (-99.04%)
PrestoThe official home of the Presto distributed SQL query engine for big data
Stars: ✭ 12,957 (+3.47%)
DrillApache Drill is a distributed MPP query layer for self describing data
Stars: ✭ 1,619 (-87.07%)
PoseidonA search engine which can hold 100 trillion lines of log data.
Stars: ✭ 1,793 (-85.68%)
Amazon S3 Find And ForgetAmazon S3 Find and Forget is a solution to handle data erasure requests from data lakes stored on Amazon S3, for example, pursuant to the European General Data Protection Regulation (GDPR)
Stars: ✭ 115 (-99.08%)
ParquetviewerSimple windows desktop application for viewing & querying Apache Parquet files
Stars: ✭ 145 (-98.84%)
Just Dashboard📊 📋 Dashboards using YAML or JSON files
Stars: ✭ 1,511 (-87.93%)
AcceleratorThe Accelerator is a tool for fast and reproducible processing of large amounts of data.
Stars: ✭ 137 (-98.91%)
AmbariMirror of Apache Ambari
Stars: ✭ 1,576 (-87.41%)
KeyviKeyvi - the key value index. It is an in-memory FST-based data structure highly optimized for size and lookup performance.
Stars: ✭ 161 (-98.71%)
BigdataclassTwo-day workshop that covers how to use R to interact databases and Spark
Stars: ✭ 110 (-99.12%)
HydrographA visual ETL development and debugging tool for big data
Stars: ✭ 144 (-98.85%)
Tennis Crystal BallUltimate Tennis Statistics and Tennis Crystal Ball - Tennis Big Data Analysis and Prediction
Stars: ✭ 107 (-99.15%)
HamaMirror of Apache Hama
Stars: ✭ 129 (-98.97%)
MahaA framework for rapid reporting API development; with out of the box support for high cardinality dimension lookups with druid.
Stars: ✭ 101 (-99.19%)
VizukaExplore high-dimensional datasets and how your algo handles specific regions.
Stars: ✭ 100 (-99.2%)
Spark.jlJulia binding for Apache Spark
Stars: ✭ 153 (-98.78%)
MetamodelMirror of Apache Metamodel
Stars: ✭ 143 (-98.86%)
GafferA large-scale entity and relation database supporting aggregation of properties
Stars: ✭ 1,642 (-86.89%)
Graph samplingGraph Sampling is a python package containing various approaches which samples the original graph according to different sample sizes.
Stars: ✭ 99 (-99.21%)
KuduMirror of Apache Kudu
Stars: ✭ 1,360 (-89.14%)