All Projects → snowplow → snowplow-s3-loader

snowplow / snowplow-s3-loader

Licence: other
Mirrors a Kinesis stream to Amazon S3 using the KCL

Programming Languages

scala
5932 projects

Snowplow S3 Loader

Build Status Release License

Overview

The Snowplow S3 Loader consumes records from an Amazon Kinesis stream and writes them to S3.

There are 2 file formats supported:

  • LZO
  • GZip

LZO

The records are treated as raw byte arrays. Elephant Bird's BinaryBlockWriter class is used to serialize them as a Protocol Buffers array (so it is clear where one record ends and the next begins) before compressing them.

The compression process generates both compressed .lzo files and small .lzo.index files (splittable LZO). Each index file contain the byte offsets of the LZO blocks in the corresponding compressed file, meaning that the blocks can be processed in parallel.

GZip

The records are treated as byte arrays containing UTF-8 encoded strings (whether CSV, JSON or TSV). New lines are used to separate records written to a file. This format can be used with the Snowplow Kinesis Enriched stream, among other streams.

Quickstart

Docker

We publish three flavours of the docker image:

  • Pull the :2.2.1 tag if you only need GZip output format
  • Pull the :2.2.1-lzo tag if you also need LZO output format
  • Pull the :2.2.1-distroless tag for an lightweight alternative to :2.2.1
docker run snowplow/snowplow-s3-loader:2.2.1 --help
docker run snowplow/snowplow-s3-loader:2.2.1-lzo --help
docker run snowplow/snowplow-s3-loader:2.2.1-distroless --help

Download jar

curl -Lo snowplow-s3-loader.jar https://github.com/snowplow/snowplow-s3-loader/releases/download/2.2.1/snowplow-s3-loader-2.2.1.jar
java -jar snowplow-s3-loader.jar --help

Build it yourself

Assuming git and SBT installed:

$ git clone https://github.com/snowplow/snowplow-s3-loader.git
$ cd snowplow-s3-loader
$ sbt assembly

Prerequisites

You must have lzop and lzop-dev installed. In Ubuntu, install them like this:

$ sudo apt-get install lzop liblzo2-dev

Command Line Interface

The Snowplow S3 Loader has the following command-line interface:

snowplow-s3-loader: Version 2.2.1

Usage: snowplow-s3-loader [options]

--config <filename>

Running

Create your own config file:

$ cp config/config.hocon.sample my.conf

You will need to edit all fields in the config. Consult the configuration reference of the setup guide on how to fill in the fields.

Next, start the sink, making sure to specify your new config file:

$ java -jar snowplow-s3-loader-2.2.1.jar --config my.conf

Find out more

Technical Docs Setup Guide Roadmap Contributing
i1 i2 i3 i4
Technical Docs Setup Guide Roadmap Contributing

Copyright and license

Snowplow S3 Loader is copyright 2014-2022 Snowplow Analytics Ltd.

Licensed under the Apache License, Version 2.0 (the "License"); you may not use this software except in compliance with the License.

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

Note that the project description data, including the texts, logos, images, and/or trademarks, for each open source project belongs to its rightful owner. If you wish to add or remove any projects, please contact us at [email protected].