{"id":111364,"date":"2025-02-26T14:39:44","date_gmt":"2025-02-26T14:39:44","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/snowplow-full-setup-with-google-analytics-tracking\/"},"modified":"2025-02-26T14:39:44","modified_gmt":"2025-02-26T14:39:44","slug":"snowplow-full-setup-with-google-analytics-tracking","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=111364","title":{"rendered":"Snowplow: Full Setup With Google Analytics Tracking"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p>A <a href=\"https:\/\/www.simoahava.com\/analytics\/automatically-fork-google-analytics-hits-snowplow\/\">recent guide<\/a> of mine introduced the <a href=\"https:\/\/snowplowanalytics.com\/blog\/2018\/01\/25\/snowplow-r99-carnac-with-google-analytics-support\/\">Google Analytics adapter<\/a> in <a href=\"https:\/\/snowplowanalytics.com\/\">Snowplow<\/a>. The idea was that you can duplicate the <a href=\"https:\/\/analytics.google.com\/\">Google Analytics<\/a> requests sent via <a href=\"https:\/\/tagmanager.google.com\/\">Google Tag Manager<\/a> and dispatch them to your Snowplow analytics pipeline, too. The pipeline then takes care of these duplicated requests, using the new adapter to automatically align the hits with their corresponding data tables, ready for data modeling and analysis.<\/p>\n<p>While testing the new adapter, I implemented a Snowplow pipeline from scratch for parsing data from my own website. This was the first time I\u2019d done the whole process from end-to-end myself, so I thought it might be prudent to document the process for the benefit of others who might want to take a jab at Snowplow but are intimidated by the moving parts.<\/p>\n<div style=\"aspect-ratio: 1919 \/ 467;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/snowplow-etl-emr.jpg\" title=\"Snowplow ETL in process\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"467\" width=\"1919\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/snowplow-etl-emr.jpg#ZgotmplZ\" alt=\"Snowplow ETL in process\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>And make no mistake. There are <strong>plenty<\/strong> of moving parts. Snowplow leverages a number of <a href=\"https:\/\/aws.amazon.com\/\">Amazon Web Services<\/a> components, in addition to a whole host of utilities of its own. It\u2019s not like setting up Google Analytics, which is, at its most basic, a fairly rudimentary plug-and-play affair.<\/p>\n<div style=\"aspect-ratio: 690 \/ 317;\" class=\"figure \">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/ga-schema.jpg\" title=\"Image source: https:\/\/goo.gl\/D2f3xi\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"317\" width=\"690\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/ga-schema.jpg#ZgotmplZ\" alt=\"Image source: https:\/\/goo.gl\/D2f3xi\"\/><\/p>\n<p>    <\/a><\/p>\n<p>    <span class=\"caption\">Image source: https:\/\/goo.gl\/D2f3xi<\/span><\/p>\n<\/div>\n<p>Take this article with a grain of salt. It definitely does <strong>not<\/strong> describe the most efficient or cost-effective way to do things, but it should help you get started with Snowplow, ending up with a data store full of usable data for modeling and analysis. As such, it\u2019s not necessarily a <strong>guide<\/strong> rather than a description of the steps I took, in good and bad.<\/p>\n<div style=\"aspect-ratio: 1433 \/ 297;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/sql-workbench-query.jpg\" title=\"SQL Workbench query against page view hits\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"297\" width=\"1433\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/sql-workbench-query.jpg#ZgotmplZ\" alt=\"SQL Workbench query against page view hits\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p><strong>WARNING:<\/strong> If you DO follow this guide step-by-step (which makes me very happy), please do note that there will be costs involved. For example, my current, very light-weight setup, is costing me a couple of dollars a day to maintain, with most costs incurred by running the collector on a virtual machine in Amazon\u2019s cloud. Just keep this in mind when working with AWS. It\u2019s unlikely to be totally free, even if you have the free tier for your account.<\/p>\n<p>                <span class=\"simmer\"><br \/>\n  <span class=\"close\">X<\/span><\/p>\n<p>\n    <span class=\"fa fa-md fa-bell\"\/><br \/>\n    <strong>The Simmer Newsletter<\/strong>\n  <\/p>\n<p>\n    Subscribe to the <a href=\"https:\/\/www.simoahava.com\/newsletter\/\">Simmer newsletter<\/a> to get the latest news and content from Simo Ahava into your email inbox!\n  <\/p>\n<p>  <\/span><\/p>\n<h2 id=\"to-start-with\">To start with<\/h2>\n<p>I ended up needing the following things to make the pipeline work:<\/p>\n<ul>\n<li>\n<p>The Linux\/Unix command line (handily accessible via the Terminal application of Mac OS X).<\/p>\n<\/li>\n<li>\n<p><a href=\"https:\/\/git-scm.com\/\">Git<\/a> client &#8211; not strictly necessary but it makes life easier to clone the Snowplow repo and work with it locally.<\/p>\n<\/li>\n<li>\n<p>A new <a href=\"https:\/\/aws.amazon.com\/\">Amazon Web Services<\/a> account with the introductory free tier (first 12 months).<\/p>\n<\/li>\n<li>\n<p>A credit card &#8211; even with the free tier the pipeline is not free.<\/p>\n<\/li>\n<li>\n<p>A <strong>domain name<\/strong> of my own (I used gtmtools.com) whose DNS records I can modify.<\/p>\n<\/li>\n<li>\n<p>A Google Analytics tag running through Google Tag Manager.<\/p>\n<\/li>\n<li>\n<p>A lot of time.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1453 \/ 470;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/command-line.jpg\" title=\"The OS X command line\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"470\" width=\"1453\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/command-line.jpg#ZgotmplZ\" alt=\"The OS X command line\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>The bullets concerning money and custom domain name might be a turn-off to some.<\/p>\n<p>You might be able to set up the pipeline without a domain name by using some combination of Amazon CloudFront and Route 53 with Amazon\u2019s own SSL certificates, but I didn\u2019t explore this option.<\/p>\n<p>And yes, this whole thing is going to cost money. As I wrote in the beginning, I didn\u2019t follow the most cost-effective path. But even if I did, it would still cost a dollar or something per day to keep this up and running. If this is enough of a red flag for you, then take a look at what <a href=\"https:\/\/snowplowanalytics.com\/products\/snowplow-insights\/\">managed solutions<\/a> Snowplow is offering. This article is for the engineers out there who want to try building the whole thing from scratch.<\/p>\n<h2 id=\"why-snowplow\">Why Snowplow?<\/h2>\n<p>Why do this exercise at all? Why even look towards Snowplow? The transition from the pre-built, top-down world of Google Analytics to the anarchy represented by Snowplow\u2019s agnostic approach to data processing can be daunting.<\/p>\n<p>Let me be frank: Snowplow is not for everyone. Even though the company itself offers managed solutions, making it as turnkey as it can get, it\u2019s still <strong>you<\/strong> building an <strong>analytics pipeline<\/strong> to suit <strong>your<\/strong> organization\u2019s needs. This involves asking very difficult questions, such as:<\/p>\n<ul>\n<li>\n<p>What is an \u201cevent\u201d?<\/p>\n<\/li>\n<li>\n<p>What constitutes a \u201csession\u201d?<\/p>\n<\/li>\n<li>\n<p>Who owns the data?<\/p>\n<\/li>\n<li>\n<p>What\u2019s the ROI of data analytics?<\/p>\n<\/li>\n<li>\n<p>How should conversions be attributed?<\/p>\n<\/li>\n<li>\n<p>How should I measure users across domains and devices?<\/p>\n<\/li>\n<\/ul>\n<p>If you\u2019ve never asked one of these (or similar) questions before, you might not want to look at Snowplow or any other custom-built setup yet. These are questions that inevitably surface at some point when using tools that give you very few configuration options.<\/p>\n<p>At this point I think I should add a disclaimer. This article is not <strong>Google Analytics versus Snowplow<\/strong>. There\u2019s no reason to bring one down to highlight the benefits of the other. Both GA and Snowplow have their place in the world of analytics, and having one is not predicated on the absence of the other.<\/p>\n<p>The whole idea behind the Google Analytics plugin, for example, is that you can duplicate tracking to both GA and to Snowplow. You might want to reserve GA tracking for marketing and advertising attribution, as Google\u2019s backend integrations are still unmatched by other platforms. You can then run Snowplow to collect this same data so that you\u2019ll have access to an unsampled, raw, relational database you can use to enrich and join with your other data sets.<\/p>\n<p><strong>Snowplow is NOT a Google Analytics killer<\/strong>. They\u2019re more like cousins fighting together for the honor of the same family line, but occasionally quarreling over the inheritance of a common, recently deceased relative.<\/p>\n<h2 id=\"what-we-are-going-to-build\">What we are going to build<\/h2>\n<p>Here\u2019s a diagram of what we\u2019re hopefully going to build in this article:<\/p>\n<div style=\"aspect-ratio: 1392 \/ 1029;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/snowplow-process.jpg\" title=\"Snowplow pipeline\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1029\" width=\"1392\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/snowplow-process.jpg#ZgotmplZ\" alt=\"Snowplow pipeline\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>I wonder why no one\u2019s hired me as a designer yet\u2026<\/p>\n<p>The process will cover the following steps.<\/p>\n<ol>\n<li>\n<p>The website will duplicate the payloads sent to Google Analytics, and send them to a collector written with <a href=\"https:\/\/clojure.org\/\">Clojure<\/a>.<\/p>\n<\/li>\n<li>\n<p>The collector runs as a web service on AWS <a href=\"https:\/\/docs.aws.amazon.com\/elasticbeanstalk\/latest\/dg\/Welcome.html\">Elastic Beanstalk<\/a>, to which traffic is routed and secured with SSL from my custom domain name using AWS <a href=\"https:\/\/docs.aws.amazon.com\/Route53\/latest\/DeveloperGuide\/Welcome.html\">Route 53<\/a>.<\/p>\n<\/li>\n<li>\n<p>The log data from the collector is stored in AWS <a href=\"https:\/\/docs.aws.amazon.com\/AmazonS3\/latest\/dev\/Welcome.html\">S3<\/a>.<\/p>\n<\/li>\n<li>\n<p>A utility is periodically executed on my local machine, which runs an ETL (<strong>e<\/strong>xtract, <strong>t<\/strong>ransform, <strong>l<\/strong>oad) process using AWS <a href=\"https:\/\/docs.aws.amazon.com\/emr\/latest\/APIReference\/Welcome.html\">EMR<\/a> to enrich and \u201cshred\u201d the data in S3.<\/p>\n<\/li>\n<li>\n<p>The same utility finally stores the processed data files into relational tables in AWS <a href=\"https:\/\/docs.aws.amazon.com\/redshift\/latest\/mgmt\/welcome.html\">Redshift<\/a>.<\/p>\n<\/li>\n<\/ol>\n<p>So the process ends with a relational database that has all your collected data populated periodically.<\/p>\n<h2 id=\"step-0-register-on-aws-and-setup-iam-roles\">Step 0: Register on AWS and setup IAM roles<\/h2>\n<h3 id=\"what-you-need-for-this-step\">What you need for this step<\/h3>\n<ol>\n<li>You\u2019ll just need a credit card to <a href=\"https:\/\/portal.aws.amazon.com\/billing\/signup#\/start\">register on AWS<\/a>. You\u2019ll get the benefits of a <a href=\"https:\/\/aws.amazon.com\/free\/\">free tier<\/a>, but you\u2019ll still need to enable billing.<\/li>\n<\/ol>\n<h3 id=\"register-on-amr\">Register on AMR<\/h3>\n<p>The very first thing to do is register on Amazon Web Services and setup an IAM (Identity and Access Management) User that will run the whole setup.<\/p>\n<p>So browse to <a href=\"https:\/\/aws.amazon.com\/,\">https:\/\/aws.amazon.com\/,<\/a> and select the option to create a free account.<\/p>\n<div style=\"aspect-ratio: 1466 \/ 388;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/create-free-account-aws.jpg\" title=\"Create free AWS account\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"388\" width=\"1466\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/create-free-account-aws.jpg#ZgotmplZ\" alt=\"Create free AWS account\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>The free account gives you access to the <a href=\"https:\/\/aws.amazon.com\/free\/\">free tier<\/a> of services, some of which will help keep costs down for this pipeline, too.<\/p>\n<h3 id=\"create-an-identity-and-account-management-iam-user\">Create an Identity and Account Management (IAM) user<\/h3>\n<p>Once you\u2019ve created the account, you can do the first important thing in setting up the pipeline: create an IAM User. We\u2019ll be following Snowplow\u2019s own excellent <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Setup-IAM-permissions-for-users-installing-Snowplow\">IAM setup guide<\/a> for these steps.<\/p>\n<p>In the <strong>Services<\/strong> menu, select IAM from the long list of items.<\/p>\n<div style=\"aspect-ratio: 1232 \/ 337;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/services-menu.jpg\" title=\"Services \/ IAM\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"337\" width=\"1232\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/services-menu.jpg#ZgotmplZ\" alt=\"Services \/ IAM\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>In the left-hand menu, select <strong>Groups<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Click the <strong>Create New Group<\/strong> in the view that opens.<\/p>\n<\/li>\n<li>\n<p>Name the group <code>snowplow-setup<\/code>.<\/p>\n<\/li>\n<li>\n<p>Skip the Attach Policy step for now by clicking the <strong>Next Step<\/strong> button.<\/p>\n<\/li>\n<li>\n<p>Click <strong>Create Group<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<p>Now in the left-hand menu, select <strong>Policies<\/strong>.<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-json\" data-lang=\"json\">{\n  <span style=\"color:#1e90ff;font-weight:bold\">\"Version\"<\/span>: <span style=\"color:#a50\">\"2012-10-17\"<\/span>,\n  <span style=\"color:#1e90ff;font-weight:bold\">\"Statement\"<\/span>: [\n    {\n      <span style=\"color:#1e90ff;font-weight:bold\">\"Effect\"<\/span>: <span style=\"color:#a50\">\"Allow\"<\/span>,\n      <span style=\"color:#1e90ff;font-weight:bold\">\"Action\"<\/span>: [\n        <span style=\"color:#a50\">\"acm:*\"<\/span>,\n        <span style=\"color:#a50\">\"autoscaling:*\"<\/span>,\n        <span style=\"color:#a50\">\"aws-marketplace:ViewSubscriptions\"<\/span>,\n        <span style=\"color:#a50\">\"aws-marketplace:Subscribe\"<\/span>,\n        <span style=\"color:#a50\">\"aws-marketplace:Unsubscribe\"<\/span>,\n        <span style=\"color:#a50\">\"cloudformation:*\"<\/span>,\n        <span style=\"color:#a50\">\"cloudfront:*\"<\/span>,\n        <span style=\"color:#a50\">\"cloudwatch:*\"<\/span>,\n        <span style=\"color:#a50\">\"ec2:*\"<\/span>,\n        <span style=\"color:#a50\">\"elasticbeanstalk:*\"<\/span>,\n        <span style=\"color:#a50\">\"elasticloadbalancing:*\"<\/span>,\n        <span style=\"color:#a50\">\"elasticmapreduce:*\"<\/span>,\n        <span style=\"color:#a50\">\"es:*\"<\/span>,\n        <span style=\"color:#a50\">\"iam:*\"<\/span>,\n        <span style=\"color:#a50\">\"rds:*\"<\/span>,\n        <span style=\"color:#a50\">\"redshift:*\"<\/span>,\n        <span style=\"color:#a50\">\"s3:*\"<\/span>,\n        <span style=\"color:#a50\">\"sns:*\"<\/span>\n      ],\n      <span style=\"color:#1e90ff;font-weight:bold\">\"Resource\"<\/span>: <span style=\"color:#a50\">\"*\"<\/span>\n    }\n  ]\n}<\/code><\/pre>\n<\/div>\n<ul>\n<li>\n<p>Next, click <strong>Review Policy<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Name the policy <code>snowplow-setup-policy-infrastructure<\/code>.<\/p>\n<\/li>\n<li>\n<p>Finally, click <strong>Create Policy<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<p>Now go back to <strong>Groups<\/strong> from the left-hand menu, and click the <strong>snowplow-setup<\/strong> group name to open its settings.<\/p>\n<div style=\"aspect-ratio: 911 \/ 325;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/iam-group.jpg\" title=\"IAM Group list\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"325\" width=\"911\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/iam-group.jpg#ZgotmplZ\" alt=\"IAM Group list\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Switch to the <strong>Permissions<\/strong> tab and click <strong>Attach Policy<\/strong>.<\/p>\n<\/li>\n<li>\n<p>From the list that opens, select <strong>snowplow-setup-policy-infrastructure<\/strong> and click <strong>Attach Policy<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 855 \/ 395;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/attach-policy.jpg\" title=\"Attach custom policy\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"395\" width=\"855\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/attach-policy.jpg#ZgotmplZ\" alt=\"Attach custom policy\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Now select <strong>Users<\/strong> from the left-hand menu, and click the <strong>Add user<\/strong> button.<\/p>\n<ul>\n<li>\n<p>Name the user <code>snowplow-setup<\/code>.<\/p>\n<\/li>\n<li>\n<p>Check the box next to <strong>Programmatic access<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Click <strong>Next: Permissions<\/strong>.<\/p>\n<\/li>\n<li>\n<p>With <strong>Add user to group<\/strong> selected, check the box next to <strong>snowplow-setup<\/strong>, and click the <strong>Next: Review<\/strong> button at the bottom of the page.<\/p>\n<\/li>\n<li>\n<p>Finally, click <strong>Create user<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<p>The following screen will show you a success message, and your user with an <strong>Access Key ID<\/strong> and <strong>Secret Access Key<\/strong> (click Show to see it) available. At this point, <strong>copy both of these somewhere safe<\/strong>. You will need them soon, and once you leave this screen you will not be able to see the secret key anymore. You can also download them as a CSV file by clicking the <strong>Download .csv<\/strong> button.<\/p>\n<p>You have now created a user with which you will set up your entire pipeline. Once everything is set up, you will create a new user with fewer permissions, who will take care of running and managing the pipeline.<\/p>\n<p>Congratulations, step 0 complete!<\/p>\n<h3 id=\"what-you-should-have-after-this-step\">What you should have after this step<\/h3>\n<ol>\n<li>\n<p>You should have successfully registered a new account on AWS.<\/p>\n<\/li>\n<li>\n<p>You should have a new IAM user named <code>snowplow-setup<\/code> with all the access privileges distributed by the custom policy you created.<\/p>\n<\/li>\n<\/ol>\n<h2 id=\"step-1-the-clojure-collector\">Step 1: The Clojure collector<\/h2>\n<h3 id=\"what-you-need-for-this-step-1\">What you need for this step<\/h3>\n<ol>\n<li>A custom domain name and access to its DNS configuration<\/li>\n<\/ol>\n<h3 id=\"getting-started\">Getting started<\/h3>\n<p>This is going to be one of the more difficult steps to do, since there\u2019s no generic guide for some of the things that need to be done with the collector endpoint. If you want, you can look into the <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Setting-up-the-Cloudfront-collector\">CloudFront Collector<\/a> instead, because that runs directly on top of S3 without needing a web service to collect the hits. However, it doesn\u2019t support the Google Analytics hit duplicator, which is why in this article we\u2019ll use the <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Setting-up-the-Clojure-collector\">Clojure collector<\/a>.<\/p>\n<p>The Clojure collector is basically a web endpoint to which you will log the requests from your site. The endpoint is a scalable web service running on Apache Tomcat, which is hosted on AWS\u2019 Elastic Beanstalk. The collector has been configured to automatically log the Tomcat access logs directly into AWS S3 storage, meaning whenever your site sends a request to the collector, it logs this request as an access log entry, ready for the ETL process that comes soon after.<\/p>\n<p>But let\u2019s not get ahead of ourselves.<\/p>\n<h3 id=\"setting-up-the-clojure-collector\">Setting up the Clojure collector<\/h3>\n<p>The first thing you\u2019ll need to do is download the Clojure collector WAR file and store it locally in a temporary location (such as your Downloads folder).<\/p>\n<p>You can download the binary by following <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Hosted-assets#21-clojure-collector-resources\">this<\/a> link. It should lead you to a file that looks something like <code>clojure-collector-1.X.X-standalone.war<\/code>.<\/p>\n<p>Once you\u2019ve downloaded it, you can set up the <strong>Elastic Beanstalk<\/strong> application.<\/p>\n<p>In AWS, open the <strong>Services<\/strong> menu again, and select <strong>Elastic Beanstalk<\/strong>.<\/p>\n<p>At this point, you\u2019ll also want to make sure that any AWS services you use in this pipeline are located in the same region. There are differences to what services are supported in each region. I built the pipeline in <strong>EU (Ireland)<\/strong>, and the Snowplow guide itself uses <strong>US West (Oregon)<\/strong>. Click the region menu and choose the region you want to use.<\/p>\n<div style=\"aspect-ratio: 879 \/ 405;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/region-selection.jpg\" title=\"Select AWS region\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"405\" width=\"879\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/region-selection.jpg#ZgotmplZ\" alt=\"Select AWS region\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Next, click the <strong>Create New Application<\/strong> link in the top-right corner of the page (just below the region selector).<\/p>\n<\/li>\n<li>\n<p>Give the application a descriptive name (e.g. <code>Snowplow Clojure Collector<\/code>) and description (e.g. <code>I love Kamaka HF-3<\/code>), and then click <strong>Next<\/strong>.<\/p>\n<\/li>\n<li>\n<p>In the New Environment view, select <strong>Create web server<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1270 \/ 481;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/beanstalk-new-web-server.jpg\" title=\"New Elastic Beanstalk web server\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"481\" width=\"1270\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/beanstalk-new-web-server.jpg#ZgotmplZ\" alt=\"New Elastic Beanstalk web server\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>In the Environment Type view, set the following:<\/li>\n<\/ul>\n<blockquote>\n<p><strong>Predefined configuration<\/strong>: Tomcat<br \/><strong>Environment type<\/strong>: Single instance<\/p>\n<\/blockquote>\n<ul>\n<li>\n<p>Then, click <strong>Next<\/strong>.<\/p>\n<\/li>\n<li>\n<p>In the Application Version view, select <strong>Upload your own<\/strong>, click <strong>Choose file<\/strong>, and find the WAR file you downloaded earlier in this chapter. Click <strong>Next<\/strong> to upload the file.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 838 \/ 315;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/upload-war-file.jpg\" title=\"Upload WAR file\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"315\" width=\"838\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/upload-war-file.jpg#ZgotmplZ\" alt=\"Upload WAR file\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>In the Environment Info view, you\u2019ll need to set the <strong>Environment name<\/strong>, which is then used to generate the <strong>Environment URL<\/strong> (which, in turn, will be the endpoint URL receiving the collector requests).<\/p>\n<\/li>\n<li>\n<p>Remember to click <strong>Check availability<\/strong> for the URL to make sure someone hasn\u2019t grabbed it yet. Click <strong>Next<\/strong> once you\u2019re done.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 804 \/ 329;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/environment-name-and-url.jpg\" title=\"Environment name and URL\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"329\" width=\"804\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/environment-name-and-url.jpg#ZgotmplZ\" alt=\"Environment name and URL\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>In Additional Resources, you can leave both options unchecked for now, and click <strong>Next<\/strong>.<\/p>\n<\/li>\n<li>\n<p>In Configuration Details, select <strong>m1.small<\/strong> as the instance type. You can leave all the other options to their default settings. Click <strong>Next<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 977 \/ 381;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/configuration-details.jpg\" title=\"Configuration details\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"381\" width=\"977\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/configuration-details.jpg#ZgotmplZ\" alt=\"Configuration details\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>No need to add any Environment Tags, so click <strong>Next<\/strong> again.<\/p>\n<\/li>\n<li>\n<p>In the Permissions view, by clicking <strong>Next<\/strong>, AWS assigns default roles to Instance profile and Service role, so that\u2019s fine.<\/p>\n<\/li>\n<li>\n<p>Finally, you can take a quick look at what you\u2019ve done in the Review view, before clicking the <strong>Launch<\/strong> button.<\/p>\n<\/li>\n<\/ul>\n<p>At this point, you\u2019ll see that AWS is firing up your environment, where the Clojure collector WAR file will start running the instant the environment has been created.<\/p>\n<div style=\"aspect-ratio: 1709 \/ 614;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-collector-starting-up.jpg\" title=\"Clojure collector starting up\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"614\" width=\"1709\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-collector-starting-up.jpg#ZgotmplZ\" alt=\"Clojure collector starting up\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Once the environment is up and running, you can copy the URL from the top of the view, paste it into the address bar of your browser, and add the path <code>\/i<\/code> to its end, so it ends up something like:<\/p>\n<p><code>http:\/\/simoahava-snowplow-collector.eu-west-1.elasticbeanstalk.com\/i<\/code><\/p>\n<p>If the collector is running correctly, you should see a pixel in the center of the screen. By clicking the right mouse button and choosing Inspect (in the Google Chrome browser), you should now find a cookie named <code>sp<\/code> in the Applications tab. If you do, it means the collector is working correctly.<\/p>\n<div style=\"aspect-ratio: 1335 \/ 891;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-test.jpg\" title=\"Testing the clojure collector\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"891\" width=\"1335\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-test.jpg#ZgotmplZ\" alt=\"Testing the clojure collector\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Congratulations! You\u2019ve set up the collector.<\/p>\n<p>However, we\u2019re not done here yet.<\/p>\n<h3 id=\"enable-logging-to-s3\">Enable logging to S3<\/h3>\n<p>Automatically logging the Tomcat access logs to S3 storage is absolutely crucial for this whole pipeline. The batch process looks for these logs when sorting out the data into queriable chunks. Access logs are your typical web server logs, detailing all the HTTP requests made to the endpoint, with request headers and payloads included.<\/p>\n<p>To enable logging, you\u2019ll need to edit the Elastic Beanstalk application you just created. So, once the endpoint is up and running, you can open it by clicking the application name while in the <strong>Elastic Beanstalk<\/strong> service front page.<\/p>\n<div style=\"aspect-ratio: 1158 \/ 658;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-collector-application.jpg\" title=\"Clojure collector application\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"658\" width=\"1158\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-collector-application.jpg#ZgotmplZ\" alt=\"Clojure collector application\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Next, select <strong>Configuration<\/strong> from the left-hand menu.<\/p>\n<\/li>\n<li>\n<p>Click the cogwheel in the box titled <strong>Software Configuration<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 838 \/ 451;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/software-configuration.jpg\" title=\"Software configuration\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"451\" width=\"838\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/software-configuration.jpg#ZgotmplZ\" alt=\"Software configuration\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>Under <strong>Log Options<\/strong>, check the box next to <strong>Enable log file rotation to Amazon S3. If checked, service logs are published to S3<\/strong>.<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 921 \/ 235;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/enable-log-rotation.jpg\" title=\"Enable log rotation\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"235\" width=\"921\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/enable-log-rotation.jpg#ZgotmplZ\" alt=\"Enable log rotation\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>This is a crucial step, because it will store all the access logs from your endpoint requests to S3, ready for ETL.<\/p>\n<p>Click <strong>Apply<\/strong> to apply the change.<\/p>\n<h3 id=\"set-up-the-load-balancer\">Set up the load balancer<\/h3>\n<p>Next, we need to configure the Elastic Beanstalk environment for SSL.<\/p>\n<p>Before you do anything else, you\u2019ll need to switch from a single instance to a load-balancing, auto-scaling environment in Elastic Beanstalk. This is necessary for securing the traffic between your domain name and the Clojure collector.<\/p>\n<p>In the AWS <strong>Services<\/strong> menu, select <strong>Elastic Beanstalk<\/strong>, and then click your application name in the view that opens.<\/p>\n<div style=\"aspect-ratio: 1158 \/ 658;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-collector-application.jpg\" title=\"Clojure collector application\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"658\" width=\"1158\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/clojure-collector-application.jpg#ZgotmplZ\" alt=\"Clojure collector application\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>In the next view, select <strong>Configuration<\/strong> in the left-hand menu.<\/p>\n<\/li>\n<li>\n<p>Click the cogwheel in the box titled <strong>Scaling<\/strong>.<\/p>\n<\/li>\n<li>\n<p>From the <strong>Environment type<\/strong> menu, select <strong>Load balancing, auto scaling<\/strong>, and then click the <strong>Apply<\/strong> button, and <strong>Save<\/strong> in the next view.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1404 \/ 380;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/change-environment-type.jpg\" title=\"Change environment type\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"380\" width=\"1404\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/change-environment-type.jpg#ZgotmplZ\" alt=\"Change environment type\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>You now have set up the load balancer.<\/p>\n<h3 id=\"route-traffic-from-your-custom-domain-name-to-the-load-balancer\">Route traffic from your custom domain name to the load balancer<\/h3>\n<p>Next, we\u2019ll get started on routing traffic from your custom domain name to this load balancer.<\/p>\n<ul>\n<li>\n<p>Open the <strong>Services<\/strong> menu in the AWS console, and select <strong>Route 53<\/strong> from the list.<\/p>\n<\/li>\n<li>\n<p>In the view that opens, click <strong>Create Hosted Zone<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Set the <strong>Domain Name<\/strong> to the domain whose DNS records you want to delegate to Amazon. I chose <strong>collector.gtmtools.com<\/strong>. Leave <strong>Type<\/strong> as <strong>Public Hosted Zone<\/strong>, and click <strong>Create<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1050 \/ 412;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/route-53-hosted-zone.jpg\" title=\"Configure hosted zone in Route 53\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"412\" width=\"1050\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/route-53-hosted-zone.jpg#ZgotmplZ\" alt=\"Configure hosted zone in Route 53\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>In the view that opens, you\u2019ll see the settings for your Hosted Zone. Make note of the four NS records AWS has assigned to your domain name. You\u2019ll need these in the next step.<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1271 \/ 297;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/hosted-zone-name-servers.jpg\" title=\"Hosted Zone name servers\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"297\" width=\"1271\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/hosted-zone-name-servers.jpg#ZgotmplZ\" alt=\"Hosted Zone name servers\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Next, you\u2019ll need to go wherever it is you manage your DNS records. I use <a href=\"https:\/\/www.godaddy.com\/\">GoDaddy<\/a>.<\/p>\n<\/li>\n<li>\n<p>You need to add the four NS addresses in the AWS Hosted Zone as <strong>NS<\/strong> records in the DNS settings of your domain name. This is what the modifications would look like in my GoDaddy control panel:<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1170 \/ 344;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/godaddy-ns-records.jpg\" title=\"GoDaddy NS records\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"344\" width=\"1170\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/godaddy-ns-records.jpg#ZgotmplZ\" alt=\"GoDaddy NS records\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>As you can see, there are four <strong>NS<\/strong> records with host <strong>collector<\/strong> (for collector.gtmtools.com), each pointing to one of the four corresponding NS addresses in the AWS Hosted Zone. I set the TTL to the shortest possible GoDaddy allows, which is 600 seconds. That means that within 10 minutes, the Hosted Zone name servers should respond to <strong>collector.gtmtools.com<\/strong>.<\/p>\n<p>You can test this with a service such as <a href=\"https:\/\/dig.whois.com.au\/dig\/,\">https:\/\/dig.whois.com.au\/dig\/,<\/a> or any similar service that lets you check DNS records. Once the DNS settings are updated, you can increase the TTL to something more sensible, such as 1 hour, or even 1 day.<\/p>\n<p>Now that you\u2019ve set up your custom domain name to point to your Route 53 Hosted Zone, there\u2019s just one step missing. You\u2019ll need to create an <strong>Alias<\/strong> record in the Hosted Zone, which points to your load balancer. That way when typing the URL <strong>collector.gtmtools.com<\/strong> into the browser address bar, the DNS record first directs it to your Hosted Zone, where a new A record shuffles the traffic to your load balanced Clojure collector endpoint. Phew!<\/p>\n<ul>\n<li>\n<p>So, in the Hosted Zone you\u2019ve created, click <strong>Create Record Set<\/strong>.<\/p>\n<\/li>\n<li>\n<p>In the overlay that opens, leave the <strong>Name<\/strong> empty, since you want to apply the name to the root domain of the NS (collector.gtmtools.com in my case). Keep <strong>A &#8211; IPv4 address<\/strong> as the Type, and select <strong>Yes<\/strong> for <strong>Alias<\/strong>.<\/p>\n<\/li>\n<li>\n<p>When you click the <strong>Target<\/strong> field, a drop-down list should appear, and your load balancer should be selectable under the <strong>ELB Classic load balancers<\/strong> heading. Select that, and then click <strong>Create<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 418 \/ 547;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/a-record-set.jpg\" title=\"A Record Set\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"547\" width=\"418\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/a-record-set.jpg#ZgotmplZ\" alt=\"A Record Set\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Now if you visit <code>http:\/\/collector.gtmtools.com\/i<\/code>, you should see the same pixel response as you got when visiting the Clojure collector endpoint directly. So your domain name routing works!<\/p>\n<p>But we\u2019re STILL not done here.<\/p>\n<h3 id=\"setting-up-https-for-the-collector\">Setting up HTTPS for the collector<\/h3>\n<p>To make sure the collector is secured with HTTPS, you will need to generate a (free) AWS SSL certificate for it, and apply it to the Load Balancer. Luckily this is easy to do now that we\u2019re working with Route 53.<\/p>\n<ul>\n<li>\n<p>The first thing to do is generate the SSL certificate. In <strong>Services<\/strong>, find and select the <strong>AWS Certificate Manager<\/strong>. Click <strong>Get started<\/strong> to, well, get started.<\/p>\n<\/li>\n<li>\n<p>Type the domain name you want to apply the certificate to in the relevant field. I wrote <strong>collector.gtmtools.com<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Click <strong>Next<\/strong> when ready.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1351 \/ 609;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/apply-certificate.jpg\" title=\"Apply AWS certificate\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"609\" width=\"1351\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/apply-certificate.jpg#ZgotmplZ\" alt=\"Apply AWS certificate\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>In the next step you need to choose a validation method. Since we\u2019ve delegated DNS of collector.gtmtools.com to Route 53, I chose <strong>DNS Validation<\/strong> without hesitation.<\/p>\n<\/li>\n<li>\n<p>Then click <strong>Review<\/strong>, and then <strong>Confirm and request<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Validation is done in the next view. With Route 53, this is really easy. Just click <strong>Create record in Route 53<\/strong>, and <strong>Create<\/strong> in the overlay that opens. Amazon takes care of validation for you!<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1781 \/ 552;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/validate-certificate.jpg\" title=\"Validate SSL certificate\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"552\" width=\"1781\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/validate-certificate.jpg#ZgotmplZ\" alt=\"Validate SSL certificate\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>After clicking Create, you should see a <strong>Success<\/strong> message, and you can click the <strong>Continue<\/strong> button in the bottom of the screen. It might take up to 30 minutes for the certificate to validate, so go grab a cup of coffee or something! We still have one more step left\u2026<\/p>\n<h3 id=\"switch-the-load-balancer-to-support-https\">Switch the load balancer to support HTTPS<\/h3>\n<p>You\u2019ll still need to switch your load balancer to listen for secure requests, too.<\/p>\n<ul>\n<li>\n<p>In <strong>Services<\/strong>, open <strong>Elastic Beanstalk<\/strong>, click your application name in the view that opens, and finally click <strong>Configuration<\/strong> in the left-hand menu. You should be able to do this stuff in your sleep by now!<\/p>\n<\/li>\n<li>\n<p>Next, scroll down to the box titled <strong>Load Balancing<\/strong> and click the cogwheel in it.<\/p>\n<\/li>\n<li>\n<p>In the view that opens, set <strong>Secure listener port<\/strong> to <strong>443<\/strong>, and select the SSL certificate you just generated from the menu next to <strong>SSL Certificate ID<\/strong>. Click <strong>Apply<\/strong> when ready.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1035 \/ 455;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/load-balancer-settings.jpg\" title=\"HTTPS for load balancer\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"455\" width=\"1035\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/load-balancer-settings.jpg#ZgotmplZ\" alt=\"HTTPS for load balancer\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h3 id=\"all-done\">All done!<\/h3>\n<p>At this point, you might also want to take a look at <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Configuring-the-Clojure-collector\">this Snowplow guide<\/a> for configuring the collector further (e.g. applying proper scaling settings to the load balancer).<\/p>\n<p>The process above might seem convoluted, but there\u2019s a certain practical logic to it all. And once you have the whole pipeline up and running, it will be easier to understand how things proceed from the S3 storage onwards.<\/p>\n<p>Setting up the custom domain name is a bit of chore, though. But if you use Route 53, most of the things are either automated for you or taken care of with the click of a button.<\/p>\n<h3 id=\"what-you-should-have-after-this-step-1\">What you should have after this step<\/h3>\n<ol>\n<li>\n<p>The Clojure collector application running in an Elastic Beanstalk environment.<\/p>\n<\/li>\n<li>\n<p>Your own custom domain name pointing at the router configured in Amazon Route 53.<\/p>\n<\/li>\n<li>\n<p>An SSL secured load balancer, to which Route 53 diverts traffic to your custom domain name.<\/p>\n<\/li>\n<li>\n<p>Automatic logging of the collector Tomcat logs to S3. The bucket name is something like <strong>elasticbeanstalk-region-id<\/strong>, and you can find it by clicking <strong>Services<\/strong> in the AWS navigation and choosing <strong>S3<\/strong>. The logs are pushed hourly.<\/p>\n<\/li>\n<\/ol>\n<h2 id=\"step-2-the-tracker\">Step 2: The tracker<\/h2>\n<p>You\u2019ll need to configure a tracker to collect data to the S3 storage.<\/p>\n<p>This is really easy, since you\u2019re of course using Google Analytics, and tracking data to it using Google Tag Manager tags.<\/p>\n<p>Navigate to my <a href=\"https:\/\/www.simoahava.com\/analytics\/automatically-fork-google-analytics-hits-snowplow\/\">recent guide<\/a> on setting up the duplicator, do what it says, and you\u2019ll be set. Remember to change the <code>endpoint<\/code> variable in the Custom JavaScript variable to the domain name you set up in the previous chapter (<code>https:\/\/collector.gtmtools.com\/<\/code> in my case).<\/p>\n<h2 id=\"step-25-test-the-tracker-and-collector\">Step 2.5: Test the tracker and collector<\/h2>\n<p>Once you\u2019ve installed the GA duplicator, you can test to see if the logs are being stored in S3 properly.<\/p>\n<p>If the duplicator is doing its job correctly, you can open the <strong>Network<\/strong> tab in your browser\u2019s developer tools, and look for requests to your collector endpoint. You should see POST requests for each GA call, with the payload of the request in the POST body:<\/p>\n<div style=\"aspect-ratio: 1476 \/ 383;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/network-developer-tools.jpg\" title=\"Network tab in developer tools\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"383\" width=\"1476\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/network-developer-tools.jpg#ZgotmplZ\" alt=\"Network tab in developer tools\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>If you don\u2019t see these requests, it means you\u2019ve misconfigured the duplicator somehow, and you should re-read the <a href=\"https:\/\/www.simoahava.com\/analytics\/automatically-fork-google-analytics-hits-snowplow\/\">instructions<\/a>.<\/p>\n<p>If you see the requests but there\u2019s an error in the HTTP response, you\u2019ll need to check the process outlined in the previous two chapters again.<\/p>\n<p>At something like 10 minutes past each hour, the Clojure collector running in Elastic Beanstalk will dump all the Tomcat access logs to S3. You should check that they are being stored, because the whole batch process hinges on the presence of these logs.<\/p>\n<p>In S3, there will be a bucket prefixed with <strong>elasticbeanstalk-region-id<\/strong>. Within that bucket, browse to folder <strong>resources \/ environments \/ logs \/ publish \/ (some ID) \/ (some ID)<\/strong>. In other words, within the <strong>publish<\/strong> folder will be a folder named something like <strong>e-ab12cd23ef<\/strong>, and within that will be a folder named something like <strong>i-1234567890<\/strong>. Within that final folder will be all your logs in gzip format.<\/p>\n<p>Look for the ones named <strong>_var_log_tomcat8_rotated_localhost_access_log.txt123456789.gz<\/strong>, as these are the logs that the ETL process will use to build the data tables.<\/p>\n<p>If you open one of those logs, you should find a bunch of GET and POST entries. Look for POST entries where the endpoint is <code>\/com.google.analytics\/v1<\/code>, and the HTTP status code is <code>200<\/code>. If you see these, it means that the Clojure collector is almost certainly doing its job. The entry will contain a bunch of interesting information, such as the IP address of the visitor, the User-Agent string of the browser, and a <em>base64<\/em> encoded string with the payload. If you decode this string, you should see the full payload of your Google Analytics hit as a query string.<\/p>\n<div style=\"aspect-ratio: 827 \/ 205;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/post-request-from-logs.jpg\" title=\"POST request from Tomcat access logs\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"205\" width=\"827\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/post-request-from-logs.jpg#ZgotmplZ\" alt=\"POST request from Tomcat access logs\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h2 id=\"step-3-configure-the-etl-process\">Step 3: Configure the ETL process<\/h2>\n<h3 id=\"what-you-need-for-this-step-2\">What you need for this step<\/h3>\n<ol>\n<li>\n<p>A collector running and dumping the logs in an S3 bucket.<\/p>\n<\/li>\n<li>\n<p>Access Key ID and Secret Access Key for the IAM user you created in <a href=\"#step-0-register-on-aws-and-setup-iam-roles\">Step 0<\/a>.<\/p>\n<\/li>\n<\/ol>\n<h3 id=\"getting-started-1\">Getting started<\/h3>\n<p>This step <strong>should<\/strong> be pretty straightforward, at least more so than the previous one.<\/p>\n<p>The process does a number of things, and you\u2019ll want to check out <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Batch-pipeline-steps\">this page<\/a> for more info.<\/p>\n<p>But, in short, here are the main steps operated by the AWS Elastic MapReduce (EMR) service.<\/p>\n<ol>\n<li>\n<p>The Tomcat logs are cleaned up so that they can be parsed more easily. Irrelevant log entries are discarded.<\/p>\n<\/li>\n<li>\n<p>Custom enrichments are applied to the data, if you so wish. Enrichments include things like geolocation from IP addresses, or adding <a href=\"https:\/\/www.simoahava.com\/analytics\/send-weather-data-to-google-analytics-in-gtm-v2\/\">weather<\/a> information to the data set.<\/p>\n<\/li>\n<li>\n<p>The enriched data is then <em>shredded<\/em>, or split into more atomic data sets, each corresponding with a hit that validates against a given schema. For example, if you are collecting data with the Google Analytics setup outlined in my <a href=\"https:\/\/www.simoahava.com\/analytics\/automatically-fork-google-analytics-hits-snowplow\/\">guide<\/a>, these hits would be automatically shredded into data sets for Page Views, Events, and other Google Analytics tables, ready for transportation to a relational database.<\/p>\n<\/li>\n<li>\n<p>Finally, the data is copied from the shredded sets to a database created in Amazon Redshift.<\/p>\n<\/li>\n<\/ol>\n<p>We\u2019ll go over these steps in detail next. It\u2019s important to understand that each step of the ETL process leaves a trace in S3 buckets you\u2019ll build along the way. This means that even if you choose to apply the current process to your raw logs, you can rerun your entire log history, if you\u2019ve decided to keep the files, with new enrichments and shredding schemas later on. All the Tomcat logs are archived, too, so you\u2019ll always be able to start the entire process from scratch, using all your historical data, if you wish!<\/p>\n<p>The way we\u2019ll work in this process is run a Java application name <strong>EmrEtlRunner<\/strong> from your local machine. This application initiates and runs the ETL process using Amazon\u2019s Elastic MapReduce. At some point, you might want to upgrade your setup, and have <strong>EmrEtlRunner<\/strong> execute in an AWS EC2 instance (basically a virtual machine in Amazon\u2019s cloud). That way you could schedule it to run, say, every 60 minutes, and then forget about it.<\/p>\n<h3 id=\"download-the-necessary-files\">Download the necessary files<\/h3>\n<p>The ETL runner is a Unix application you can download from <a href=\"http:\/\/dl.bintray.com\/snowplow\/snowplow-generic\/\">this<\/a> link. To grab the latest version, look for the file that begins with <code>snowplow_emr_rXX<\/code>, where XX is the highest number you can find. At the time of writing, the latest binary is <code>snowplow_emr_r97_knossos.zip<\/code>.<\/p>\n<ul>\n<li>\n<p>Download this ZIP file, and copy the <code>snowplow-emr-etl-runner<\/code> Unix executable into a new folder on your hard drive. This folder will be your base of operations.<\/p>\n<\/li>\n<li>\n<p>At this point, you\u2019ll want to also clone the <a href=\"https:\/\/github.com\/snowplow\/snowplow\">Snowplow Github repo<\/a> in that folder, because it has all the config file templates and SQL files you\u2019ll need later on.<\/p>\n<\/li>\n<li>\n<p>So browse to the directory to where you copied the <code>snowplow-emr-etl-runner<\/code> file, and run the following command:<\/p>\n<\/li>\n<\/ul>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-bash\" data-lang=\"bash\">git clone https:\/\/github.com\/snowplow\/snowplow.git<\/code><\/pre>\n<\/div>\n<p>If you don\u2019t have Git installed, now would be a good time to <a href=\"https:\/\/git-scm.com\/book\/en\/v2\/Getting-Started-Installing-Git\">do it<\/a>.<\/p>\n<div style=\"aspect-ratio: 1015 \/ 217;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/git-clone.jpg\" title=\"Git clone the Snowplow repo\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"217\" width=\"1015\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/git-clone.jpg#ZgotmplZ\" alt=\"Git clone the Snowplow repo\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Now you should have both the <code>snowplow-emr-etl-runner<\/code> file and the <strong>snowplow<\/strong> folder in the same directory.<\/p>\n<\/li>\n<li>\n<p>Next, create a new folder named <code>config<\/code>, and in that, a new folder named <code>targets<\/code>.<\/p>\n<\/li>\n<li>\n<p>Then, perform the following copy operations:<\/p>\n<ol>\n<li>\n<p>Copy <code>snowplow\/3-enrich\/emr-etl-runner\/config\/config.yml.sample<\/code> to <code>config\/config.yml<\/code>.<\/p>\n<\/li>\n<li>\n<p>Copy <code>snowplow\/3-enrich\/config\/iglu_resolver.json<\/code> to <code>config\/iglu_resolver.json<\/code>.<\/p>\n<\/li>\n<li>\n<p>Copy <code>snowplow\/4-storage\/config\/targets\/redshift.json<\/code> to <code>config\/targets\/redshift.json<\/code>.<\/p>\n<\/li>\n<\/ol>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 617 \/ 104;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/copy-config-files.jpg\" title=\"Copy config files\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"104\" width=\"617\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/copy-config-files.jpg#ZgotmplZ\" alt=\"Copy config files\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>In the end, you should end up with a folder and file structure like this:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-bash\" data-lang=\"bash\">.\n|-- snowplow-emr-etl-runner\n|-- snowplow\n|   |-- -SNOWPLOW GIT REPO HERE-\n|-- config\n|   |-- iglu_resolver.json\n|   |-- config.yml\n|   |-- targets\n|       |-- redshift.json <\/code><\/pre>\n<\/div>\n<h3 id=\"create-an-ec2-key-pair\">Create an EC2 key pair<\/h3>\n<p>At this point, you\u2019ll also need to create a private key pair in Amazon EC2. The ETL process will run on virtual machines in the Amazon cloud, and these machines are powered by Amazon EC2. For the runner to have full privileges to create and manage these machines, you will need to provide it with access control rights, and that\u2019s what the key pair is for.<\/p>\n<ul>\n<li>\n<p>In AWS, select <strong>Services<\/strong> from the top navigation, and click on <strong>EC2<\/strong>. In the left-hand menu, browse down to <strong>Key Pairs<\/strong>, and click the link.<\/p>\n<\/li>\n<li>\n<p>At this point, make sure you are in the Region where you\u2019ll be running all the proceeding jobs. For consistency\u2019s sake, it makes sense to stay in the same Region you\u2019ve been in all along. Remember, you can choose the Region from the top navigation.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 879 \/ 405;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/region-selection.jpg\" title=\"Region selection\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"405\" width=\"879\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/region-selection.jpg#ZgotmplZ\" alt=\"Region selection\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Once you\u2019ve made sure you\u2019re in the correct Region, click <strong>Create Key Pair<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Give the key pair a name you\u2019ll remember. My key pair is named <code>simoahava<\/code>.<\/p>\n<\/li>\n<li>\n<p>Once you\u2019re done, you\u2019ll see your new key pair in the list, and the browser automatically downloads the file <code><key pair=\"\" name=\"\">.pem<\/key><\/code> to your computer.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 647 \/ 194;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/create-key-pair.jpg\" title=\"Create EC2 key pair\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"194\" width=\"647\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/create-key-pair.jpg#ZgotmplZ\" alt=\"Create EC2 key pair\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h3 id=\"create-the-s3-buckets\">Create the S3 buckets<\/h3>\n<p>At this time, you\u2019ll need to create a bunch of buckets (storage locations) in Amazon S3. These will be used by the batch process to manage all your data files through various stages of the ETL process.<\/p>\n<p>You will need buckets for the following:<\/p>\n<ul>\n<li>\n<p><code>:raw:in<\/code> &#8211; you already have this, actually. It\u2019s the <strong>elasticbeanstalk-region-id<\/strong> created by the Clojure collector running in Elastic Beanstalk.<\/p>\n<\/li>\n<li>\n<p><code>:processing<\/code> &#8211; intermediate bucket for the log files before they are enriched.<\/p>\n<\/li>\n<li>\n<p><code>:archive<\/code> &#8211; you\u2019ll need three different archive buckets: <code>:raw<\/code> (for the raw log files), <code>:enriched<\/code> (for the enriched files), <code>:shredded<\/code> (for the shredded files).<\/p>\n<\/li>\n<li>\n<p><code>:enriched<\/code> &#8211; you\u2019ll need two buckets for enriched data: <code>:good<\/code> (for data sets successfully enriched), <code>:bad<\/code> (for those that failed enrich).<\/p>\n<\/li>\n<li>\n<p><code>:shredded<\/code> &#8211; you\u2019ll likewise need two buckets for shredded data: <code>:good<\/code> (for data sets successfully shredded), <code>:bad<\/code> (for those that failed shredding).<\/p>\n<\/li>\n<li>\n<p><code>:log<\/code> &#8211; a bucket for logs produced by the ETL process.<\/p>\n<\/li>\n<\/ul>\n<p>To create these buckets, head on over to S3 by clicking <strong>Services<\/strong> in the AWS top navigation, and choosing <strong>S3<\/strong>.<\/p>\n<p>You should already have your <code>:raw:in<\/code> bucket here, it\u2019s the one whose name starts with <strong>elasticbeanstalk-<\/strong>.<\/p>\n<p>Let\u2019s start with creating a new bucket that will contain all the \u201csub-buckets\u201d for ETL.<\/p>\n<p>Click <strong>+Create bucket<\/strong>, and name the bucket something like <strong>simoahava-snowplow-data<\/strong>. The bucket name must be unique across all of S3, so you can\u2019t just name it <strong>snowplow<\/strong>. Click <strong>Next<\/strong> a couple of times and then finally <strong>Create bucket<\/strong> to create this root bucket.<\/p>\n<div style=\"aspect-ratio: 722 \/ 514;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/create-root-bucket.jpg\" title=\"create root bucket in S3\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"514\" width=\"722\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/create-root-bucket.jpg#ZgotmplZ\" alt=\"create root bucket in S3\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Now click the new bucket name to open the bucket. You should see a screen like this:<\/p>\n<div style=\"aspect-ratio: 1807 \/ 754;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/empty-bucket.jpg\" title=\"Empty S3 bucket\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"754\" width=\"1807\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/empty-bucket.jpg#ZgotmplZ\" alt=\"Empty S3 bucket\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Click <strong>+ Create folder<\/strong>, and create the following three folders into this empty bucket:<\/p>\n<ol>\n<li>\n<p><strong>archive<\/strong><\/p>\n<\/li>\n<li>\n<p><strong>enriched<\/strong><\/p>\n<\/li>\n<li>\n<p><strong>shredded<\/strong><\/p>\n<\/li>\n<\/ol>\n<div style=\"aspect-ratio: 1033 \/ 447;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/s3-create-folder.jpg\" title=\"Create folder in S3\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"447\" width=\"1033\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/s3-create-folder.jpg#ZgotmplZ\" alt=\"Create folder in S3\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Then, in <strong>archive<\/strong>, create the following three folders:<\/p>\n<ol>\n<li>\n<p><strong>raw<\/strong><\/p>\n<\/li>\n<li>\n<p><strong>enriched<\/strong><\/p>\n<\/li>\n<li>\n<p><strong>shredded<\/strong><\/p>\n<\/li>\n<\/ol>\n<p>Next, in both <strong>enriched<\/strong> and <strong>shredded<\/strong>, create the following two folders:<\/p>\n<ol>\n<li>\n<p><strong>good<\/strong><\/p>\n<\/li>\n<li>\n<p><strong>bad<\/strong><\/p>\n<\/li>\n<\/ol>\n<p>Thus, you should end up with a bucket that has the following structure:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-bash\" data-lang=\"bash\">.\n|-- elasticbeanstalk-region-id\n|-- simoahava-snowplow-data\n|   |-- archive\n|   |   |-- raw\n|   |   |-- enriched\n|   |   |-- shredded\n|   |-- encriched\n|   |   |-- good\n|   |   |-- bad\n|   |-- shredded\n|   |   |-- good\n|   |   |-- bad<\/code><\/pre>\n<\/div>\n<p>Finally, create one more bucket in the root of S3 named something like <strong>simoahava-snowplow-log<\/strong>. You\u2019ll use this for the logs produced by the batch process.<\/p>\n<h3 id=\"prepare-for-configuring-the-emretlrunner\">Prepare for configuring the EmrEtlRunner<\/h3>\n<p>Now you\u2019ll need to configure the EmrEtlRunner. You\u2019ll use the <code>config.yml<\/code> file you copied from the Snowplow repo to the <code>config\/<\/code> folder. For the config, you\u2019ll need the following things:<\/p>\n<ol>\n<li>\n<p>The Access Key ID and Secret Access Key for the <code>snowplow-setup<\/code> user you created all the way back in <a href=\"#step-0-register-on-aws-and-setup-iam-roles\">Step 0<\/a>. If you didn\u2019t save these, you can generate a new Access Key through AWS <strong>IAM<\/strong>.<\/p>\n<\/li>\n<li>\n<p>You will need to download and install the <a href=\"https:\/\/aws.amazon.com\/cli\/\">AWS Command Line Interface<\/a>. You can use the <a href=\"https:\/\/docs.aws.amazon.com\/cli\/latest\/userguide\/cli-install-macos.html\">official guide<\/a> to install it with Python\/pip, but if you\u2019re running Mac OS X, I recommend using <a href=\"https:\/\/brew.sh\/\">Homebrew<\/a> instead. Once you\u2019ve installed Homebrew, you just need to run <code>brew install awscli<\/code> to install the AWS client.<\/p>\n<\/li>\n<\/ol>\n<p>Once you\u2019ve installed <code>awscli<\/code>, you need to run <code>aws configure<\/code> in your terminal, and do what it instructs you to do. You\u2019ll need your <strong>Access Key ID<\/strong> and <strong>Secret Access Key<\/strong> handy, as well as the region name (e.g. <code>eu-west-1<\/code>) where you\u2019ll be running your EC2 (again, I recommend to use the same region for all parts of this pipeline process).<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-bash\" data-lang=\"bash\">$ aws configure\nAWS Access Key ID: <enter your=\"\" iam=\"\" user=\"\" access=\"\" key=\"\" id=\"\" here=\"\">\nAWS Secret Access Key: <enter you=\"\" iam=\"\" user=\"\" secret=\"\" access=\"\" key=\"\" here=\"\">\nDefault region name: <enter the=\"\" region=\"\" name=\"\" e.g.=\"\" eu-west-1=\"\" here=\"\">\nDefault output format: <just press=\"\" enter=\"\"\/><\/enter><\/enter><\/enter><\/code><\/pre>\n<\/div>\n<p>This is what it looked like when I ran <code>aws configure<\/code>.<\/p>\n<div style=\"aspect-ratio: 548 \/ 142;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/aws-configure.jpg\" title=\"aws configure\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"142\" width=\"548\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/aws-configure.jpg#ZgotmplZ\" alt=\"aws configure\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>After running <code>aws configure<\/code>, the next command you\u2019ll need to run is <code>aws emr create-default-roles<\/code>. This will create default roles for the EmrEtlRunner, so that it can perform the necessary tasks in EC2 for you.<\/p>\n<p>Once you\u2019ve done these steps (remember to still keep your Access Key ID and Secret Access Key close by), you\u2019re ready to configure EmrEtlRunner!<\/p>\n<h3 id=\"configure-emretlrunner\">Configure EmrEtlRunner<\/h3>\n<p><strong>EmrEtlRunner<\/strong> is the name of the utility you downloaded earlier, with the filename <code>snowplow-emr-etl-runner<\/code>.<\/p>\n<p>EmrEtlRunner does a LOT of things. See <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Batch-pipeline-steps#dataflow-diagram\">this diagram<\/a> to see an overview of the process. At this point, we\u2019ll do all the steps except for step 13, <strong>rdb_load<\/strong>. That\u2019s the step where the enriched and shredded data are copied into a relational database. We\u2019ll take care of that in the next step.<\/p>\n<p>Anyway, EmrEtlRunner is operated by <code>config.yml<\/code>, which you\u2019ve copied into the <code>config\/<\/code> directory. I\u2019ll show you the config I use, and highlight the parts you\u2019ll need to change.<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-txt\" data-lang=\"txt\">aws:\n  access_key_id: AKIAIBAWU2NAYME55123\n  secret_access_key: iEmruXM7dSbOemQy63FhRjzhSboisP5TcJlj9123\n  s3:\n    region: eu-west-1\n    buckets:\n      assets: s3:\/\/snowplow-hosted-assets\n      jsonpath_assets:\n      log: s3:\/\/simoahava-snowplow-log\n      raw:\n        in:\n          - s3:\/\/elasticbeanstalk-eu-west-1-375284143851\/resources\/environments\/logs\/publish\/e-f4pdn8dtsg\n        processing: s3:\/\/simoahava-snowplow-data\/processing\n        archive: s3:\/\/simoahava-snowplow-data\/archive\/raw\n      enriched:\n        good: s3:\/\/simoahava-snowplow-data\/enriched\/good\n        bad: s3:\/\/simoahava-snowplow-data\/enriched\/bad\n        errors: \n        archive: s3:\/\/simoahava-snowplow-data\/archive\/enriched\n      shredded:\n        good: s3:\/\/simoahava-snowplow-data\/shredded\/good\n        bad: s3:\/\/simoahava-snowplow-data\/shredded\/bad\n        errors:\n        archive: s3:\/\/simoahava-snowplow-data\/archive\/shredded\n  emr:\n    ami_version: 5.9.0\n    region: eu-west-1\n    jobflow_role: EMR_EC2_DefaultRole\n    service_role: EMR_DefaultRole\n    placement:\n    ec2_subnet_id: subnet-d6e91a9e\n    ec2_key_name: simoahava\n    bootstrap: []\n    software:\n      hbase:\n      lingual:\n    jobflow:\n      job_name: Snowplow ETL\n      master_instance_type: m1.medium\n      core_instance_count: 2\n      core_instance_type: m1.medium\n      core_instance_ebs:\n        volume_size: 100\n        volume_type: \"gp2\"\n        volume_iops: 400\n        ebs_optimized: false\n      task_instance_count: 0\n      task_instance_type: m1.medium\n      task_instance_bid: 0.015\n    bootstrap_failure_tries: 3\n    configuration:\n      yarn-site:\n        yarn.resourcemanager.am.max-attempts: \"1\"\n      spark:\n        maximizeResourceAllocation: \"true\"\n    additional_info:\ncollectors:\n  format: clj-tomcat\nenrich:\n  versions:\n    spark_enrich: 1.12.0\n  continue_on_unexpected_error: false\n  output_compression: NONE\nstorage:\n  versions:\n    rdb_loader: 0.14.0\n    rdb_shredder: 0.13.0\n    hadoop_elasticsearch: 0.1.0\nmonitoring:\n  tags: {}\n  logging:\n    level: DEBUG<\/code><\/pre>\n<\/div>\n<p>The keys you need to edit are listed next, with a comment on how to edit them. All the keys not listed below you can leave with their default values. I really recommend you read through the <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/Common-configuration\">configuration documentation<\/a> for ideas on how to modify the rest of the keys to make your setup more powerful.<\/p>\n<table>\n<thead>\n<tr>\n<th>Key<\/th>\n<th>Comment<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><code>:aws:access_key_id<\/code><\/td>\n<td>Copy-paste the Access Key ID of your IAM user here.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:secret_access_key<\/code><\/td>\n<td>Copy-paste the Secret Access Key of your IAM user here.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:region<\/code><\/td>\n<td>Set this to the region where your S3 buckets are located in.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:log<\/code><\/td>\n<td>Change this to the name of the S3 bucket you created for the ETL logs.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:raw:in<\/code><\/td>\n<td>This is the bucket where the Tomcat logs are automatically pushed to. Do <strong>not<\/strong> include the last folder in the path, because this might change with an auto-scaling environment. <strong>Note!<\/strong> Keep the hyphen in the beginning of the line as in the config file example!<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:raw:processing<\/code><\/td>\n<td>Set this to the respective processing bucket.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:raw:archive<\/code><\/td>\n<td>Set this to the archive bucket for raw data.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:enriched:good<\/code><\/td>\n<td>Set this to the enriched\/good bucket.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:enriched:bad<\/code><\/td>\n<td>Set this to the enriched\/bad bucket.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:enriched:errors<\/code><\/td>\n<td>Leave this empty.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:enriched:archive<\/code><\/td>\n<td>Set this to the archive bucket for enriched data.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:shredded:good<\/code><\/td>\n<td>Set this to the shredded\/good bucket.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:shredded:bad<\/code><\/td>\n<td>Set this to the shredded\/bad bucket.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:shredded:errors<\/code><\/td>\n<td>Leave this empty.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:s3:buckets:shredded:archive<\/code><\/td>\n<td>Set this to the archive bucket for shredded data.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:emr:region<\/code><\/td>\n<td>This should be the region where the EC2 job will run.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:emr:placement<\/code><\/td>\n<td>Leave this empty.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:emr:ec2_subnet_id<\/code><\/td>\n<td>The subnet ID of the Virtual Private Cloud the job will run in. You can use the same subnet ID used by the EC2 instance running your collector.<\/td>\n<\/tr>\n<tr>\n<td><code>:aws:emr:ec2_key_name<\/code><\/td>\n<td>The name of the EC2 Key Pair you created earlier.<\/td>\n<\/tr>\n<tr>\n<td><code>:collectors:format<\/code><\/td>\n<td>Set this to <strong>clj-tomcat<\/strong>.<\/td>\n<\/tr>\n<tr>\n<td><code>:monitoring:snowplow<\/code><\/td>\n<td>Remove this key and all its children (<code>:method<\/code>, <code>:app_id<\/code>, and <code>:collector<\/code>).<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Just two things to point out.<\/p>\n<p>First, when copying the <code>:aws:s3:buckets:raw:in<\/code> path, do not copy the last folder name. This is the instance ID. With an auto-scaling environment set for the collector, there might be multiple instance folders in this bucket. If you only name one folder, you\u2019ll risk missing out on a lot of data.<\/p>\n<div style=\"aspect-ratio: 2704 \/ 892;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/s3-raw-in-bucket.jpg\" title=\"Path to the raw:in bucket\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"892\" width=\"2704\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/s3-raw-in-bucket.jpg#ZgotmplZ\" alt=\"Path to the raw:in bucket\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>You can get the <code>:aws:emr:ec2_subnet_id<\/code> by opening the <strong>Services<\/strong> menu in AWS and clicking <strong>EC2<\/strong>. Click the link titled <strong>Running Instances<\/strong> (there should be 1 running instance &#8211; your collector). Scroll down the <strong>Description<\/strong> tab contents until you find the <strong>Subnet ID<\/strong> entry. Copy-paste that into the <code>aws:emr:ec2_subnet_id<\/code> field.<\/p>\n<div style=\"aspect-ratio: 2878 \/ 1072;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/subnet-id.jpg\" title=\"EC2 subnet ID\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1072\" width=\"2878\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/subnet-id.jpg#ZgotmplZ\" alt=\"EC2 subnet ID\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>If you\u2019ve followed all the steps in this chapter, you should now be set.<\/p>\n<p>You can verify everything works by running the following command in the directory where the <code>snowplow-emr-etl-runner<\/code> executable is, and where the <code>config<\/code> folder is located.<\/p>\n<p><code>.\/snowplow-emr-etl-runner run -c config\/config.yml -r config\/iglu_resolver.json<\/code><\/p>\n<div style=\"aspect-ratio: 1482 \/ 202;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/emr-etl-runner.jpg\" title=\"Successfully ran that emr-etl-runner\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"202\" width=\"1482\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/emr-etl-runner.jpg#ZgotmplZ\" alt=\"Successfully ran that emr-etl-runner\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>You can also follow the process in real-time by opening the <strong>Services<\/strong> menu in AWS and clicking <strong>EMR<\/strong>. There, you should see the job named <strong>Snowplow ETL<\/strong>. By clicking it, you can see all the steps it is running through. If the process ends in an error, you can also debug quite handily via this view, since you can see the exact step where the process failed.<\/p>\n<div style=\"aspect-ratio: 2878 \/ 998;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/aws-emr-view.jpg\" title=\"AWS EMR steps\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"998\" width=\"2878\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/aws-emr-view.jpg#ZgotmplZ\" alt=\"AWS EMR steps\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Once the ETL has successfully completed, you can check your S3 buckets again. Your Snowplow data buckets should now contain a lot of new stuff. The folder with the interesting data is <strong>archive<\/strong> \/ <strong>shredded<\/strong>. This is where the good shredded datasets are copied to, and corresponds to what would have been copied to the relational database had you set this up in this step.<\/p>\n<p>Anyway, with the ETL process up and running, just one more step remains in this monster of a guide: setting up AWS Redshift as the relational database where you\u2019ll be warehousing your analytics data!<\/p>\n<h3 id=\"what-you-should-have-after-this-step-2\">What you should have after this step<\/h3>\n<ol>\n<li>\n<p>The <code>snowplow-emr-etl-runner<\/code> executable configured with your <code>config.yml<\/code> file.<\/p>\n<\/li>\n<li>\n<p>New buckets in S3 to store all the files created by the batch process.<\/p>\n<\/li>\n<li>\n<p>The ETL job running without errors all the way to completion, enriching and shredding the raw Tomcat logs into relevant S3 buckets.<\/p>\n<\/li>\n<\/ol>\n<h2 id=\"step-4-load-the-data-into-redshift\">Step 4: Load the data into Redshift<\/h2>\n<h3 id=\"what-you-need-for-this-step-3\">What you need for this step<\/h3>\n<ol>\n<li>\n<p>The ETL process configured and available to run at your whim.<\/p>\n<\/li>\n<li>\n<p>Shredded files being stored in the <strong>archive<\/strong> \/ <strong>shredded<\/strong> S3 bucket.<\/p>\n<\/li>\n<li>\n<p>An SQL query client. I recommend <a href=\"http:\/\/www.sql-workbench.net\/\">SQL Workbench\/J<\/a>, which is free. That\u2019s the one I\u2019ll be using in this guide.<\/p>\n<\/li>\n<\/ol>\n<h3 id=\"getting-started-2\">Getting started<\/h3>\n<p>In this final step of this tutorial, we\u2019ll load the shredded data into Redshift tables. Redshift is a data warehouse service offered by AWS. We\u2019ll use it to build a relational database, where each table stores the information shredded from the Tomcat logs in a format easy to query with SQL. By the way, if you\u2019re unfamiliar with SQL, look no further than this great <a href=\"https:\/\/www.codecademy.com\/learn\/learn-sql\">Codecademy course<\/a> to get you started with the query language!<\/p>\n<p>The steps you\u2019ll take in this chapter are roughly these:<\/p>\n<ol>\n<li>\n<p>Create a new cluster and database in Redshift.<\/p>\n<\/li>\n<li>\n<p>Add users and all the necessary tables to the database.<\/p>\n<\/li>\n<li>\n<p>Configure the EmrEtlRunner to automatically load the shredded data into Redshift tables.<\/p>\n<\/li>\n<\/ol>\n<p>Once you\u2019re done, each time you run EmrEtlRunner, the tables will be populated with the parsed tracker data. You can then run SQL queries against this data, and use it to proceed with the two remaining steps (not covered in this guide) of the Snowplow pipeline: <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/getting-started-with-data-modeling\">data modeling<\/a> and <a href=\"https:\/\/github.com\/snowplow\/snowplow\/wiki\/getting-started-analyzing-snowplow-data\">analysis<\/a>.<\/p>\n<h3 id=\"create-the-cluster\">Create the cluster<\/h3>\n<p>In AWS, select <strong>Services<\/strong> from the top navigation and choose the <strong>Amazon Redshift<\/strong> service.<\/p>\n<p>Again, double-check that you are in the correct region (the same one where you\u2019ve been working on all along, or, at the very least, the one where your S3 logs are). Then click the <strong>Launch cluster<\/strong> button.<\/p>\n<div style=\"aspect-ratio: 1918 \/ 790;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/redshift-launch-cluster.jpg\" title=\"Launch Redshift cluster\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"790\" width=\"1918\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/redshift-launch-cluster.jpg#ZgotmplZ\" alt=\"Launch Redshift cluster\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Give the cluster an identifier. I named my cluster <code>simoahava<\/code>. Give a name to the database, too. The name I chose was <code>snowplow<\/code>.<\/p>\n<p>Keep the database port at its default value (<strong>5439<\/strong>).<\/p>\n<p>Add a username and password to your master user. This is the user you\u2019ll initially log in with, and it\u2019s the one you\u2019ll create the rest of the users and all the necessary tables with. Remember to write these down somewhere &#8211; you\u2019ll need them in just a bit.<\/p>\n<p>Click <strong>Continue<\/strong> when ready.<\/p>\n<div style=\"aspect-ratio: 2252 \/ 1114;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/cluster-details.jpg\" title=\"Redshift Cluster details\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1114\" width=\"2252\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/cluster-details.jpg#ZgotmplZ\" alt=\"Redshift Cluster details\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>In the next view, leave the two options at their default values. Node type should be <strong>dc2.large<\/strong>, and Cluster type should be <strong>Single Node<\/strong> (with <code>1<\/code> as the number of compute nodes used). Click <strong>Continue<\/strong> when ready.<\/p>\n<div style=\"aspect-ratio: 2190 \/ 1294;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/node-configuration.jpg\" title=\"Configure cluster node\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1294\" width=\"2190\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/node-configuration.jpg#ZgotmplZ\" alt=\"Configure cluster node\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>In the Additional Configuration view, you can leave most of the options at their default values. For the <strong>VPC security group<\/strong>, select the <strong>default<\/strong> group. The settings should thus be something like these:<\/p>\n<p><strong>Cluster parameter group<\/strong>: default-redshift-1.0<br \/><strong>Encrypt database<\/strong>: None<br \/><strong>Choose a VPC<\/strong>: Default VPC (\u2026)<br \/><strong>Cluster subnet group<\/strong>: default<br \/><strong>Publicly accessible<\/strong>: Yes<br \/><strong>Choose a public IP address<\/strong>: No<br \/><strong>Enhanced VPC Routing<\/strong>: No<br \/><strong>Availability zone<\/strong>: No Preference<br \/><strong>VPC security groups<\/strong>: default (sg-\u2026)<br \/><strong>Create CloudWatch Alarm<\/strong>: No<br \/><strong>Available roles<\/strong>: No selection<\/p>\n<p>Once done, click <strong>Continue<\/strong>.<\/p>\n<p>You can double-check your settings, and then just click <strong>Launch cluster<\/strong>.<\/p>\n<p>The cluster will take some minutes to launch. You can check the status of the cluster in the Redshift dashboard.<\/p>\n<div style=\"aspect-ratio: 2838 \/ 454;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/launching-cluster.jpg\" title=\"Launching Redshift cluster\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"454\" width=\"2838\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/launching-cluster.jpg#ZgotmplZ\" alt=\"Launching Redshift cluster\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Once the cluster has been launched, you are ready to log in and configure it!<\/p>\n<h3 id=\"configure-the-cluster-and-connect-to-it\">Configure the cluster and connect to it<\/h3>\n<p>The first thing you\u2019ll need to do is make sure the cluster accepts incoming connections from your local machine.<\/p>\n<p>So after clicking <strong>Services<\/strong> in the AWS top navigation and choosing <strong>Amazon Redshift<\/strong>, go to <strong>Clusters<\/strong> and then click the cluster name in the dashboard.<\/p>\n<p>Under <strong>Cluster Properties<\/strong>, click the link to the <strong>VPC security group<\/strong> (should be named something like <code>default (sg-1234abcd)<\/code>).<\/p>\n<div style=\"aspect-ratio: 748 \/ 502;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/vpc-security-group.jpg\" title=\"VPC Security Group\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"502\" width=\"748\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/vpc-security-group.jpg#ZgotmplZ\" alt=\"VPC Security Group\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>You should be transported to the <strong>EC2<\/strong> dashboard, and <strong>Security Groups<\/strong> should be active in the navigation menu.<\/p>\n<p>In the bottom of the screen, the settings for the security group you clicked should be visible.<\/p>\n<p>Select the <strong>Inbound<\/strong> tab, and make sure it shows a <strong>TCP<\/strong> connection with Port Range <strong>5439<\/strong> and <strong>0.0.0.0\/0<\/strong> as the Source. This means that all incoming TCP connections are permitted (you can change this to a more stricter policy later on).<\/p>\n<p>If the values are different, you can <strong>Edit<\/strong> the Inbound rule to match these.<\/p>\n<div style=\"aspect-ratio: 1912 \/ 837;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/edit-security-group.jpg\" title=\"Edit security group\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"837\" width=\"1912\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/edit-security-group.jpg#ZgotmplZ\" alt=\"Edit security group\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Now it\u2019s time to connect to the cluster. Go back to <strong>Amazon Redshift<\/strong>, and open your cluster settings as before. Copy the link to the cluster from the top of the settings list.<\/p>\n<div style=\"aspect-ratio: 727 \/ 190;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/redshift-cluster-url.jpg\" title=\"Redshift cluster address\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"190\" width=\"727\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/redshift-cluster-url.jpg#ZgotmplZ\" alt=\"Redshift cluster address\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Next, open the SQL query tool of your choice. I\u2019m using <a href=\"http:\/\/www.sql-workbench.net\/\">SQL Workbench\/J<\/a>. Select <strong>File<\/strong> \/ <strong>Connect Window<\/strong>, and create a new connection with the following settings changed from defaults:<\/p>\n<p><strong>Driver<\/strong>: Amazon Redshift (com.amazon.redshift.jdbc.Driver)<br \/><strong>URL<\/strong>: jdbc:redshift:\/\/cluster_url:cluster_port\/database_name<br \/><strong>Username<\/strong>: master_username<br \/><strong>Password<\/strong>: master_password<br \/><strong>Autocommit<\/strong>: Check<\/p>\n<p>In <strong>URL<\/strong>, copy-paste the Redshift URL with port after the colon and database name after the slash.<\/p>\n<p>As <strong>Username<\/strong> and <strong>Password<\/strong>, add the master username and master password you chose when creating the cluster.<\/p>\n<p>Make sure <strong>Autocommit<\/strong> is checked. These are settings I have:<\/p>\n<div style=\"aspect-ratio: 978 \/ 385;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/connect-cluster-settings.jpg\" title=\"Cluster connection settings\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"385\" width=\"978\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/connect-cluster-settings.jpg#ZgotmplZ\" alt=\"Cluster connection settings\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Once done, you can click <strong>OK<\/strong>, and the tool will connect to your cluster and database.<\/p>\n<p>Once connected, you can feed the command <code>SELECT current_database();<\/code> and click the <strong>Execute<\/strong> button to check if everything works. This is what you should see:<\/p>\n<div style=\"aspect-ratio: 980 \/ 419;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/test-sql.jpg\" title=\"Test SQL query\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"419\" width=\"980\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/test-sql.jpg#ZgotmplZ\" alt=\"Test SQL query\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>If the query returns the name of the database, you\u2019re good to go!<\/p>\n<h3 id=\"create-the-database-tables\">Create the database tables<\/h3>\n<p>First, we\u2019ll need to create the tables that will store the Google Analytics tracker data within them. The tables are loaded as <code>.sql<\/code> files, and these files contain DDL (data-definition language) constructions that build all the necessary schemas and tables.<\/p>\n<p>For this, you\u2019ll need access to the schema <code>.sql<\/code> files, which you\u2019ll find in the following locations within the snowplow Git repo:<\/p>\n<p>Load <code>atomic-def.sql<\/code> first, and run the file in your SQL query tool while connected to your Redshift database. You should see a message that the schema <code>atomic<\/code> and table <code>atomic.events<\/code> were created successfully.<\/p>\n<div style=\"aspect-ratio: 1540 \/ 600;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/atomic-def-run.jpg\" title=\"atomic-def.sql run successfully\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"600\" width=\"1540\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/atomic-def-run.jpg#ZgotmplZ\" alt=\"atomic-def.sql run successfully\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Next, run <code>manifest-def.sql<\/code> while connected to the database. You should see a message that the table <code>atomic.manifest<\/code> was created successfully.<\/p>\n<p>Now you need to load all the DDLs for the Google Analytics schemas. If you don\u2019t create these tables, then the ETL process will run into an error, where the utility tries to copy shredded events into non-existent tables.<\/p>\n<p>You can find all the required <code>.sql<\/code> files in the following three directories:<\/p>\n<p>You need to load all the <code>.sql<\/code> files in these three directories and run them while connected to your database. This will create a whole bunch of tables you\u2019ll need if you want to collect Google Analytics tracker data.<\/p>\n<p>It might be easiest to clone the <strong>iglu-central<\/strong> repo, and then load the <code>.sql<\/code> files into your query tool from the local directories.<\/p>\n<p>Once you\u2019re done, you should be able to run the following SQL query and see a list of all the tables you just created as a result (should be 40 in total):<\/p>\n<p><code>SELECT * FROM pg_tables WHERE schemaname=\"atomic\";<\/code><\/p>\n<div style=\"aspect-ratio: 2008 \/ 880;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/all-tables-created.jpg\" title=\"All tables created\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"880\" width=\"2008\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/all-tables-created.jpg#ZgotmplZ\" alt=\"All tables created\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h3 id=\"create-the-database-users\">Create the database users<\/h3>\n<p>Next thing we\u2019ll do is create three users:<\/p>\n<ul>\n<li>\n<p><code>storageloader<\/code>, who will be in charge of the ETL process.<\/p>\n<\/li>\n<li>\n<p><code>power_user<\/code>, who will have admin privileges, so you no longer have to log in with the master credentials.<\/p>\n<\/li>\n<li>\n<p><code>read_only<\/code>, who can query data and create his\/her own tables.<\/p>\n<\/li>\n<\/ul>\n<p>Make sure you\u2019re still connected to the database, and copy-paste the following SQL queries into the query window. For each <code>$password<\/code>, change it to a proper password, and make sure you write these user + password combinations down somewhere.<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-sql\" data-lang=\"sql\"><span style=\"color:#00a\">CREATE<\/span> <span style=\"color:#00a\">USER<\/span> storageloader PASSWORD <span style=\"color:#a50\">'$password'<\/span>;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">USAGE<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">SCHEMA<\/span> <span style=\"color:#00a\">atomic<\/span> <span style=\"color:#00a\">TO<\/span> storageloader;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">INSERT<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">ALL<\/span> TABLES <span style=\"color:#00a\">IN<\/span> <span style=\"color:#00a\">SCHEMA<\/span> <span style=\"color:#00a\">atomic<\/span> <span style=\"color:#00a\">TO<\/span> storageloader;\n\n<span style=\"color:#00a\">CREATE<\/span> <span style=\"color:#00a\">USER<\/span> read_only PASSWORD <span style=\"color:#a50\">'$password'<\/span>;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">USAGE<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">SCHEMA<\/span> <span style=\"color:#00a\">atomic<\/span> <span style=\"color:#00a\">TO<\/span> read_only;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">SELECT<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">ALL<\/span> TABLES <span style=\"color:#00a\">IN<\/span> <span style=\"color:#00a\">SCHEMA<\/span> <span style=\"color:#00a\">atomic<\/span> <span style=\"color:#00a\">TO<\/span> read_only;\n<span style=\"color:#00a\">CREATE<\/span> <span style=\"color:#00a\">SCHEMA<\/span> scratchpad;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">ALL<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">SCHEMA<\/span> scratchpad <span style=\"color:#00a\">TO<\/span> read_only;\n\n<span style=\"color:#00a\">CREATE<\/span> <span style=\"color:#00a\">USER<\/span> power_user PASSWORD <span style=\"color:#a50\">'$password'<\/span>;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">ALL<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">DATABASE<\/span> snowplow <span style=\"color:#00a\">TO<\/span> power_user;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">ALL<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">SCHEMA<\/span> <span style=\"color:#00a\">atomic<\/span> <span style=\"color:#00a\">TO<\/span> power_user;\n<span style=\"color:#00a\">GRANT<\/span> <span style=\"color:#00a\">ALL<\/span> <span style=\"color:#00a\">ON<\/span> <span style=\"color:#00a\">ALL<\/span> TABLES <span style=\"color:#00a\">IN<\/span> <span style=\"color:#00a\">SCHEMA<\/span> <span style=\"color:#00a\">atomic<\/span> <span style=\"color:#00a\">TO<\/span> power_user;<\/code><\/pre>\n<\/div>\n<p>Again, remember to change the three <code>$password<\/code> values to proper SQL user passwords.<\/p>\n<p>If all goes well, you should see 12 \u201cCOMMAND executed successfully\u201d statements.<\/p>\n<p>Finally, you need to grant ownership of all tables in schema <code>atomic<\/code> to <strong>storageloader<\/strong>, because this user will need to run some commands (specifically, <code>vacuum<\/code> and <code>analyze<\/code>) that only table owners can run.<\/p>\n<p>So, first run the following query in the database.<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-sql\" data-lang=\"sql\"><span style=\"color:#00a\">SELECT<\/span> <span style=\"color:#a50\">'ALTER TABLE atomic.'<\/span> || tablename ||<span style=\"color:#a50\">' OWNER TO storageloader;'<\/span>\n<span style=\"color:#00a\">FROM<\/span> pg_tables <span style=\"color:#00a\">WHERE<\/span> schemaname=<span style=\"color:#a50\">'atomic'<\/span> <span style=\"color:#00a\">AND<\/span> <span style=\"color:#00a\">NOT<\/span> tableowner=<span style=\"color:#a50\">'storageloader'<\/span>;<\/code><\/pre>\n<\/div>\n<p>In the query results, you should see a bunch of <code>ALTER TABLE atomic.* OWNER TO storageloader;<\/code> queries. Copy all of these, and paste them into the statement field as new queries. Then run the statements.<\/p>\n<div style=\"aspect-ratio: 2080 \/ 1148;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/change-owner.jpg\" title=\"Change ownership of all tables\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1148\" width=\"2080\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/change-owner.jpg#ZgotmplZ\" alt=\"Change ownership of all tables\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Now, if you run <code>SELECT * FROM pg_tables WHERE schemaname=\"atomic\" AND tableowner=\"storageloader\";<\/code>, you should see all the tables in the atomic schema as a result.<\/p>\n<p>You have successfully created the users and the tables in the database. All that\u2019s left is to configure the EmrEtlRunner to execute the final step of the ETL process, where the <code>storageloader<\/code> user copies all the data from the shredded files into the corresponding Redshift tables.<\/p>\n<h3 id=\"create-new-iam-role-for-database-loader\">Create new IAM role for database loader<\/h3>\n<p>The EmrEtlRunner will copy the files to Redshift using a utility called RDB Loader (Relational Database Loader). For this tool to work with sufficient privileges, you\u2019ll need to create a new <strong>IAM Role<\/strong>, which grants the Redshift cluster read-only access to your S3 buckets.<\/p>\n<ul>\n<li>\n<p>So, in AWS, click <strong>Services<\/strong> and select <strong>IAM<\/strong>.<\/p>\n<\/li>\n<li>\n<p>Select <strong>Roles<\/strong> from the left-hand navigation. Click the <strong>Create role<\/strong> button.<\/p>\n<\/li>\n<li>\n<p>In the <strong>Select type of trusted entity<\/strong> view, keep the default <strong>AWS Service<\/strong> selected, and choose <strong>Redshift<\/strong> from the list of services. In the <strong>Select your use case<\/strong> list, choose <strong>Redshift &#8211; Customizable<\/strong>, and then click <strong>Next: Permissions<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1522 \/ 1454;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/iam-role-redshift.jpg\" title=\"Grant Redshift permissions to IAM role\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1454\" width=\"1522\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/iam-role-redshift.jpg#ZgotmplZ\" alt=\"Grant Redshift permissions to IAM role\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>In the next view, find the policy named <strong>AmazonS3ReadOnlyAccess<\/strong>, and check the box next to it. Click <strong>Next: Review<\/strong>.<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1996 \/ 1220;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/amazon-s3-read-only.jpg\" title=\"AmazonS3ReadOnlyAccess rights\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"1220\" width=\"1996\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/amazon-s3-read-only.jpg#ZgotmplZ\" alt=\"AmazonS3ReadOnlyAccess rights\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Name the role something useful, such as <code>RedshiftS3Access<\/code> and click <strong>Create Role<\/strong> when ready.<\/p>\n<\/li>\n<li>\n<p>You should be back in the list of roles. Click the newly created <code>RedshiftS3Access<\/code> role to see its configuration. Copy the value in the <strong>Role ARN<\/strong> field to the clipboard. You\u2019ll need it very soon.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 2458 \/ 950;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/role-arn.jpg\" title=\"Role ARN for redshift IAM role\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"950\" width=\"2458\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/role-arn.jpg#ZgotmplZ\" alt=\"Role ARN for redshift IAM role\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>\n<p>Finally, select <strong>Services<\/strong> from AWS top navigation and choose the <strong>Amazon Redshift<\/strong> service. Click <strong>Clusters<\/strong> in the left-hand navigation to see the list of running clusters.<\/p>\n<\/li>\n<li>\n<p>Check the box next to your Snowplow cluster, and click <strong>Manage IAM Roles<\/strong>.<\/p>\n<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 2246 \/ 672;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/manage-iam-roles.jpg\" title=\"Manage IAM Roles\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"672\" width=\"2246\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/manage-iam-roles.jpg#ZgotmplZ\" alt=\"Manage IAM Roles\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<ul>\n<li>In the <strong>Available roles<\/strong> list, choose the role you just created, and then click <strong>Apply changes<\/strong> to apply the role to your cluster.<\/li>\n<\/ul>\n<div style=\"aspect-ratio: 1212 \/ 630;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/apply-role.jpg\" title=\"Apply IAM role\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"630\" width=\"1212\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/apply-role.jpg#ZgotmplZ\" alt=\"Apply IAM role\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>The Cluster Status should change to <strong>modifying<\/strong>. Once it\u2019s done, the status will change to <strong>available<\/strong>, and you can check if the role you assigned is labeled as <strong>in-sync<\/strong> by clicking <strong>Manage IAM Roles<\/strong> again.<\/p>\n<h3 id=\"edit-the-redshift-target-configuration\">Edit the Redshift target configuration<\/h3>\n<p>If you copied all necessary files back in <a href=\"#download-the-necessary-files\">Step 3<\/a>, your project <code>config\/<\/code> directory should include a <code>targets\/<\/code> folder with the file <code>redshift.json<\/code> in it. If you don\u2019t have it, go back to Step 3 and make sure you copy the <code>redshift.json<\/code> template to the correct folder.<\/p>\n<p>Once you\u2019ve found the template, open it for editing, and make sure it looks something like this:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-yaml\" data-lang=\"yaml\">{<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">    <\/span><span style=\"color:#00a\">\"schema\": <\/span><span style=\"color:#a50\">\"iglu:com.snowplowanalytics.snowplow.storage\/redshift_config\/jsonschema\/2-1-0\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">    <\/span><span style=\"color:#00a\">\"data\": <\/span>{<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"name\": <\/span><span style=\"color:#a50\">\"AWS Redshift enriched events storage\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"host\": <\/span><span style=\"color:#a50\">\"simoahava.coyhone1deuh.eu-west-1.redshift.amazonaws.com\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"database\": <\/span><span style=\"color:#a50\">\"snowplow\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"port\": <\/span><span style=\"color:#099\">5439<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"sslMode\": <\/span><span style=\"color:#a50\">\"DISABLE\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"username\": <\/span><span style=\"color:#a50\">\"storageloader\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"password\": <\/span><span style=\"color:#a50\">\"...\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"roleArn\": <\/span><span style=\"color:#a50\">\"arn:aws:iam::375284143851:role\/RedshiftS3Access\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"schema\": <\/span><span style=\"color:#a50\">\"atomic\"<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"maxError\": <\/span><span style=\"color:#099\">1<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"compRows\": <\/span><span style=\"color:#099\">20000<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"sshTunnel\": <\/span><span style=\"color:#00a\">null<\/span>,<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">        <\/span><span style=\"color:#00a\">\"purpose\": <\/span><span style=\"color:#a50\">\"ENRICHED_EVENTS\"<\/span><span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\">    <\/span>}<span style=\"color:#bbb\">\n<\/span><span style=\"color:#bbb\"\/>}<\/code><\/pre>\n<\/div>\n<p>Here are the fields you need to edit:<\/p>\n<ul>\n<li><strong>host<\/strong>: The URL of your Redshift cluster<\/li>\n<li><strong>database<\/strong>: The database name<\/li>\n<li><strong>username<\/strong>: storageloader<\/li>\n<li><strong>password<\/strong>: storageloader password<\/li>\n<li><strong>roleArn<\/strong>: The Role ARN of the IAM role you created in the previous step<\/li>\n<\/ul>\n<p>All the other options you can leave with their default values.<\/p>\n<h3 id=\"re-run-emretlrunner-through-the-whole-process\">Re-run EmrEtlRunner through the whole process<\/h3>\n<p>Now that you\u2019ve configured <strong>everything<\/strong>, you\u2019re ready to run the EmrEtlRunner with all the steps in the ETL process included. This means <strong>enrichment<\/strong> of the log data, <strong>shredding<\/strong> of the log data to atomic datasets, and <strong>loading<\/strong> these datasets into your Redshift tables.<\/p>\n<p>The command you\u2019ll need to run in the root of your project folder (where the <code>snowplow-emr-etl-runner<\/code> executable is) is this:<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-bash\" data-lang=\"bash\">.\/snowplow-emr-etl-runner run -c config\/config.yml -r config\/iglu_resolver.json -t config\/targets<\/code><\/pre>\n<\/div>\n<p>This command will process all the data in the <code>:raw:in<\/code> bucket (the one with all your Tomcat logs), and proceed to extract and transform them, and finally load them into your Redshift tables. The process will take a while, so go grab a coffee. Remember that you can check the status of the job by browsing to <strong>EMR<\/strong> via the AWS <strong>Services<\/strong> navigation.<\/p>\n<p>Once complete, you should see something like this in the command line:<\/p>\n<div style=\"aspect-ratio: 1722 \/ 338;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/emr-etl-complete.jpg\" title=\"EmrEtlRunner finished\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"338\" width=\"1722\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/emr-etl-complete.jpg#ZgotmplZ\" alt=\"EmrEtlRunner finished\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<h3 id=\"test-it\">Test it<\/h3>\n<p>Now you should be able to login to the database using the new <code>read_only<\/code> user. If you run the following query, it should return a list of timestamps and events for each Client ID visiting your site.<\/p>\n<div class=\"highlight\">\n<pre style=\"background-color:#fff;-moz-tab-size:4;-o-tab-size:4;tab-size:4\"><code class=\"language-sql\" data-lang=\"sql\"><span style=\"color:#00a\">SELECT<\/span> u.root_tstamp, u.client_id, h.<span style=\"color:#00a\">type<\/span>\n<span style=\"color:#00a\">FROM<\/span> <span style=\"color:#00a\">atomic<\/span>.com_google_analytics_measurement_protocol_user_1 <span style=\"color:#00a\">AS<\/span> u \n<span style=\"color:#00a\">JOIN<\/span> <span style=\"color:#00a\">atomic<\/span>.com_google_analytics_measurement_protocol_hit_1 <span style=\"color:#00a\">AS<\/span> h\n<span style=\"color:#00a\">ON<\/span> u.root_id = h.root_id\n<span style=\"color:#00a\">ORDER<\/span> <span style=\"color:#00a\">BY<\/span> root_tstamp <span style=\"color:#00a\">ASC<\/span><\/code><\/pre>\n<\/div>\n<div style=\"aspect-ratio: 1648 \/ 836;\" class=\"figure nocaption\">\n<p>    <a href=\"https:\/\/www.simoahava.com\/images\/2018\/02\/sql-example-query.jpg\" title=\"SQL Example Query\"><\/p>\n<p>    <img decoding=\"async\" class=\"fig-img\" height=\"836\" width=\"1648\" loading=\"lazy\" src=\"https:\/\/www.simoahava.com\/images\/2018\/02\/sql-example-query.jpg#ZgotmplZ\" alt=\"SQL Example Query\"\/><\/p>\n<p>    <\/a><\/p>\n<\/div>\n<p>Considering how much time you have probably put into making everything work (if following this guide diligently), I really hope it all works correctly.<\/p>\n<h2 id=\"wrapping-it-all-up\">Wrapping it all up<\/h2>\n<p>By following this guide, you should be able to set up an end-to-end Snowplow batch pipeline.<\/p>\n<ol>\n<li>\n<p>Google Tag Manager duplicates the payloads sent to Google Analytics, and sends these to your Amazon endpoint, using a custom domain name whose DNS records you have delegated to AWS.<\/p>\n<\/li>\n<li>\n<p>The endpoint is a collector which logs all the HTTP requests to Tomcat logs, and stores them in an S3 bucket.<\/p>\n<\/li>\n<li>\n<p>An ETL process is then run, enriching the stored data, and shredding it to atomic datasets, again stored in S3.<\/p>\n<\/li>\n<li>\n<p>Finally, the ETL runner copies these datasets into tables you\u2019ve set up in a new relational database running on an AWS Redshift cluster.<\/p>\n<\/li>\n<\/ol>\n<p>There are <strong>SO<\/strong> many moving parts here, that it\u2019s possible you\u2019ll get something wrong at some point. Just try to patiently walk through the steps in this guide to see if you\u2019ve missed anything.<\/p>\n<p>Feel free to ask questions in the comments, and maybe I or my readers will be able to help you along.<\/p>\n<p>You can also join the discussions in Snowplow\u2019s <a href=\"https:\/\/discourse.snowplowanalytics.com\/\">Discourse site<\/a> &#8211; I\u2019m certain the folks there are more than happy to help you if you run into trouble setting up the pipeline.<\/p>\n<p>Do note also that the setup outlined in this guide is very rudimentary. There are many ways you can, and should, optimize the process, such as:<\/p>\n<ol>\n<li>\n<p>Add SSL support to your Redshift cluster.<\/p>\n<\/li>\n<li>\n<p>Scale the instances (collector and the ETL process) correctly to account for peaks in traffic and dataset size.<\/p>\n<\/li>\n<li>\n<p>Move the EmrEtlRunner to AWS, too. There\u2019s no need to run it on your local machine.<\/p>\n<\/li>\n<li>\n<p>Schedule the EmrEtlRunner to run (at least) once a day, so that your database is refreshed with new data periodically.<\/p>\n<\/li>\n<\/ol>\n<p>Good luck!<\/p>\n<\/p><\/div>\n\n","protected":false},"excerpt":{"rendered":"<p>A recent guide of mine introduced the Google Analytics adapter in Snowplow. The idea was that you can duplicate the Google Analytics requests sent via Google Tag Manager and dispatch them to your Snowplow analytics pipeline, too. The pipeline then takes care of these duplicated requests, using the new adapter to automatically align the hits [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":111365,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[12765,247,1531,7003,43260,1706],"dealstore":[],"offerexpiration":[],"class_list":["post-111364","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-analytics","tag-full","tag-google","tag-setup","tag-snowplow","tag-tracking"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Snowplow: Full Setup With Google Analytics Tracking - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=111364\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Snowplow: Full Setup With Google Analytics Tracking - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"A recent guide of mine introduced the Google Analytics adapter in Snowplow. The idea was that you can duplicate the Google Analytics requests sent via Google Tag Manager and dispatch them to your Snowplow analytics pipeline, too. The pipeline then takes care of these duplicated requests, using the new adapter to automatically align the hits [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=111364\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-02-26T14:39:44+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1919\" \/>\n\t<meta property=\"og:image:height\" content=\"467\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"46 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=111364#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=111364\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Snowplow: Full Setup With Google Analytics Tracking\",\"datePublished\":\"2025-02-26T14:39:44+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=111364\"},\"wordCount\":8360,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=111364#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg\",\"keywords\":[\"Analytics\",\"Full\",\"Google\",\"Setup\",\"Snowplow\",\"Tracking\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=111364#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=111364\",\"url\":\"https:\/\/fivemor.com\/?p=111364\",\"name\":\"Snowplow: Full Setup With Google Analytics Tracking - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=111364#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=111364#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg\",\"datePublished\":\"2025-02-26T14:39:44+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=111364#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=111364\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=111364#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg\",\"width\":1919,\"height\":467},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=111364#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Snowplow: Full Setup With Google Analytics Tracking\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Snowplow: Full Setup With Google Analytics Tracking - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=111364","og_locale":"en_US","og_type":"article","og_title":"Snowplow: Full Setup With Google Analytics Tracking - Som2ny Network","og_description":"A recent guide of mine introduced the Google Analytics adapter in Snowplow. The idea was that you can duplicate the Google Analytics requests sent via Google Tag Manager and dispatch them to your Snowplow analytics pipeline, too. The pipeline then takes care of these duplicated requests, using the new adapter to automatically align the hits [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=111364","og_site_name":"Som2ny Network","article_published_time":"2025-02-26T14:39:44+00:00","og_image":[{"width":1919,"height":467,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg","type":"image\/jpeg"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"46 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=111364#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=111364"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Snowplow: Full Setup With Google Analytics Tracking","datePublished":"2025-02-26T14:39:44+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=111364"},"wordCount":8360,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=111364#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg","keywords":["Analytics","Full","Google","Setup","Snowplow","Tracking"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=111364#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=111364","url":"https:\/\/fivemor.com\/?p=111364","name":"Snowplow: Full Setup With Google Analytics Tracking - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=111364#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=111364#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg","datePublished":"2025-02-26T14:39:44+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=111364#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=111364"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=111364#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/02\/snowplow-etl-emr.jpg","width":1919,"height":467},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=111364#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Snowplow: Full Setup With Google Analytics Tracking"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/111364","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=111364"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/111364\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/111365"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=111364"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=111364"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=111364"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=111364"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=111364"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}