Big Data Integration and Processing

About Course

There are 7 modules in this course

At the end of the course, you will be able to:

*Retrieve data from example database and big data management systems *Describe the connections between data management operations and the big data processing patterns needed to utilize them in large-scale analytical applications *Identify when a big data problem needs data integration *Execute simple big data integration and processing on Hadoop and Spark platforms This course is for those new to data science. Completion of Intro to Big Data is recommended. No prior programming experience is needed, although the ability to install applications and utilize a virtual machine is necessary to complete the hands-on assignments. Refer to the specialization technical requirements for complete hardware and software specifications. Hardware Requirements: (A) Quad Core Processor (VT-x or AMD-V support recommended), 64-bit; (B) 8 GB RAM; (C) 20 GB disk free. How to find your hardware information: (Windows): Open System by clicking the Start button, right-clicking Computer, and then clicking Properties; (Mac): Open Overview by clicking on the Apple menu and clicking “About This Mac.” Most computers with 8 GB RAM purchased in the last 3 years will meet the minimum requirements.You will need a high speed internet connection because you will be downloading files up to 4 Gb in size. Software Requirements: This course relies on several open-source software tools, including Apache Hadoop. All required software can be downloaded and installed free of charge (except for data charges from your internet provider). Software requirements include: Windows 7+, Mac OS X 10.10+, Ubuntu 14.04+ or CentOS 6+ VirtualBox 5+.

Course Content

Module 1: Welcome to Big Data Integration and Processing

What is in this Course?

00:00
Summary of Big Data Modeling and Management

00:00
Why is Big Data Processing Different?

00:00

6 readings

1 discussion prompt

Module 2: Retrieving Big Data (Part 1)

2 readings

Module 3: Retrieving Big Data (Part 2)

3 readings

2 quizzes

1 discussion prompt

Module 4: Big Data Integration

Overview of Information Integration

00:00
A Data Integration Scenario

00:00
Integration for Multichannel Customer Analytics

00:00
Big Data Management and Processing Using Splunk and Datameer

00:00
Why Splunk?

00:00
Connected Cars with Ford’s OpenXC and Splunk

00:00
Big Data Management and Processing using Datameer

00:00
Installing Splunk Enterprise on Windows

00:00
Installing Splunk Enterprise on Linux

00:00
Exploring Splunk Queries

00:00
Optional: Creating Pivot Reports in Splunk

00:00

4 readings

2 quizzes

1 discussion prompt

Module 5: Processing Big Data

Big Data Processing Pipelines

00:00
Some High-Level Processing Operations in Big Data Pipelines

00:00
Aggregation Operations in Big Data Pipelines

00:00
Typical Analytical Operations in Big Data Pipelines

00:00
Overview of Big Data Processing Systems

00:00
The Integration and Processing Layer

00:00
Introduction to Apache Spark

00:00
Getting Started with Spark

00:00
WordCount in Spark

00:00

4 readings

2 quizzes

3 discussion prompts

Module 6: Big Data Analytics using Spark

Spark Core: Programming In Spark using RDDs in Pipelines

00:00
Spark Core: Transformations

00:00
Spark Core: Actions

00:00
Spark SQL

00:00
Spark Streaming

00:00
Spark MLLib

00:00
Spark GraphX

00:00
Exploring SparkSQL and Spark DataFrames

00:00
Analyzing Sensor Data with Spark Streaming

00:00

5 readings

2 quizzes

1 discussion prompt

Module 7: Learning by Doing: Putting MongoDB and Spark to Work

Student Ratings & Reviews

No Review Yet

About Course

There are 7 modules in this course

What Will You Learn?

Course Content

Module 1: Welcome to Big Data Integration and Processing

What is in this Course?

Summary of Big Data Modeling and Management

Why is Big Data Processing Different?

6 readings

Slides: Summary & Why Is Big Data Processing Different

Downloading and Installing the Cloudera VM Instructions (Windows)

Downloading and Installing the Cloudera VM Instructions (Mac)

Software Installation Frequently Asked Questions (FAQ)

Instructions for Downloading Hands On Datasets

Instructions for Starting Jupyter

1 discussion prompt

Getting to know you: Tell us about yourself and why you are taking this course.

Module 2: Retrieving Big Data (Part 1)

What is Data Retrieval? Part 1

What is Data Retrieval? Part 2

Querying Two Relations

Subqueries

Querying Relational Data with Postgres

2 readings

Slides: What is Data Retrieval?

Querying Relational Data with Postgres

Module 3: Retrieving Big Data (Part 2)

Querying JSON Data with MongoDB

Aggregation Functions

Querying Aerospike

Querying Documents in MongoDB

Exploring Pandas DataFrames

3 readings

Slides: Querying Data Part 2

Querying Documents in MongoDB

Exploring Pandas DataFrames

2 quizzes

Retrieving Big Data Quiz

Postgres, MongoDB, and Pandas

1 discussion prompt

Let’s Discuss: MongoDB

Module 4: Big Data Integration

Overview of Information Integration

A Data Integration Scenario

Integration for Multichannel Customer Analytics

Big Data Management and Processing Using Splunk and Datameer

Why Splunk?

Connected Cars with Ford’s OpenXC and Splunk

Big Data Management and Processing using Datameer

Installing Splunk Enterprise on Windows

Installing Splunk Enterprise on Linux

Exploring Splunk Queries

Optional: Creating Pivot Reports in Splunk

4 readings

Slides: Information Integration

Downloading Splunk Enterprise

Exploring Splunk Queries

Optional: Instructions for Splunk Pivot Tutorial

2 quizzes

Information Integration – Quiz

Hands-On With Splunk

1 discussion prompt

Let’s Discuss: Big Data Integration

Module 5: Processing Big Data

Big Data Processing Pipelines

Some High-Level Processing Operations in Big Data Pipelines

Aggregation Operations in Big Data Pipelines

Typical Analytical Operations in Big Data Pipelines

Overview of Big Data Processing Systems

The Integration and Processing Layer

Introduction to Apache Spark

Getting Started with Spark

WordCount in Spark

4 readings

Big Data Processing Pipelines Slides

Big Data Workflow Management

Slides for Big Data Processing Tools and Systems

WordCount in Spark

2 quizzes

Pipeline and Tools