Spark for Developers Certificate for Krzysztof Zylak
Certificate ID:
569415
Authentication Code:
1b71d
Certified Person Name:
Krzysztof Zylak
Trainer Name:
Richard Naoufal
Duration Days:
3
Duration Hours:
21
Course Name:
Spark for Developers
Course Date:
4 December 2018 09:30 to 6 December 2018 16:30
Venue:
Cambridge
Course Outline:
- Scala primer
- A quick introduction to Scala
- Labs : Getting know Scala
- Spark Basics
- Background and history
- Spark and Hadoop
- Spark concepts and architecture
- Spark eco system (core, spark sql, mlib, streaming)
- Labs : Installing and running Spark
- First Look at Spark
- Running Spark in local mode
- Spark web UI
- Spark shell
- Analyzing dataset – part 1
- Inspecting RDDs
- Labs: Spark shell exploration
- RDDs
- RDDs concepts
- Partitions
- RDD Operations / transformations
- RDD types
- Key-Value pair RDDs
- MapReduce on RDD
- Caching and persistence
- Labs : creating & inspecting RDDs; Caching RDDs
- Spark API programming
- Introduction to Spark API / RDD API
- Submitting the first program to Spark
- Debugging / logging
- Configuration properties
- Labs : Programming in Spark API, Submitting jobs
- Spark SQL
- SQL support in Spark
- Dataframes
- Defining tables and importing datasets
- Querying data frames using SQL
- Storage formats : JSON / Parquet
- Labs : Creating and querying data frames; evaluating data formats
- MLlib
- MLlib intro
- MLlib algorithms
- Labs : Writing MLib applications
- GraphX
- GraphX library overview
- GraphX APIs
- Labs : Processing graph data using Spark
- Spark Streaming
- Streaming overview
- Evaluating Streaming platforms
- Streaming operations
- Sliding window operations
- Labs : Writing spark streaming applications
- Spark and Hadoop
- Hadoop Intro (HDFS / YARN)
- Hadoop + Spark architecture
- Running Spark on Hadoop YARN
- Processing HDFS files using Spark
- Spark Performance and Tuning
- Broadcast variables
- Accumulators
- Memory management & caching
- Spark Operations
- Deploying Spark in production
- Sample deployment templates
- Configurations
- Monitoring
- Troubleshooting