BookFrontier
Learning Spark by Jules Damji

Book

Learning Spark

Lightning-fast Data Analytics

Jules Damji, Brooke Wenig, Tathagata Das

O'Reilly Media · Print & ebook · August 25, 2020

Reading lane: Data Mining

Data is bigger, arrives faster, and comes in a variety of formatsâ??and it all needs to be processed at scale for analytics or machine learning.

At a Glance

Who It's For

Data engineers and data scientists learning Apache SparkReaders seeking hands-on Spark analytics and machine learning guidance

Book Details

Authors
Jules Damji, Brooke Wenig, Tathagata Das
Publisher
O'Reilly Media
Published
August 25, 2020
Format
Print & ebook
Theme
Data Mining · Data Warehousing
Reading lane
Data Mining

Affinity

Publisher Categories

  • Productivity Software

  • Data Science

  • Java Programming

  • Open Source Development

Show all 7 publisher categories
  • Math & Stats Software

  • Mathematical Analysis

  • Complex Analysis

About This Book

Data is bigger, arrives faster, and comes in a variety of formatsâ??and it all needs to be processed at scale for analytics or machine learning. But how can you process such varied workloads efficiently? Enter Apache Spark. Updated to include Spark 3.0, this second edition shows data engineers and data scientists why structure and unification in Spark matters. Specifically, this book explains how to perform simple and complex data analytics and employ machine learning algori...

Read full description

Data is bigger, arrives faster, and comes in a variety of formatsâ??and it all needs to be processed at scale for analytics or machine learning. But how can you process such varied workloads efficiently? Enter Apache Spark. Updated to include Spark 3.0, this second edition shows data engineers and data scientists why structure and unification in Spark matters. Specifically, this book explains how to perform simple and complex data analytics and employ machine learning algorithms. Through step-by-step walk-throughs, code snippets, and notebooks, youâ??ll be able to: - Learn Python, SQL, Scala, or Java high-level Structured APIs - Understand Spark operations and SQL Engine - Inspect, tune, and debug Spark operations with Spark configurations and Spark UI - Connect to data sources: JSON, Parquet, CSV, Avro, ORC, Hive, S3, or Kafka - Perform analytics on batch and streaming data using Structured Streaming - Build reliable data pipelines with open source Delta Lake and Spark - Develop machine learning pipelines with MLlib and productionize models using MLflow

Similar Books