Reader small image

You're reading from  Modern Data Architecture on AWS

Product typeBook
Published inAug 2023
PublisherPackt
ISBN-139781801813396
Edition1st Edition
Concepts
Right arrow
Author (1)
Behram Irani
Behram Irani
author image
Behram Irani

Behram Irani is currently a technology leader with Amazon Web Services (AWS) specializing in data, analytics and AI/ML. He has spent over 18 years in the tech industry helping organizations, from start-ups to large-scale enterprises, modernize their data platforms. In the last 6 years working at AWS, Behram has been a thought leader in the data, analytics and AI/ML space; publishing multiple papers and leading the digital transformation efforts for many organizations across the globe. Behram has completed his Bachelor of Engineering in Computer Science from the University of Pune and has an MBA degree from the University of Florida.
Read more about Behram Irani

Right arrow

Data processing using AWS Glue

If you recall our conversations from the last few chapters, we kept bringing up AWS Glue for multiple use cases, including for data catalogs, crawlers, classifiers, and batch ingestion using connectors. Now, we come to Glue ETL, which is the most distinct feature of Glue. Since Glue is a fully managed and serverless service, it excels in data transformation types of tasks, usually undertaken by data engineering personas in an organization. You can create Glue ETL jobs using Spark, Python, or Ray. Spark is a common platform for creating distributed computing-based ETL jobs. Since EMR also provides Spark and Glue also has Spark, in the following table, let’s try to simplify certain scenarios where you would prefer to use one over the other:

lock icon
The rest of the page is locked
Previous PageNext Page
You have been reading a chapter from
Modern Data Architecture on AWS
Published in: Aug 2023Publisher: PacktISBN-13: 9781801813396

Author (1)

author image
Behram Irani

Behram Irani is currently a technology leader with Amazon Web Services (AWS) specializing in data, analytics and AI/ML. He has spent over 18 years in the tech industry helping organizations, from start-ups to large-scale enterprises, modernize their data platforms. In the last 6 years working at AWS, Behram has been a thought leader in the data, analytics and AI/ML space; publishing multiple papers and leading the digital transformation efforts for many organizations across the globe. Behram has completed his Bachelor of Engineering in Computer Science from the University of Pune and has an MBA degree from the University of Florida.
Read more about Behram Irani

EMR typical usage

Glue ETL typical usage

Since EMR alleviates all the infrastructure and operational heavy...