You're reading from Apache Hadoop 3 Quick Start Guide

Product typeBook

Published inOct 2018

Reading LevelIntermediate

PublisherPackt

ISBN-139781788999830

Edition1st Edition

Languages

Java

Tools

Hadoop

Concepts

Data Analysis

Author (1)

Hrishikesh Vijay Karambelkar

Demystifying Hadoop Ecosystem Components

We have gone through the Apache Hadoop subsystem in detail in previous chapters. Although Hadoop is extensively known for its core components such as HDFS, MapReduce and YARN, it also offers a whole ecosystem that is supported by various components to ensure all your business needs are addressed end-to-end. One key reason behind this evolution is because Hadoop's core components offer processing and storage in a raw form, which requires an extensive amount of investment when building software from a grass-roots level.

The ecosystem components on top of Hadoop can therefore provide the rapid development of applications, ensuring better fault-tolerance, security, and performance over custom development done on Hadoop.

In this chapter, we cover the following topics:

Understanding Hadoop's Ecosystem
Working with Apache Kafka
Writing...

Technical requirements

You will need Eclipse development environment and Java 8 installed on your system where you can run/tweak these examples. If you prefer to use maven, then you will need maven installed to compile the code. To run the example, you also need Apache Hadoop 3.1 setup on Linux system. Finally, to use the Git repository of this book, you need to install Git.

The code files of this chapter can be found on GitHub:
https://github.com/PacktPublishing/Apache-Hadoop-3-Quick-Start-Guide/tree/master/Chapter7

Check out the following video to see the code in action:

http://bit.ly/2SBdnr4

Understanding Hadoop's Ecosystem

Hadoop is often used for historical data analytics, although a new trend is emerging where it is used for real-time data streaming as well. Considering the offerings of Hadoop's ecosystem, we have broadly categorized them into the following categories:

Data flow: This includes components that can transfer data to and from different subsystems to and from Hadoop including real-time, batch, micro-batching, and event-driven data processing.
Data engine and frameworks: This provides programming capabilities on top of Hadoop YARN or MapReduce.
Data storage: This category covers all types of data storage on top of HDFS.
Machine learning and analytics: This category covers big data analytics and machine learning on top of Apache Hadoop.
Search engine: This category covers search engines in both structured and unstructured Hadoop data.
Management...

Working with Apache Kafka

Apache Kafka provides a data streaming pipeline across the cluster through its message service. It ensures a high degree of fault tolerance and message reliability through its architecture, and it also guarantees to maintain message ordering from a producer. A record in Kafka is a (key-value) pair along with a timestamp and it usually contains a topic name. A topic is a category of records on which the communication takes place.

Kafka supports producer-consumer-based messaging, which means producers can produce messages that can be sent to consumers. It maintains a queue of messages, where there is also an offset that represents its position or index. Kafka can be deployed on a multi-node cluster, as shown in the following diagram, where two producers and three consumers have been used as an example:

Producers produce multiple topics through producer...

Writing Apache Pig scripts

Apache Pig allows users to write custom scripts on top of the MapReduce framework. Pig was founded to offer flexibility in terms of data programming over large data sets and non-Java programmers. Pig can apply multiple transformations on input data in order to produce output on top of a Java virtual machine or an Apache Hadoop multi-node cluster. Pig can be used as a part of ETL (Extract Transform Load) implementations for any big data project.

Setting up Apache Pig in your Hadoop environment is relatively easy compared to other software; all you need to do is download the Pig source and build it to a pig.jar file, which can be used for your programs. Pig-generated compiled artifacts can be deployed on a standalone JVM, Apache Spark, Apache Tez, and MapReduce, and Pig supports six different execution environments (both local and distributed). The respective...

Transferring data with Sqoop

The beauty of Apache Hadoop lies in its ability to work with multiple data formats. HDFS can reliably store information flowing from a variety of data sources, whereas Hadoop requires external interfaces to interact with storage repositories outside of HDFS. Sqoop helps you to address part of this problem by allowing users to extract structured data from a relational database to Apache Hadoop. Similarly, raw data can be processed in Hadoop, and the final results can be shared with traditional databases thanks to Sqoop's bidirectional interfacing capabilities.

Sqoop can be downloaded from the Apache site directly, and it supports client-server-based architecture. A server can be installed on one of the nodes, which then acts as a gateway for all Sqoop activities. A client can be installed on any machine, which will eventually connect with the server...

Writing Flume jobs

Apache Flume offers the service to feed logs containing unstructured information back to Hadoop. Flume works across any type of data source. Flume can receive both log data or continuous event data, and it consumes events, incremental logs from sources such as the application server, and social media events.

The following diagram illustrates how Flume works. When flume receives an event, it is persisted in a channel (or data store), such as a local file system, before it is removed and pushed to the target by Sink. In the case of Flume, a target can be HDFS storage, Amazon S3, or another custom application:

Flume also supports multipleFlume agents, as shown in the preceding data flow. Data can be collected, aggregated together, and then processed through a multi-agent complex workflow that is completely customizable by the end user. Flume provides message reliability...

Understanding Hive

Apache Hive was developed at Facebook to primarily address the data warehousing requirements of the Hadoop platform. It was created to utilize analysts with strong SQL capabilities to run queries on the Hadoop cluster for data analytics. Although we often talk about going unstructured and using NoSQL, Apache Hive still fits in with today's information landscape regarding big data.

Apache Hive provides an SQL-like query language called HiveQL. Hive queries can be deployed on MapReduce, Apache Tez, and Apache Spark as jobs, which in turn can utilize the YARN engine to run programs. Just like RDBMS, Apache Hive provides indexing support with different index types, such as bitmap, on your HDFS data storage. Data can be stored in different formats, such as ORC, Parquet, Textfile, SequenceFile, and so on.

Hive querying also supports extended User Defined Functions...

Using HBase for NoSQL storage

Apache HBase provides a distributed, columnar key-value-based storage on Apache Hadoop. It is best suited when you need to perform read-writes randomly on large and varying data stores. HBase is capable of distributing and sharding its data across multiple nodes of Apache Hadoop, and it also provides high availability through its automatic failover from one region server to another. Apache HBase can be run in two modes: standalone and distributed. In the standalone mode, HBase does not use HDFS and instead uses a local directory by default, whereas the distributed mode works on HDFS.

Apache HBase stores its data across multiple rows and columns, where each row consists of a row key and a column containing one or more values. A value can be one or more attributes. Column families are sets of columns that are collocated together for performance reasons...

Summary

In this chapter, we studied the different components of Hadoop's overall ecosystem and their tools for solving many complex industrial problems. We went through a brief overview of the tools and software that run on Hadoop, specifically Apache Kafka, Apache PIG, Apache Sqoop, and Apache Flume. We also covered SQL and NoSQL-based databases on Hadoop, which included Hive and HBase respectively.

In the next chapter, we will take a look at some analytics components along with more advanced topics in Hadoop.

The rest of the chapter is locked

You have been reading a chapter from

Apache Hadoop 3 Quick Start Guide

Published in: Oct 2018Publisher: PacktISBN-13: 9781788999830

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

undefined

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $15.99/month. Cancel anytime

Author (1)

Hrishikesh Vijay Karambelkar

Hrishikesh Vijay Karambelkar is an innovator and an enterprise architect with 16 years of software design and development experience, specifically in the areas of big data, enterprise search, data analytics, text mining, and databases. He is passionate about architecting new software implementations for the next generation of software solutions for various industries, including oil and gas, chemicals, manufacturing, utilities, healthcare, and government infrastructure. In the past, he has authored three books for Packt Publishing: two editions of Scaling Big Data with Hadoop and Solr and one of Scaling Apache Solr. He has also worked with graph databases, and some of his work has been published at international conferences such as VLDB and ICDE.
Read more about Hrishikesh Vijay Karambelkar

Other recommended products

Related to this chapter

Hadoop 2.x Administration Cookbook

A practical and use case driven approach to Hadoop administration with coverage on a vast array of topics including Hadoop cluster installation, performance tuning, cluster planning, security, and much more. This book covers Hadoop from the perspective of running clusters in critical and large environments with complex data and at scale.

BookMay 2017348 pages

Mastering Hadoop 3

This is a comprehensive guide to understand advanced concepts of Hadoop ecosystem. You will learn how Hadoop works internally, and build solutions to some of real world use cases. Finally, you will have a solid understanding of how components in the Hadoop ecosystem are effectively integrated to implement a fast and reliable Big Data pipeline

BookFeb 2019544 pages

Mastering Apache Storm

With real-world examples and clear explanations, this book will ensure you will have a thorough mastery Apache Storm.You’ll get an understanding of deploying Storm on clusters. Introduce yourself to topics such as trident topology, monitoring, Storm Parallelism, scheduler and log processing. Learn how to integrate Storm with other well-known Big Data technologies such as HBase, Redis, Kafka, and Hadoop to realize the full potential of Storm.You will be able to use the knowledge to develop efficient, distributed real-time applications to cater to your business needs.

BookAug 2017284 pages

Apache Hive Essentials

Apache Hive helps you deal with data summarization, queries, and analysis for huge amounts of data. This book will give you a background in big data, and familiarize you with your Hive working environment. Next you will cover advanced topics like performance and security in Hive and how to work efficiently to find solutions to big data problems.

BookJun 2018210 pages

HBase High Performance Cookbook

BookJan 2017350 pages

Big Data Analytics with Hadoop 3

Apache Hadoop is the most popular platform for big data processing to build powerful analytics solutions. This book shows you how to do just that, with the help of practical examples. You will be well-versed with the analytical capabilities of Hadoop ecosystem with Apache Spark and Apache Flink to perform big data analytics by the end of this book.

BookMay 2018482 pages

Data Lake for Enterprises

The term 'Data Lake' has recently emerged as a prominent term in the big data industry. Data scientists can make use of it in deriving meaningful insights which can be used by businesses to redefine or transform the way they operate. Lambda architecture is also emerging as one of the very eminent patterns in the big data landscape, as it helps to derive useful information from not only the historical data but also correlates real-time data to enable business for taking critical decisions. This book tries to bring these two important aspects into one, namely data lake and lambda architecture.

BookMay 2017596 pages

Modern Big Data Processing with Hadoop

This book presents unique techniques to conquer different Big Data processing and analytics challenges using Hadoop. Practical examples are provided to boost your understanding of Big Data concepts and their implementation. By the end of the book, you will have all the knowledge and skills you need to become a true Big Data expert.

BookMar 2018394 pages

Hands-on DevOps

VideoDec 20170

Practical Big Data Analytics

Big Data analytics relates to the strategies used by enterprises to process and analyze large amounts of data to bring out hidden insights. With the help of open source and enterprise tools, such as R, Python, Hadoop, and Spark, you will learn how to effectively mine your Big Data. By the end of this book, you will have a clear understanding of how you can develop your own Big Data analytics solutions using different tools and methods.

BookJan 2018412 pages

Practical Real-time Data Processing and Analytics

Real-time data processing involves continuous input, processing and output of data, with the condition that the time required for processing is as short as possible. This book covers the majority of the existing and evolving open source technology stack for real-time processing and analytics. You will get to know about all the real-time solution aspects, from the source to the presentation to persistence. Through this practical book, you’ll be equipped with a clear understanding of how to solve challenges on your own.

BookSep 2017360 pages

Apache Spark 2.x for Java Developers

Apache Spark is the buzzword in the big data industry right now, especially with the increasing need for real-time streaming and data processing. While Spark is built on Scala, the Spark Java API exposes all the Spark features available in the Scala version for Java developers. This book will show you how you can implement various functionalities of the Apache Spark framework in Java, without stepping out of your comfort zone.

BookJul 2017350 pages

Personalised recommendations for you

Based on your interests and search pattern

Et al.

Ever wonder why speech recognition systems don't understand the Scottish accent, or what would happen if an astronaut only ate mac 'n' cheese, or other spurious reflections you'd have at a bar? We did, then collated those deliberations into absurd research articles with fake figures and methodologies inspired by even more fictionally absurd studies.

BookAug 2023230 pages5

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages4

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages5

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages1

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages5

Mastering Tableau 2023

This book is a comprehensive resource to mastering your Tableau skills and becoming a BI expert. As you progress, you will learn how to build advanced dashboards and improve your storytelling to derive key business insight, as well as make you well-versed with advanced functionalities of Tableau in the business intelligence domain.

BookAug 2023684 pages

Building AI Applications with ChatGPT APIs

This guide covers all ChatGPT API features for effortless creation of robust AI powered apps. With its help, you’ll be able to leverage ChatGPT’s cutting-edge NLP models to take your app development skills to the next level. You’ll also work on ten exciting projects that will give you the practical know-how that you can apply to your existing applications.

BookSep 2023258 pages5

Building AI Applications with ChatGPT APIs

This guide covers all ChatGPT API features for effortless creation of robust AI powered apps. With its help, you’ll be able to leverage ChatGPT’s cutting-edge NLP models to take your app development skills to the next level. You’ll also work on ten exciting projects that will give you the practical know-how that you can apply to your existing applications.

BookSep 2023258 pages2

Data Engineering with AWS

Embark on a journey to master data engineering pipelines on AWS! Our book offers a hands-on experience of AWS services for ingesting, transforming, and consuming data. Whether you're an absolute beginner or someone with basic data engineering experience, this guide is an indispensable resource.

BookOct 2023636 pages5

Modern Data Architecture on AWS

Every organization wants an agile, performant, and cost-effective data platform that meets all their current and future business needs. Purpose-built AWS analytics services and their features play a big part in building such a modern data platform. This book brings to you all the design and architectural patterns that’ll help you achieve this goal.

BookAug 2023420 pages5

Practical Guide to Applied Conformal Prediction in Python

Discover the power of Conformal Prediction with the "Practical Guide to Applied Conformal Prediction in Python." Master the latest techniques to quantify uncertainty in machine learning and computer vision models, and seamlessly apply them to your industry applications.

BookDec 2023240 pages

TinyML Cookbook

With over 70 project-based recipes, the TinyML Cookbook is a practical guide that will help you to get the most out of your microcontrollers. It provides a comprehensive understanding of the theoretical foundations while giving you hands-on experience training ML models for deployment on Arduino Nano 33 BLE Sense, Raspberry Pi Pico, and SparkFun RedBoard Artemis Nano microcontrollers.

BookNov 2023664 pages