Hadoop Cluster Deployment

Construct a modern Hadoop data platform effortlessly and gain insights into how to manage clusters efficiently

Hadoop Cluster Deployment

Progressing
Danil Zburivsky

Construct a modern Hadoop data platform effortlessly and gain insights into how to manage clusters efficiently
$10.00
$34.99
RRP $20.99
RRP $34.99
eBook
Print + eBook
$12.99 p/month

Get Access

Get Unlimited Access to every Packt eBook and Video course

Enjoy full and instant access to over 3000 books and videos – you’ll find everything you need to stay ahead of the curve and make sure you can always get the job done.

Code Files
+ Collection
Free Sample

Book Details

ISBN 139781783281718
Paperback126 pages

About This Book

  • Choose the hardware and Hadoop distribution that best suits your needs
  • Get more value out of your Hadoop cluster with Hive, Impala, and Sqoop
  • Learn useful tips for performance optimization and security

Who This Book Is For

This book is ideal for database administrators, data engineers, and system administrators, and it will act as an invaluable reference if you are planning to use the Hadoop platform in your organization. It is expected that you have basic Linux skills since all the examples in this book use this operating system. It is also useful if you have access to test hardware or virtual machines to be able to follow the examples in the book.

Table of Contents

Chapter 1: Setting Up Hadoop Cluster – from Hardware to Distribution
Choosing Hadoop cluster hardware
Hadoop distributions
Choosing OS for the Hadoop cluster
Summary
Chapter 2: Installing and Configuring Hadoop
Configuring OS for Hadoop cluster
Setting up NameNode
Summary
Chapter 3: Configuring the Hadoop Ecosystem
Hosting the Hadoop ecosystem
Sqoop
Hive
Impala
Summary
Chapter 4: Securing Hadoop Installation
Hadoop security overview
HDFS security
MapReduce security
Hadoop Service Level Authorization
Hadoop and Kerberos
Summary
Chapter 5: Monitoring Hadoop Cluster
Monitoring strategy overview
Hadoop Metrics
Monitoring MapReduce
Monitoring Hadoop with Ganglia
Summary
Chapter 6: Deploying Hadoop to the Cloud
Amazon Elastic MapReduce
Using Whirr
Summary

What You Will Learn

  • Choose the optimal hardware configuration for your Hadoop cluster
  • Decipher the differences between various Hadoop versions and distributions
  • Make your cluster crash-proof with Namenode High Availability
  • Learn tips and tricks for Jobtracker, Tasktracker, and Datanodes
  • Discover the most important Hadoop ecosystem projects
  • Get more value out of your cluster by using SQL with Hive and real-time query processing with Impala
  • Set up a proper permissions model for your cluster
  • Secure Hadoop with Kerberos
  • Deploy a Hadoop cluster in a cloud environment

 

In Detail

Big Data is the hottest trend in the IT industry at the moment. Companies are realizing the value of collecting, retaining, and analyzing as much data as possible. They are therefore rushing to implement the next generation of data platform, and Hadoop is the centerpiece of these platforms.

This practical guide is filled with examples which will show you how to successfully build a data platform using Hadoop. Step-by-step instructions will explain how to install, configure, and tie all major Hadoop components together. This book will allow you to avoid common pitfalls, follow best practices, and go beyond the basics when building a Hadoop cluster.

This book will walk you through the process of building a Hadoop cluster from the ground up. By using practical examples and command samples, you will be able to get a cluster up and running in no time, and you will also gain a deep understanding of how various Hadoop components work and interact with each other.

You will learn how to pick the right hardware for different types of Hadoop clusters and about the differences between various Hadoop distributions. By the end of this book, you will be able to install and configure several of the most popular Hadoop ecosystem projects including Hive, Impala, and Sqoop, and you will also be given a sneak peek into the pros and cons of using Hadoop in the cloud.

Authors

Table of Contents

Chapter 1: Setting Up Hadoop Cluster – from Hardware to Distribution
Choosing Hadoop cluster hardware
Hadoop distributions
Choosing OS for the Hadoop cluster
Summary
Chapter 2: Installing and Configuring Hadoop
Configuring OS for Hadoop cluster
Setting up NameNode
Summary
Chapter 3: Configuring the Hadoop Ecosystem
Hosting the Hadoop ecosystem
Sqoop
Hive
Impala
Summary
Chapter 4: Securing Hadoop Installation
Hadoop security overview
HDFS security
MapReduce security
Hadoop Service Level Authorization
Hadoop and Kerberos
Summary
Chapter 5: Monitoring Hadoop Cluster
Monitoring strategy overview
Hadoop Metrics
Monitoring MapReduce
Monitoring Hadoop with Ganglia
Summary
Chapter 6: Deploying Hadoop to the Cloud
Amazon Elastic MapReduce
Using Whirr
Summary

Book Details

ISBN 139781783281718
Paperback126 pages
Read More