Apache Flume: Distributed Log Collection for Hadoop

If your role includes moving datasets into Hadoop, this book will help you do it more efficiently using Apache Flume. From installation to customization, it’s a complete step-by-step guide on making the service work for you.
Code Files

Apache Flume: Distributed Log Collection for Hadoop

Steve Hoffman

If your role includes moving datasets into Hadoop, this book will help you do it more efficiently using Apache Flume. From installation to customization, it’s a complete step-by-step guide on making the service work for you.
Mapt Subscription
FREE
$0.00/m after trial
eBook
$10.00
RRP $21.99
Save 54%
Print + eBook
$36.99
RRP $36.99
What do I get with a Mapt subscription?
  • Unlimited access to all Packt’s 6,000+ eBooks and Videos
  • 100+ new titles a month, learning paths, assessments & code files
  • 1 Free eBook or Video to download and keep every month after trial
What do I get with an eBook?
  • Download this book in EPUB, PDF, MOBI formats
  • DRM FREE - read and interact with your content when you want, where you want, and how you want
  • Access this title in the Mapt reader
What do I get with Print & eBook?
  • Get a paperback copy of the book delivered to you
  • Download this book in EPUB, PDF, MOBI formats
  • DRM FREE - read and interact with your content when you want, where you want, and how you want
  • Access this title in the Mapt reader
What do I get with a Video?
  • Download this Video course in MP4 format
  • DRM FREE - read and interact with your content when you want, where you want, and how you want
  • Access this title in the Mapt reader
$0.00
$10.00
$36.99
$29.99 p/m after trial
RRP $21.99
RRP $36.99
Subscription
eBook
Print + eBook
Start 14 Day Trial

Frequently bought together


Apache Flume: Distributed Log Collection for Hadoop Book Cover
Apache Flume: Distributed Log Collection for Hadoop
$ 21.99
$ 10.00
Scala Design Patterns Book Cover
Scala Design Patterns
$ 43.99
$ 10.00
Buy 2 for $20.00
Save $45.98
Add to Cart

Book Details

ISBN 139781782167914
Paperback108 pages

Book Description

Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log data. Its main goal is to deliver data from applications to Apache Hadoop's HDFS. It has a simple and flexible architecture based on streaming data flows. It is robust and fault tolerant with many failover and recovery mechanisms.

Apache Flume: Distributed Log Collection for Hadoop covers problems with HDFS and streaming data/logs, and how Flume can resolve these problems. This book explains the generalized architecture of Flume, which includes moving data to/from databases, NO-SQL-ish data stores, as well as optimizing performance. This book includes real-world scenarios on Flume implementation.

Apache Flume: Distributed Log Collection for Hadoop starts with an architectural overview of Flume and then discusses each component in detail. It guides you through the complete installation process and compilation of Flume.

It will give you a heads-up on how to use channels and channel selectors. For each architectural component (Sources, Channels, Sinks, Channel Processors, Sink Groups, and so on) the various implementations will be covered in detail along with configuration options. You can use it to customize Flume to your specific needs. There are pointers given on writing custom implementations as well that would help you learn and implement them.

By the end, you should be able to construct a series of Flume agents to transport your streaming data and logs from your systems into Hadoop in near real time.

What You Will Learn

  • Understand the Flume architecture
  • Download and install open source Flume from Apache
  • Discover when to use a memory or file-backed channel
  • Understand and configure the Hadoop File System (HDFS) sink
  • Learn how to use sink groups to create redundant data flows
  • Configure and use various sources for ingesting data
  • Inspect data records and route to different or multiple destinations based on payload content
  • Transform data en-route to Hadoop
  • Monitor your data flows

Authors

Book Details

ISBN 139781782167914
Paperback108 pages
Read More

Read More Reviews

Recommended for You

Scala Design Patterns Book Cover
Scala Design Patterns
$ 43.99
$ 10.00
Apache Spark 2.x for Java Developers Book Cover
Apache Spark 2.x for Java Developers
$ 39.99
$ 10.00
Practical Data Wrangling Book Cover
Practical Data Wrangling
$ 23.99
$ 10.00
Learning Apache Cassandra - Second Edition Book Cover
Learning Apache Cassandra - Second Edition
$ 35.99
$ 10.00
Burp Suite Essentials Book Cover
Burp Suite Essentials
$ 17.99
$ 10.00
Learning Puppet for Windows Server Book Cover
Learning Puppet for Windows Server
$ 35.99
$ 10.00