Packt+ | Advance your knowledge in tech

You're reading from Learning Hunk

Product typeBook

Published inDec 2015

Reading LevelIntermediate

Publisher

ISBN-139781782174820

Edition1st Edition

Languages

Java

Tools

Hadoop

Concepts

Data Analysis

Authors (2):

Dmitry Anoshin

Sergey Sheypak

View More author details

Chapter 4. Adding Speed to Reports

One of the attributes of big data analytics is its velocity. In the modern world of information technology, speed is one of the crucial factors of any successful organization because even delays measured in seconds can cost money. Big data must move at extremely high velocities no matter how much we scale or what workloads our store must handle. The data handling hoops of Hadoop or NoSQL solutions put a serious drag on performance. That's why Hunk has a powerful feature that can speed up analytics and help immediately derive business insight from a vast amount of data.

In this chapter, we will learn about the report acceleration technique of Hunk, create new virtual indexes, and compare the performance of the same search with and without acceleration.

Big data performance issues

Despite the fact that, with modern technology, we can handle any big data issue, we still to have spend some time waiting for our questions to be answered. For example, we collect data and store it in Hadoop, then we deploy Hunk and configure a data provider, create a virtual index, and start to ask business questions by creating a query and running search commands. We should wait before the MapReduce job is finished. The following diagram illustrates this situation:

Moreover, if we want to ask the question over and over again by modifying the initial query, we will lose much time and money.

It would be superb if we could just run the search and immediately get the answer, as in the following diagram:

Yes, this is possible with Hunk, because it allows us to accelerate the report and get an answer to our business question very quickly. Let's learn how to do it.

Hunk report acceleration

We can easily accelerate our searches, which is critical for business. The idea behind Hunk is easy: the same search on the same data always gives the same result. In other words, same search + same data = same results. In the case of acceleration, Hunk caches the results and returns them on demand. Moreover, it gives us the opportunity to choose a data range for a particular data summary. In other words, if the data change is due to a fresh portion of events, then the accelerated report will rebuild the data summary in order to meet the requirements of the particular data range. Technically, we just cache the map phase in HDFS. When we run the accelerated search, Hunk just returns straight to us. There are four main steps in running an accelerated search:

The scheduled job builds a cache.
Find cache hits.
Stream the results to a search head.
Reduce on the search head.

Tip

There is more information about search heads at: http://docs.splunk.com/Splexicon:Searchhead.

The...

Hunk accelerations limits

Hunk is superb, but it still has some drawbacks:

Hardware limitations related to memory consumption by the search head. They are solved by adjusting the memory configuration.
Software limitations related to the cache. Sometimes we should delete the old cache using command rm -rf <vix.splunk.home.hdfs>/cash.
Human factor—this is a popular issue, especially in analytics. We learnt that, for Hunk, acceleration means: same search + same data = same result. But it won't work if we change the KV extraction rules. Be careful with this.

Summary

In this chapter we learnt how to accelerate searches in Hunk. This feature is easy to use and maintain. It helps to reduce resources and improve user experience. In addition, we learnt how acceleration works and created our own accelerated report. Moreover, we compared it with a normal report and figured out how it became faster. Finally, we learnt how to manage Hunk report summaries.

We also learnt much about the base functionality of Hunk. In the next chapter, we are going to extend Hunk functionality via Hunk SDK and Rest API.

The rest of the chapter is locked

You have been reading a chapter from

Learning Hunk

Published in: Dec 2015Publisher: ISBN-13: 9781782174820

A free Packt account unlocks extra newsletters, articles, discounted offers, and much more. Start advancing your knowledge today.

undefined

Unlock this book and the full library FREE for 7 days

Get unlimited access to 7000+ expert-authored eBooks and videos courses covering every tech area you can think of

Start free trial

Renews at $15.99/month. Cancel anytime

Authors (2)

Dmitry Anoshin

Dmitry Anoshin is a data-centric technologist and a recognized expert in building and implementing big data and analytics solutions. He has a successful track record when it comes to implementing business and digital intelligence projects in numerous industries, including retail, finance, marketing, and e-commerce. Dmitry possesses in-depth knowledge of digital/business intelligence, ETL, data warehousing, and big data technologies. He has extensive experience in the data integration process and is proficient in using various data warehousing methodologies. Dmitry has constantly exceeded project expectations when he has worked in the financial, machine tool, and retail industries. He has completed a number of multinational full BI/DI solution life cycle implementation projects. With expertise in data modeling, Dmitry also has a background and business experience in multiple relation databases, OLAP systems, and NoSQL databases. He is also an active speaker at data conferences and helps people to adopt cloud analytics.
Read more about Dmitry Anoshin

Sergey Sheypak

Sergey Sheypak started his so-called big data practice in 2010 as a Teradata PS consultant. His was leading the Teradata Master Data Management deployment in Sberbank, Russia (which has 110 billion customers). Later Sergey switched to AsterData and Hadoop practices. Sergey joined the Research and Development team at MegaFon (one of the top three telecom companies in Russia with 70 billion customers) in 2012. While leading the Hadoop team at MegaFon, Sergey built ETL processes from existing Oracle DWH to HDFS. Automated end-to-end tests and acceptance tests were introduced as a mandatory part of the Hadoop development process. Scoring geospatial analysis systems based on specific telecom data were developed and launched. Now, Sergey works as independent consultant in Sweden.
Read more about Sergey Sheypak

Personalised recommendations for you

Based on your interests and search pattern

Et al.

Ever wonder why speech recognition systems don't understand the Scottish accent, or what would happen if an astronaut only ate mac 'n' cheese, or other spurious reflections you'd have at a bar? We did, then collated those deliberations into absurd research articles with fake figures and methodologies inspired by even more fictionally absurd studies.

BookAug 2023230 pages5

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages4

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages5

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages1

Generative AI with LangChain

This book is a comprehensive introduction to LLMs and LangChain, demystifying the basic mechanics of LangChain, its functionalities, and the myriad of applications it can be integrated into.

BookDec 2023360 pages5

Mastering Tableau 2023

This book is a comprehensive resource to mastering your Tableau skills and becoming a BI expert. As you progress, you will learn how to build advanced dashboards and improve your storytelling to derive key business insight, as well as make you well-versed with advanced functionalities of Tableau in the business intelligence domain.

BookAug 2023684 pages

Building AI Applications with ChatGPT APIs

This guide covers all ChatGPT API features for effortless creation of robust AI powered apps. With its help, you’ll be able to leverage ChatGPT’s cutting-edge NLP models to take your app development skills to the next level. You’ll also work on ten exciting projects that will give you the practical know-how that you can apply to your existing applications.

BookSep 2023258 pages5

Building AI Applications with ChatGPT APIs

This guide covers all ChatGPT API features for effortless creation of robust AI powered apps. With its help, you’ll be able to leverage ChatGPT’s cutting-edge NLP models to take your app development skills to the next level. You’ll also work on ten exciting projects that will give you the practical know-how that you can apply to your existing applications.

BookSep 2023258 pages2

Data Engineering with AWS

Embark on a journey to master data engineering pipelines on AWS! Our book offers a hands-on experience of AWS services for ingesting, transforming, and consuming data. Whether you're an absolute beginner or someone with basic data engineering experience, this guide is an indispensable resource.

BookOct 2023636 pages5

Modern Data Architecture on AWS

Every organization wants an agile, performant, and cost-effective data platform that meets all their current and future business needs. Purpose-built AWS analytics services and their features play a big part in building such a modern data platform. This book brings to you all the design and architectural patterns that’ll help you achieve this goal.

BookAug 2023420 pages5

Practical Guide to Applied Conformal Prediction in Python

Discover the power of Conformal Prediction with the "Practical Guide to Applied Conformal Prediction in Python." Master the latest techniques to quantify uncertainty in machine learning and computer vision models, and seamlessly apply them to your industry applications.

BookDec 2023240 pages

TinyML Cookbook

With over 70 project-based recipes, the TinyML Cookbook is a practical guide that will help you to get the most out of your microcontrollers. It provides a comprehensive understanding of the theoretical foundations while giving you hands-on experience training ML models for deployment on Arduino Nano 33 BLE Sense, Raspberry Pi Pico, and SparkFun RedBoard Artemis Nano microcontrollers.

BookNov 2023664 pages