Reader small image

You're reading from  Data Wrangling with R

Product typeBook
Published inFeb 2023
PublisherPackt
ISBN-139781803235400
Edition1st Edition
Concepts
Right arrow
Author (1)
Gustavo R Santos
Gustavo R Santos
author image
Gustavo R Santos

Gustavo R Santos has worked in the Technology Industry for 13 years, improving processes, and analyzing datasets and creating dashboards. Since 2020, he has been working as a Data Scientist in the retail industry, wrangling, analyzing, visualizing and modeling data with the most modern tools like R, Python and Databricks. Gustavo also gives lectures from time to time at an online school about Data Science concepts. He has a background in Marketing, is certified as Data Scientist by the Data Science Academy Brazil and pursues his specialist MBA in Data Science at the University of São Paulo
Read more about Gustavo R Santos

Right arrow

Replacing and filling data

A dataset can and certainly will be acquired with imperfections. An example of imperfection is the use of the ? sign instead of the default NA for missing values for the Census Income dataset. This problem will require the question mark to be replaced with NA first, and then filled with another value, such as the mean, the most frequent observation, or using more complex methods, even machine learning.

This case clearly illustrates the necessity of replacing and filling data points from a dataset. Using tidyr, there are specific functions to replace and fill in missing data.

First, the ? sign needs to be replaced with NA, before we can think of filling the missing values. As seen in Chapter 7, there are only missing values for the workclass (1836), occupation (1843), and native_country (583) columns. To confirm that, a loop through the variables searching for ? would be the fastest resource:

# Loop through variables looking for cells == "...
lock icon
The rest of the page is locked
Previous PageNext Page
You have been reading a chapter from
Data Wrangling with R
Published in: Feb 2023Publisher: PacktISBN-13: 9781803235400

Author (1)

author image
Gustavo R Santos

Gustavo R Santos has worked in the Technology Industry for 13 years, improving processes, and analyzing datasets and creating dashboards. Since 2020, he has been working as a Data Scientist in the retail industry, wrangling, analyzing, visualizing and modeling data with the most modern tools like R, Python and Databricks. Gustavo also gives lectures from time to time at an online school about Data Science concepts. He has a background in Marketing, is certified as Data Scientist by the Data Science Academy Brazil and pursues his specialist MBA in Data Science at the University of São Paulo
Read more about Gustavo R Santos