Reader small image

You're reading from  Developing Kaggle Notebooks

Product typeBook
Published inDec 2023
Reading LevelIntermediate
PublisherPackt
ISBN-139781805128519
Edition1st Edition
Languages
Right arrow
Author (1)
Gabriel Preda
Gabriel Preda
author image
Gabriel Preda

Dr. Gabriel Preda is a Principal Data Scientist for Endava, a major software services company. He has worked on projects in various industries, including financial services, banking, portfolio management, telecom, and healthcare, developing machine learning solutions for various business problems, including risk prediction, churn analysis, anomaly detection, task recommendations, and document information extraction. In addition, he is very active in competitive machine learning, currently holding the title of a three-time Kaggle Grandmaster and is well-known for his Kaggle Notebooks.
Read more about Gabriel Preda

Right arrow

Building a baseline model

As a result of our data analysis, we were able to identify some of the features with predictive value. We can now build a model by using this knowledge to select relevant features. We will start with a model that will use just two out of the many features we investigated. This is called a baseline model and it is used as a starting point for the incremental refinement of the solution.

For the baseline model, we chose a RandomForestClassifier model. The model is simple to use, gives good results with the default parameters, and can be interpreted easily, using feature importance.

Let’s begin with the following code block to implement the model. First, we import a few libraries that are needed to prepare the model. Then, we convert the categorical data to numerical. We need to do this since the model we chose deals with numbers only. The operation of converting the categorical feature values to numbers is called label encoding. Then, we split...

lock icon
The rest of the page is locked
Previous PageNext Page
You have been reading a chapter from
Developing Kaggle Notebooks
Published in: Dec 2023Publisher: PacktISBN-13: 9781805128519

Author (1)

author image
Gabriel Preda

Dr. Gabriel Preda is a Principal Data Scientist for Endava, a major software services company. He has worked on projects in various industries, including financial services, banking, portfolio management, telecom, and healthcare, developing machine learning solutions for various business problems, including risk prediction, churn analysis, anomaly detection, task recommendations, and document information extraction. In addition, he is very active in competitive machine learning, currently holding the title of a three-time Kaggle Grandmaster and is well-known for his Kaggle Notebooks.
Read more about Gabriel Preda