Reader small image

You're reading from  Developing Kaggle Notebooks

Product typeBook
Published inDec 2023
Reading LevelIntermediate
PublisherPackt
ISBN-139781805128519
Edition1st Edition
Languages
Right arrow
Author (1)
Gabriel Preda
Gabriel Preda
author image
Gabriel Preda

Dr. Gabriel Preda is a Principal Data Scientist for Endava, a major software services company. He has worked on projects in various industries, including financial services, banking, portfolio management, telecom, and healthcare, developing machine learning solutions for various business problems, including risk prediction, churn analysis, anomaly detection, task recommendations, and document information extraction. In addition, he is very active in competitive machine learning, currently holding the title of a three-time Kaggle Grandmaster and is well-known for his Kaggle Notebooks.
Read more about Gabriel Preda

Right arrow

Analyzing the comments text

NLP is a field of AI that involves the use of computational techniques to enable computers to understand, interpret, transform, and even generate human language. NLP uses several techniques, algorithms, and models to process and analyze large datasets of text. Among these techniques, we can mention:

  • Tokenization: Breaks down text into smaller units, like words, parts of words, or characters
  • Lemmatization or stemming: Reduces the words to dictionary form or removes the last few characters to get to a common form (stem)
  • Part-of-Speech (POS) tagging: Assigns a grammatical category (for example, nouns, verbs, proper nouns, and adjectives) to each word in a sequence
  • Named Entity Recognition (NER): Identifies and classifies entities (for example, names of people, organizations, and places)
  • Word embeddings: Use a high-dimensional space to represent the words, a space in which the position of each word is determined by its relationship...
lock icon
The rest of the page is locked
Previous PageNext Page
You have been reading a chapter from
Developing Kaggle Notebooks
Published in: Dec 2023Publisher: PacktISBN-13: 9781805128519

Author (1)

author image
Gabriel Preda

Dr. Gabriel Preda is a Principal Data Scientist for Endava, a major software services company. He has worked on projects in various industries, including financial services, banking, portfolio management, telecom, and healthcare, developing machine learning solutions for various business problems, including risk prediction, churn analysis, anomaly detection, task recommendations, and document information extraction. In addition, he is very active in competitive machine learning, currently holding the title of a three-time Kaggle Grandmaster and is well-known for his Kaggle Notebooks.
Read more about Gabriel Preda