From the course: Natural Language Processing in Python

Unlock this course with a free trial

Join today to access over 26,400 courses taught by industry experts.

Assignment: CountVectorizer

Assignment: CountVectorizer

“

For your next assignment on Count Vectorizer, you have another message from Lexie Khan. And she says, hello, now that you've cleaned and normalized the book descriptions using pandas and spaCy, can you create a quick visualization to show the top 10 most common terms in the descriptions? Could you also share some of the less common terms that appear in multiple book descriptions? Thanks, Lexie. Your key objectives for this assignment are first to vectorize your cleaned and normalized text using count vectorizer. So you're going to take the output from the previous assignment and then use count vectorizer here to create a document term matrix. In this first step, you're just going to be using the default parameters. Then once you do that, the next step will be to modify those parameters to reduce the total number of columns. First, you're going to remove stop words, and then also you're going to set a minimum document frequency. From there, once you've updated your count vectorizer…

Contents