From the course: Natural Language Processing in Python

Unlock this course with a free trial

Join today to access over 26,400 courses taught by industry experts.

Solution: CountVectorizer

Solution: CountVectorizer

“

For this assignment, our first step is to vectorize the cleaned and normalized text. If you remember from the last assignment, we cleaned and normalized a column of text and we saved it in a data frame called df. So down here, let me add a few more cells. And let's first take a look at our clean text. You can see here we have this description clean column. And it's been cleaned and normalized using pandas and spaCy. So now we're ready to vectorize this column. I'm going to start by changing this to a two so I can still see this text column as I'm coding. And the first thing we have to do is import count vectorizer. I'm going to go to scikit-learn and then specifically the feature extraction section focused on text data. And here I'm going to import count vectorizer with a capital C and a capital V. Okay, now next we need to instantiate a new count vectorizer object. I'm going to say count vectorizer once again, and then put parentheses after, and I'm going to call that CV. And now…

Contents