Machine Learning :: Text feature extraction (tf-idf) – Part I

18/09/201119/01/2020 by Christian S. Perone

Short introduction to Vector Space Model (VSM)

In information retrieval or text mining, the term frequency – inverse document frequency (also called tf-idf), is a well know method to evaluate how important is a word in a document. tf-idf are is a very interesting way to convert the textual representation of information into a Vector Space Model (VSM), or into sparse features, we’ll discuss more about it later, but first, let’s try to understand what is tf-idf and the VSM.

VSM has a very confusing past, see for example the paper The most influential paper Gerard Salton Never Wrote that explains the history behind the ghost cited paper which in fact never existed; in sum, VSM is an algebraic model representing textual information as a vector, the components of this vector could represent the importance of a term (tf–idf) or even the absence or presence (Bag of Words) of it in a document; it is important to note that the classical VSM proposed by Salton incorporates local and global parameters/information (in a sense that it uses both the isolated term being analyzed as well the entire collection of documents). VSM, interpreted in a lato sensu, is a space where text is represented as a vector of numbers instead of its original string textual representation; the VSM represents the features extracted from the document.

Let’s try to mathematically define the VSM and tf-idf together with concrete examples, for the concrete examples I’ll be using Python (as well the amazing scikits.learn Python module).

Going to the vector space

The first step in modeling the document into a vector space is to create a dictionary of terms present in documents. To do that, you can simple select all terms from the document and convert it to a dimension in the vector space, but we know that there are some kind of words (stop words) that are present in almost all documents, and what we’re doing is extracting important features from documents, features do identify them among other similar documents, so using terms like “the, is, at, on”, etc.. isn’t going to help us, so in the information extraction, we’ll just ignore them.

Let’s take the documents below to define our (stupid) document space:

Train Document Set:

d1: The sky is blue.

d2: The sun is bright.

Test Document Set:

d3: The sun in the sky is bright.

d4: We can see the shining sun, the bright sun.

Train Document Set: d1: The sky is blue. d2: The sun is bright. Test Document Set: d3: The sun in the sky is bright. d4: We can see the shining sun, the bright sun.

Train Document Set:

d1: The sky is blue.
d2: The sun is bright.

Test Document Set:

d3: The sun in the sky is bright.
d4: We can see the shining sun, the bright sun.

Now, what we have to do is to create a index vocabulary (dictionary) of the words of the train document set, using the documents $d1$ and $d2$ from the document set, we’ll have the following index vocabulary denoted as $\mathrm{E}(t)$ where the $t$ is the term:

$\mathrm{E}(t) = \begin{cases} 1, & \mbox{if } t\mbox{ is ``blue''} \\ 2, & \mbox{if } t\mbox{ is ``sun''} \\ 3, & \mbox{if } t\mbox{ is ``bright''} \\ 4, & \mbox{if } t\mbox{ is ``sky''} \\ \end{cases}$

Note that the terms like “is” and “the” were ignored as cited before. Now that we have an index vocabulary, we can convert the test document set into a vector space where each term of the vector is indexed as our index vocabulary, so the first term of the vector represents the “blue” term of our vocabulary, the second represents “sun” and so on. Now, we’re going to use the term-frequency to represent each term in our vector space; the term-frequency is nothing more than a measure of how many times the terms present in our vocabulary $\mathrm{E}(t)$ are present in the documents $d3$ or $d4$ , we define the term-frequency as a couting function:

$\mathrm{tf}(t,d) = \sum\limits_{x\in d} \mathrm{fr}(x, t)$

where the $\mathrm{fr}(x, t)$ is a simple function defined as:

$\mathrm{fr}(x,t) = \begin{cases} 1, & \mbox{if } x = t \\ 0, & \mbox{otherwise} \\ \end{cases}$

So, what the $tf(t,d)$ returns is how many times is the term $t$ is present in the document $d$ . An example of this, could be $tf(``sun'', d4) = 2$ since we have only two occurrences of the term “sun” in the document $d4$ . Now you understood how the term-frequency works, we can go on into the creation of the document vector, which is represented by:

$\displaystyle \vec{v_{d_n}} =(\mathrm{tf}(t_1,d_n), \mathrm{tf}(t_2,d_n), \mathrm{tf}(t_3,d_n), \ldots, \mathrm{tf}(t_n,d_n))$

Each dimension of the document vector is represented by the term of the vocabulary, for example, the $\mathrm{tf}(t_1,d_2)$ represents the frequency-term of the term 1 or $t_1$ (which is our “blue” term of the vocabulary) in the document $d_2$ .

Let’s now show a concrete example of how the documents $d_3$ and $d_4$ are represented as vectors:

$\vec{v_{d_3}} = (\mathrm{tf}(t_1,d_3), \mathrm{tf}(t_2,d_3), \mathrm{tf}(t_3,d_3), \ldots, \mathrm{tf}(t_n,d_3)) \\ \vec{v_{d_4}} = (\mathrm{tf}(t_1,d_4), \mathrm{tf}(t_2,d_4), \mathrm{tf}(t_3,d_4), \ldots, \mathrm{tf}(t_n,d_4))$

which evaluates to:

$\vec{v_{d_3}} = (0, 1, 1, 1) \\ \vec{v_{d_4}} = (0, 2, 1, 0)$

As you can see, since the documents $d_3$ and $d_4$ are:

d3: The sun in the sky is bright.

d4: We can see the shining sun, the bright sun.

d3: The sun in the sky is bright. d4: We can see the shining sun, the bright sun.

d3: The sun in the sky is bright.
d4: We can see the shining sun, the bright sun.

The resulting vector $\vec{v_{d_3}}$ shows that we have, in order, 0 occurrences of the term “blue”, 1 occurrence of the term “sun”, and so on. In the $\vec{v_{d_3}}$ , we have 0 occurences of the term “blue”, 2 occurrences of the term “sun”, etc.

But wait, since we have a collection of documents, now represented by vectors, we can represent them as a matrix with $|D| \times F$ shape, where $|D|$ is the cardinality of the document space, or how many documents we have and the $F$ is the number of features, in our case represented by the vocabulary size. An example of the matrix representation of the vectors described above is:

$M_{|D| \times F} = \begin{bmatrix} 0 & 1 & 1 & 1\\ 0 & 2 & 1 & 0 \end{bmatrix}$

As you may have noted, these matrices representing the term frequencies tend to be very sparse (with majority of terms zeroed), and that’s why you’ll see a common representation of these matrix as sparse matrices.

Python practice

Environment Used: Python v.2.7.2, Numpy 1.6.1, Scipy v.0.9.0, Sklearn (Scikits.learn) v.0.9.

Since we know the theory behind the term frequency and the vector space conversion, let’s show how easy is to do that using the amazing scikit.learn Python module.

Scikit.learn comes with lots of examples as well real-life interesting datasets you can use and also some helper functions to download 18k newsgroups posts for instance.

Since we already defined our small train/test dataset before, let’s use them to define the dataset in a way that scikit.learn can use:

train_set = ("The sky is blue.", "The sun is bright.")

test_set = ("The sun in the sky is bright.",

"We can see the shining sun, the bright sun.")

train_set = ("The sky is blue.", "The sun is bright.") test_set = ("The sun in the sky is bright.", "We can see the shining sun, the bright sun.")

train_set = ("The sky is blue.", "The sun is bright.")
test_set = ("The sun in the sky is bright.",
    "We can see the shining sun, the bright sun.")

In scikit.learn, what we have presented as the term-frequency, is called CountVectorizer, so we need to import it and create a news instance:

from sklearn.feature_extraction.text import CountVectorizer

vectorizer = CountVectorizer()

from sklearn.feature_extraction.text import CountVectorizer vectorizer = CountVectorizer()

from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer()

The CountVectorizer already uses as default “analyzer” called WordNGramAnalyzer, which is responsible to convert the text to lowercase, accents removal, token extraction, filter stop words, etc… you can see more information by printing the class information:

print vectorizer

CountVectorizer(analyzer__min_n=1,

analyzer__stop_words=set(['all', 'six', 'less', 'being', 'indeed', 'over', 'move', 'anyway', 'four', 'not', 'own', 'through', 'yourselves', (...)

print vectorizer CountVectorizer(analyzer__min_n=1, analyzer__stop_words=set(['all', 'six', 'less', 'being', 'indeed', 'over', 'move', 'anyway', 'four', 'not', 'own', 'through', 'yourselves', (...)

print vectorizer

CountVectorizer(analyzer__min_n=1,
analyzer__stop_words=set(['all', 'six', 'less', 'being', 'indeed', 'over', 'move', 'anyway', 'four', 'not', 'own', 'through', 'yourselves', (...)

Let’s create now the vocabulary index:

vectorizer.fit_transform(train_set)

print vectorizer.vocabulary

{'blue': 0, 'sun': 1, 'bright': 2, 'sky': 3}

vectorizer.fit_transform(train_set) print vectorizer.vocabulary {'blue': 0, 'sun': 1, 'bright': 2, 'sky': 3}

vectorizer.fit_transform(train_set)
print vectorizer.vocabulary
{'blue': 0, 'sun': 1, 'bright': 2, 'sky': 3}

See that the vocabulary created is the same as $E(t)$ (except because it is zero-indexed).

Let’s use the same vectorizer now to create the sparse matrix of our test_set documents:

smatrix = vectorizer.transform(test_set)

print smatrix

(0, 1) 1

(0, 2) 1

(0, 3) 1

(1, 1) 2

(1, 2) 1

smatrix = vectorizer.transform(test_set) print smatrix (0, 1) 1 (0, 2) 1 (0, 3) 1 (1, 1) 2 (1, 2) 1

smatrix = vectorizer.transform(test_set)

print smatrix

(0, 1)        1
(0, 2)        1
(0, 3)        1
(1, 1)        2
(1, 2)        1

Note that the sparse matrix created called smatrix is a Scipy sparse matrix with elements stored in a Coordinate format. But you can convert it into a dense format:

smatrix.todense()

matrix([[0, 1, 1, 1],

........[0, 2, 1, 0]], dtype=int64)

smatrix.todense() matrix([[0, 1, 1, 1], ........[0, 2, 1, 0]], dtype=int64)

smatrix.todense()

matrix([[0, 1, 1, 1],
........[0, 2, 1, 0]], dtype=int64)

Note that the sparse matrix created is the same matrix $M_{|D| \times F}$ we cited earlier in this post, which represents the two document vectors $\vec{v_{d_3}}$ and $\vec{v_{d_4}}$ .

We’ll see in the next post how we define the idf (inverse document frequency) instead of the simple term-frequency, as well how logarithmic scale is used to adjust the measurement of term frequencies according to its importance, and how we can use it to classify documents using some of the well-know machine learning approaches.

I hope you liked this post, and if you really liked, leave a comment so I’ll able to know if there are enough people interested in these series of posts in Machine Learning topics.

As promised, here is the second part of this tutorial series.

Cite this article as: Christian S. Perone, "Machine Learning :: Text feature extraction (tf-idf) – Part I," in Terra Incognita, 18/09/2011, https://blog.christianperone.com/2011/09/machine-learning-text-feature-extraction-tf-idf-part-i/.

References

The classic Vector Space Model

The most influential paper Gerard Salton never wrote

Wikipedia: tf-idf

Wikipedia: Vector space model

Scikits.learn Examples

Updates

21 Sep 11 – fixed some typos and the vector notation
22 Sep 11 – fixed import of sklearn according to the new 0.9 release and added the environment section
02 Oct 11 – fixed Latex math typos
18 Oct 11 – added link to the second part of the tutorial series
04 Mar 11 – Fixed formatting issues

113 thoughts on “Machine Learning :: Text feature extraction (tf-idf) – Part I”

justrreadrrr says:

18/09/2011 at 16:19

latex path not specified.
all over the text

Reply
1. Christian S. Perone says:
  
  18/09/2011 at 16:22
  
  I’m using the latex from wordpress.com service, for me it is working, maybe they service is down for a while =( thanks for reporting.
  
  Reply
Patrick Durusau says:

18/09/2011 at 16:49

The link to the The most influential paper Gerard Salton Never Wrote fails. Try the cached copy at CiteSeer: The most influential paper Gerard Salton Never Wrote.

Very enjoyable post! I have pointed to it from my blog: http://tm.durusau.net/?p=15199

Reply
1. Christian S. Perone says:
  
  18/09/2011 at 16:56
  
  Thank you Patrick, I’m glad you liked it. I updated the link with the CiteSeer copy.
  
  Reply
Helio Perroni Filho says:

18/09/2011 at 22:31

Very interesting read. Keep the good work.

Reply
Anand Jeyahar says:

19/09/2011 at 06:11

Thanks, the mix of actual examples with theory is very handy to see the theory in action and helps retain the theory better. Though in this particular post, i was a little disappointed as i felt it ended too soon. I would like more longer articles. But i guess longer articles turn off majority of the readers.

Reply
dvdgrs says:

19/09/2011 at 12:06

Very interesting blogpost, I’m sure up for more on the topic :)!

I recently had to handle VSM & TF-IDF in Python too, in a text-processing task of returning most similar strings of an input-string. I haven’t looked at scikits.learn, but it sure looks useful and straightforward.

I use Gensim (VSM for human beings: http://radimrehurek.com/gensim/) together with NLTK for preparing the data (aka word tokenizing, lowering words, and removing stopwords). I can highly recommend both libraries!

For some more (slightly out of date) details of my approach, see: http://graus.nu/blog/simple-keyword-extraction-in-python/

Thanks for the post, and looking forward to part II :).

Reply
1. Rahul says:
  
  11/09/2017 at 17:15
  
  Informative Blog Post, helped me a lot in understanding the concept. Please, keep the series going.
  
  Reply
Trey says:

19/09/2011 at 18:17

Thanks for posting this, would love to see more.

Reply
Johan says:

20/09/2011 at 04:57

Very well written and interesting!

Reply
Led says:

21/09/2011 at 11:22

Thanks for this, most interesting. I look forward to reading your future posts on the subject.

Reply
Pingback: Machine Learning :: Text feature extraction (tf-idf) – Part II | Pyevolve
Niu says:

24/11/2011 at 12:30

It is very useful and easy for start and is well organized. Thanks.

Reply
1. Christian S. Perone says:
  
  24/11/2011 at 13:31
  
  Thanks, I’m glad you liked it.
  
  Reply
baali says:

21/12/2011 at 07:12

Thanks a lot for this writeup. At times its really good to know what is cooking backstage behind all fancy and magical functions.

Reply
GEETHA r says:

04/01/2012 at 04:34

Thank you! It is very useful for me to learn about the vector space model. But I am having some doubts, please make me clear.
1. In my work I have added terms,Inverse document frequency. I want to achieve more accuracy. So I want to add some more, please suggest me…

Reply
alex says:

10/01/2012 at 15:52

Great post, I will certainly try this out.

I would be interested to see a similar detailed break down on using something like svmlight in conjunction with these techniques.

Thanks!

Reply
Jaques Grobler says:

26/03/2012 at 10:16

Hey again – my outputs are slighlty different to yours.. Think there may have been changes to the module on scikit-learn.

Will let you know what i find out.

Take care

Reply
Jaques Grobler says:

27/03/2012 at 06:54

Hello there,
so basically the class feature_selection.text.Vectorizer in Sklearn is now deprecated and replaced by feature_selection.text.TfidfVectorizer.

The whole module has been completely re-factored –
here’s the change-log from the Scikit-learn website:
http://scikit-learn.org/dev/whats_new.html

See under ‘API changes summary’ for what’s changed

Just thought I’d give you a headsup about this.
Enjoyed your post regardless!
Take care

Reply
1. Christian S. Perone says:
  
  27/03/2012 at 10:11
  
  Hello Jaques, great thanks for the feedback !
  
  Reply
Anita Mazur says:

20/04/2012 at 06:36

Thank you for your post. I am currently working on a way how to index documents, but with vocabulary terms taken from a thesaurus in SKOS format.
Your posts are interesting and very helpful to me.

Reply
1. Christian S. Perone says:
  
  20/04/2012 at 15:06
  
  Thanks for the feedback Anita, I’m glad you liked it.
  
  Reply
Zach says:

06/06/2012 at 16:24

Hey thanks for the very insightful post! I had no idea modules existed in Python that could do that for you ( I calculated it the hard way :/)

Just curious did you happen to know about using tf-idf weighting as a feature selection or text categorization method. I’ve been looking at many papers (most from China for some reason) but am finding numerous ways of approaching this question.

If there’s any advice or direction to steer me towards as far as additional resources, that would be greatly appreciated.

Reply
Andres Soto says:

03/08/2012 at 22:01

Hi
I am using python-2.7.3, numpy-1.6.2-win32-superpack-python2.7, scipy-0.11.0rc1-win32-superpack-python2.7, scikit-learn-0.11.win32-py2.7
I tried to repeat your steps but I couldn´t print the vectorizer.vocabulary (see below).
Any suggestions?
Regards
Andres Soto
>>> train_set = (“The sky is blue.”, “The sun is bright.”)
>>> test_set = (“The sun in the sky is bright.”,
“We can see the shining sun, the bright sun.”)
>>> from sklearn.feature_extraction.text import CountVectorizer
>>> vectorizer = CountVectorizer()
>>> print vectorizer
CountVectorizer(analyzer=word, binary=False, charset=utf-8,
charset_error=strict, dtype=, input=content,
lowercase=True, max_df=1.0, max_features=None, max_n=1, min_n=1,
preprocessor=None, stop_words=None, strip_accents=None,
token_pattern=bww+b, tokenizer=None, vocabulary=None)
>>> vectorizer.fit_transform(train_set)
<2×6 sparse matrix of type '’
with 8 stored elements in COOrdinate format>
>>> print vectorizer.vocabulary

Traceback (most recent call last):
File “”, line 1, in
print vectorizer.vocabulary
AttributeError: ‘CountVectorizer’ object has no attribute ‘vocabulary’
>>>

Reply
1. Koos vanderwilt says:
  
  17/09/2016 at 09:36
  
  use underscore:
  
  print vectorizer.vocabulary_
  
  Reply
Andres Soto says:

06/08/2012 at 14:32

I tried to fix the parameters of CountVectorizer (analyzer = WordNGramAnalyzer, vocabulary = dict) but it didn’t work. Therefore I decided to install sklearn 0.9 and it works, so we could say that everything is OK but I still would like to know what is wrong with version sklearn 0.11

Reply
1. Christian S. Perone says:
  
  06/08/2012 at 17:48
  
  Hello Andres, what I know is that this API has changed a lot on the sklearn 0.10/0.11, I heard some discussions about these changes but I can’t remember where right now.
  
  Reply
Gavin Igor says:

16/08/2012 at 23:41

Thanks for the great overview, looks like the part 2 link is broken. It would be great if you could fix it. Thank You.

Reply
1. Christian S. Perone says:
  
  17/08/2012 at 09:55
  
  Thanks for the feedback Gavin, the link is ok, it seems that the problem is sourceforge hosting that is throwing some errors.
  
  Reply
Gavin Igor says:

24/08/2012 at 17:48

I am using a mac and running 0.11 version but I got the following error I wonder how i change this according to the latest api

>> train_set
(‘The sky is blue.’, ‘The sun is bright.’)
>>> vectorizer.fit_transform(train_set)
<2×6 sparse matrix of type '’
with 8 stored elements in COOrdinate format>
>>> print vectorizer.vocabulary
Traceback (most recent call last):
File “”, line 1, in
AttributeError: ‘CountVectorizer’ object has no attribute ‘vocabulary’
>>> vocabulary

Reply
creativega says:

03/10/2012 at 05:18

Hello, Mr. Perone! Thank you very much, I’m newbie in TF-IDF and your posts have helped me a lot to understand it. Greetings from Japan^^

Reply
1. Christian S. Perone says:
  
  03/10/2012 at 16:21
  
  Great thanks for the feedback, I’m very glad you liked and that the post helped you !
  
  Reply
Mohit says:

20/10/2012 at 08:18

wonderful post… It helps me to understand VSM concept

Reply
Ryan says:

30/01/2013 at 04:08

Thanks for this awesome post! Eminently readable introduction to the topic.

Reply
Thomas says:

15/09/2013 at 07:42

very well written. Good example presented in a form that makes it easy to follow and understand. Curious now to read more…
Thank you for sharing your knowledge
Thomas, Germany

Reply
1. Christian S. Perone says:
  
  15/09/2013 at 12:47
  
  Thanks Thomas, I appreciate your feedback.
  
  Reply
UA says:

17/09/2013 at 16:59

Hi. I am having trouble understanding how to compute tf-idf weights for a text file I have which contains 300k lines of text. Each line is considered as a document. Example excerpt from the text file:

hacking hard
jeetu smart editor
shyamal vizualizr
setting demo hacks
vivek mans land
social routing guys minute discussions learn photography properly
naseer ahmed yahealer
sridhar vibhash yahoo search mashup
vaibhav chintan facebook friend folio
vaibhav chintan facebook friend folio
judgess comments
slickrnot

I’m pretty confused as to what I should do. Thanks

Reply
Nishant says:

08/10/2013 at 02:49

Thanks. It was helpful. Was looking for a good Python Vectorizer tutorial.

Reply
Igor says:

09/10/2013 at 15:59

That’s really interesting post, thanks a lot!

Reply
1. Christian S. Perone says:
  
  15/10/2013 at 12:56
  
  Thanks for the feedback Igor.
  
  Reply
Navdeep says:

08/11/2013 at 08:30

Really helpful post

Reply
Haneef says:

06/12/2013 at 07:02

Really a very good effort in explaining in such a simple way.

Reply
Williame Rocha says:

20/01/2014 at 18:19

it is a very good text.

Thanks for explanation.

Reply
need info says:

10/03/2014 at 15:25

Its giving idf vector as zero if test is same as train. why???

Reply
chamso says:

10/04/2014 at 18:24

thank you very much i encourage you to continue this is very helpful post <3

Reply
Matthieu says:

05/05/2014 at 11:49

Thanks
A well written clear explanation

Definitely a reference when taking the first steps in text mining.

Reply
1. Christian S. Perone says:
  
  06/05/2014 at 09:56
  
  Thank you Matthieu !
  
  Reply
purna says:

27/06/2014 at 01:26

Good post with example

Reply
Hannah says:

30/07/2014 at 11:21

thank you very much. this post helped me alot. very well written and clear explanation.
Hannah

Reply
Tim I. says:

02/09/2014 at 08:04

Awesome stuff. I really appreciate the simplicity and clarity of the information. A great, great help.

Reply
1. Christian S. Perone says:
  
  02/09/2014 at 10:29
  
  Great thanks Tim, I’m glad you liked it.
  
  Reply
Ganesh says:

13/11/2014 at 13:48

Thanks for the great post. You have explained it in simple words, so that a novice like me can understand. Will move on to read the next part!

Reply
Usher says:

25/03/2015 at 16:39

Solution to question of Andres and Gavin:

>>> print vectorizer.vocabulary_

(with underscore at the end in new versions of scikit!)
will output:

{u’blue’: 0, u’bright’: 1, u’sun’: 4, u’is’: 2, u’sky’: 3, u’the’: 5}

Reply
1. Sultan says:
  
  10/11/2015 at 23:55
  
  Hey brother,
  
  Do you know exactly what is the difference between (vectorizer.vocabulary_) and (vectorizer.get_feature_names() )?
  
  Reply
sanja7s says:

01/06/2015 at 09:17

Thanks, a great one and useful!

Reply
jpf says:

24/06/2015 at 20:16

” … you can simple select …” -> “>>> you can simply select …”

Reply
Narendra Rawat says:

04/09/2015 at 09:25

Very nice post!
You made tf-idf look really interesting. I am looking forward to some more such posts.

Reply
sandanaanil says:

19/09/2015 at 12:31

Hi Christian,

Thank u for sharing. Seeing more updates from you

Reply
Anil says:

21/10/2015 at 07:08

This is Great!
Really helpful for starters like me!
Thanks & Keep up the Good Work!

Cheers!

Reply
Sultan says:

10/11/2015 at 23:53

Could you tell me please what is the difference between feature names and vocabulary_ ?

I printed them both after vectorizing,, they seem having different words??

Also, I need to print out the most informative words in each class, could you suggest me a way please?

thanks

Reply
Sambit says:

23/11/2015 at 18:00

Thank you. This was a very informative post.

Reply
farhan khan says:

25/12/2015 at 10:28

nice one….helped a lot…!!

Reply
Shashi says:

30/12/2015 at 11:10

Thanks…very nicely explained. I feel I could understand the concept and now I will experiment. It will be very helpful in my work

Reply
Pingback: An Attempt at Quantifying Changes to Genre Medium
Pingback: Inverse Document Frequency Issue: Unexpected Results »
Pinki talukdar says:

04/03/2016 at 22:46

Thank u for u r post..it is very helpful.if possible can you tell in matlab how it will work

Reply
Eduardo says:

01/05/2016 at 20:25

Great and simple. Thanks,very helpful!

Reply
Hazem says:

27/07/2016 at 08:37

Thanks .. it was very inspiring tutorial for me

Reply
Swat says:

11/10/2016 at 23:57

Well written blog..I really loved it..:) Thank you..

Reply
Reza says:

07/11/2016 at 20:36

Hi, you have a nice blog.
You mentioned by text mining, stop words like “the, is, at, on”, etc.. isn’t going to help us”. This is partly true, for example in case of analyzing webpages, you want to ignore the Advertisements on a webpage, one good thing is to ignore the those sentences that do no have stop words. compared to normal sentences which do have these words.

Reply
Devang says:

11/11/2016 at 04:53

Very helpful. Like your writing style.

Reply
pan says:

21/11/2016 at 17:54

I tried out this, did not quite get the expected result:
Please see below:
train_set = (“The sky is blue.”, “The sun is bright.”)
test_set = (“The sun in the sky is bright.”, “We can see the shining sun, the bright sun.”)
from sklearn.feature_extraction.text import CountVectorizer
vectorizer = CountVectorizer()
stopwords = nltk.corpus.stopwords.words(‘english’)
vectorizer.stop_words = stopwords
print vectorizer
vectorizer.fit_transform(train_set)
print vectorizer.vocabulary

And I get ‘None’

Reply
1. Christian Hubbs says:
  
  01/01/2017 at 11:24
  
  Instead, try:
  
  print(vectorizer.vocabulary_)
  
  Reply
  1. Anonymous says:
    
    06/02/2017 at 13:40
    
    great, it works. Thanks!
    
    Reply
ankur says:

02/12/2016 at 22:09

very helpful post!

Reply
Kathereine says:

07/12/2016 at 01:58

Great post! Thanks.

Reply
Priya says:

15/12/2016 at 04:23

Great tutorial!! Thankz

Reply
Alex says:

04/01/2017 at 14:32

Really nice tutorial. Very helpful to get some context additional to the official skikit-learn tutorial and user guide. Thanks.

Reply
yizhen says:

21/01/2017 at 23:45

It cool work

Reply
Kiran Reddy says:

13/02/2017 at 07:58

Thank you for helping in understanding.

Reply
Michael says:

06/03/2017 at 14:16

Thank you so much, Christian! This post helped me a lot!

Reply
isco sarita says:

19/03/2017 at 16:17

i need a java program for indexing a set of files by computing tf and idf please help me

Reply
Krishna Priya says:

22/03/2017 at 07:17

Thanks a lot……..Post really helped me a lot!!!!!!!!!

Reply
Kalai says:

19/04/2017 at 04:24

The article is helpful. Thanks.

Reply
Rens says:

24/04/2017 at 07:08

As a PhD candidate in sociology who is diving into the world of machine learning, this post was also very helpful for me. Thanks!

Reply
akhil says:

03/05/2017 at 19:48

This is very helpful … it gave me thorough understanding of the concepts …

Reply
C17 says:

17/05/2017 at 13:07

Great article!! Read much but this belongs definitely to the “good stuff”!!!

Reply
Rajat says:

12/06/2017 at 14:01

It was cool man

Reply
nen says:

18/06/2017 at 22:02

Hello there!
I have a question regarding natural language processing. There are two terms in this field ‘feature extraction’ and ‘feature selection’. I don’t exactly understand the difference between them and whether we only use one of them or is it possible to use both for text classification?
My second question is whether ‘tf’ and ‘tfidf’ are considered feature extraction methods in NLP?

Reply
Rosangela Oliveira says:

06/07/2017 at 15:47

TypeError: __init__() got an unexpected keyword argument ‘analyzer__stop_words’

Reply
1. Rosangela Oliveira says:
  
  06/07/2017 at 15:48
  
  Could you hel with this error?
  
  Reply
Pedram says:

19/07/2017 at 16:39

It is an interesting article indeed. Personally, I know everything that has been mentioned in this post and I did all of them before, but sometimes it is worth spending little time to review some stuff that you already know. Keep up the good work!

Reply
Afsan Gujarati says:

19/10/2017 at 19:49

Really appreciate you taking the time to write this post.
Pretty detailed and well explained.
Appreciated.

Reply
Itxel Zavala says:

24/10/2017 at 19:33

I tried with print(vectorizer.vocabulary_) and it’s works, but my output is:
{‘the’: 5, ‘sky’: 3, ‘is’: 2, ‘blue’: 0, ‘sun’: 4, ‘bright’: 1}

Do you know why doesn’t ignored “is’ and “the” ?

Reply
1. ldag says:
  
  17/11/2017 at 04:22
  
  I was also facing the same issue but got solution. You can initialize the vectorizer as follow:
  vectorizer = CountVectorizer(stop_words=”english”)
  
  above will escape all english stop words.
  cheers..
  
  Reply
ldag says:

17/11/2017 at 04:45

Nice work Christian…

Reply
Manoj says:

24/11/2017 at 18:03

This is by far the best article on TF-IDF and Vector spaces. Thank you for posting such a helpful article. Please keep writing more articles on Machine learning basics and concepts. Thank you!

Reply
ali mohammad says:

01/12/2017 at 06:22

very nice explanation allah bless you

Reply
Manoj says:

29/12/2017 at 04:33

Good tutorial. It explains things in a simple and clear way to new bees like me… Thanks for sharing it…

Reply
Ameera says:

12/02/2018 at 07:53

Thank you
You made it so easy to understand!

Reply
Anonymous says:

12/02/2018 at 10:46

excellent article…very informative and way of explanation is very good.
Thank you so much!!!

Reply
Bren says:

06/03/2018 at 03:37

This post helped a lot….waiting for next article….

Reply
Harish says:

03/04/2018 at 00:05

print vectorizer.vocabulary_ (_) is missing.

Reply
Harish says:

03/04/2018 at 00:16

CountVectorizer() method for stopword removal does not seem to be clear, please complete the function with correct syntax

Reply
Aswitha Visvesvaran says:

22/05/2018 at 12:29

Great post..Very clean explanation of the concept. Python codes are an added bonus. Thank you

Reply
Márcio Jesus says:

05/06/2018 at 13:55

Thanks Christian, very good. Too bad it took me to start studying about this.

Reply
amanda says:

10/08/2018 at 11:05

hi! i’m currently make a search engine for journals with tfidf method for my undergraduate. but my professor said that my method is too old. could you recommend some new method in this past 5 years? or maybe method to optimize the tfidf? additional research paper about the method will be great. thankyou very much!

Reply
Insect Spring says:

03/10/2018 at 17:21

Very interesting and succinct read! I am ramping on to ML and it really helped. Going to read your other posts too.

Reply
ana says:

10/02/2019 at 02:33

this post is soo great keep the good work

Reply
Mehtab says:

12/04/2019 at 10:25

Very well explained with examples step by step. Easy to understand and really very helpful. Thanks a lot for such efforts. Please post further also.

Reply
Gopi Prashanth says:

15/07/2019 at 06:49

Thanks for detailed explanation

Reply
Bivas says:

04/01/2020 at 13:44

Detailed and simplified explanation .
Thank you so much !!! Keep up the good work

Reply
Enameguolor says:

10/07/2022 at 12:11

Hello Christian,
I want to thank you for the great work you are doing.Please i was give a project on similarity scoring system that uses modern plagiarism checking technology and returns a similarity score for the student submissions based on how similar the answers provided by 2 students.My question is that,i have started learning python already,which NLP language do i learn to achieve this.I already have c# and asp.net experience.Please help me please.

Reply