Latent Dirichlet Allocation

The most common technique currently in use for topic modeling of text, and the one that the Facebook researchers used in their 2013 paper, is called Latent Dirichlet Allocation (LDA).

Tip

Many people wonder how to pronounce Dirichlet in English. The most common pronunciation I have heard is DEER-uh-shlay, and I have also heard DEER-uh-klay a few times.

LDA was first proposed for text topic extraction by David Blei, Andrew Ng, and Michael Jordan in a 2003 paper entitled simply Latent Direchlet Allocation, available from the Journal of Machine Learning Research at http://www.jmlr.org/papers/volume3/blei03a/blei03a.pdf. Blei also wrote a good follow-up article in 2012 for the Communications of the ACM about LDA and some ...

Get Mastering Data Mining with Python – Find patterns hidden in your data now with the O’Reilly learning platform.

O’Reilly members experience books, live events, courses curated by job role, and more from O’Reilly and nearly 200 top publishers.