21.19 Exploring the Document Term Matrix
We can obtain the term frequencies as a vector by converting the document term matrix into a matrix and summing the column counts:
## [1] 6508
By ordering the frequencies we can list the most frequent terms and the least frequent terms:
## acnntex dmitl microsystem ventur adra attributeori
## 1 1 1 1 1 1
Notice these terms appear just once and are probably not really terms that are of interest to us. Indeed they are likely to be spurious terms introduced through the translation of the original document from PDF to text.
## can dataset pattern use mine data
## 709 776 887 1366 1446 3101
These terms are much more likely to be of interest to us. Not surprising, given the choice of documents in the corpus, the most frequent terms are: data, mine, use, pattern, dataset, can.
If you find this curated material useful then you can consider a donation to support it's ongoing availability and give you access to the PDF version of this book. The material has been scoped up by Generative AI without permission or any kind of recompense so do consider a donation if you can afford it. Unlike Generative AI your access to this materials is freely given. Desktop Survival Guides include Data Science, GNU/Linux, and MLHub. Books available on Amazon include Data Mining with Rattle and Essentials of Data Science. Togaware has a 30 year tradition of making popular open source software which includes sold privacy preserving productivity apps, rattle, wajig, and mlhub. Hosted by Togaware, a pioneer of free and open source software since 1984. Copyright © 1995-2022 Graham.Williams@togaware.com Creative Commons Attribution-ShareAlike 4.0