Working through text analysis from the ground up: regular expressions, tokenisation, vectorisation, and sentiment.
Written in 2020 during my MSc Data Analytics at London Metropolitan University.
| Report | What it covers |
|---|---|
| Regular expressions | Patterns for pulling structured information out of unstructured text |
| Tokenization, stemming, lemmatization | Normalising text, and why stemming and lemmatisation are not interchangeable |
| Bag of words and TF-IDF | Vectorising documents with tidytext, sparse matrix representation, and weighting terms by inverse document frequency |
| Sentiment analysis | Lexicon-based sentiment over the Twitter US Airline Sentiment dataset |
TF-IDF is the idea worth taking away: raw counts rank every document by how chatty it is, and weighting by how rare a term is across the corpus is what makes the representation carry meaning.
Each report has its R Markdown source alongside it, plus Extracting twitter data.Rmd for collecting the sentiment corpus.
| File | Used by |
|---|---|
Tweets.csv |
Twitter US Airline Sentiment — the sentiment analysis |
Womens Clothing ECommerce Reviews.csv |
Product review corpus — the bag-of-words and TF-IDF work |
R · tidytext · dplyr · ggplot2 · R Markdown
Both corpora are committed, so the four reports run as-is.
Rscript install.R # tidytext, tm, textstem, SnowballC, wordcloud, widyr, stringr, ...Then knit any of the .Rmd files.
Extracting twitter data.Rmd is the exception: it calls the live Twitter API through rtweet and needs your own credentials. The sentiment analysis does not depend on it — it reads the committed Tweets.csv.
These are the classical foundations. For current retrieval work — hybrid semantic and keyword search, groundedness and hallucination measurement — see rag-eval-harness.