Skip to content

Latest commit

 

History

32 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Natural language processing in R

Working through text analysis from the ground up: regular expressions, tokenisation, vectorisation, and sentiment.

Written in 2020 during my MSc Data Analytics at London Metropolitan University.

Topics

Report What it covers
Regular expressions Patterns for pulling structured information out of unstructured text
Tokenization, stemming, lemmatization Normalising text, and why stemming and lemmatisation are not interchangeable
Bag of words and TF-IDF Vectorising documents with tidytext, sparse matrix representation, and weighting terms by inverse document frequency
Sentiment analysis Lexicon-based sentiment over the Twitter US Airline Sentiment dataset

TF-IDF is the idea worth taking away: raw counts rank every document by how chatty it is, and weighting by how rare a term is across the corpus is what makes the representation carry meaning.

Source

Each report has its R Markdown source alongside it, plus Extracting twitter data.Rmd for collecting the sentiment corpus.

Data

File Used by
Tweets.csv Twitter US Airline Sentiment — the sentiment analysis
Womens Clothing ECommerce Reviews.csv Product review corpus — the bag-of-words and TF-IDF work

Stack

R · tidytext · dplyr · ggplot2 · R Markdown

Running it

Both corpora are committed, so the four reports run as-is.

Rscript install.R      # tidytext, tm, textstem, SnowballC, wordcloud, widyr, stringr, ...

Then knit any of the .Rmd files.

Extracting twitter data.Rmd is the exception: it calls the live Twitter API through rtweet and needs your own credentials. The sentiment analysis does not depend on it — it reads the committed Tweets.csv.


These are the classical foundations. For current retrieval work — hybrid semantic and keyword search, groundedness and hallucination measurement — see rag-eval-harness.

About

Text analysis in R: regular expressions, tokenisation, stemming, bag of words, TF-IDF and sentiment analysis.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages