Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Introduction

Fake News Classifier is a BERT-based model. The main hypothesis behind this project was that the bodies of news articles have textual characteristics (e.g., formality, sentimental information, etc.) which might be useful to identify whether an article is unreliable or not.

How to Use

Dataset

The data used in this repo were taken from the Kaggle. As I am not the owner of the data, I cannot store it in the repo. However, one can download the data by following the steps given below:

  1. Create a directory named input under the root directory.
  2. Create folders named kaggle1, kaggle2, kaggle3, and kaggle4 under the input directory.
  3. Download the csv files from the Kaggle links shared below and save them under the corresponding directories:

The final folder structure should look like this:

  • inputs/
    • kaggle1/
      • Fake.csv
      • True.csv
    • kaggle2/
      • Test.csv
    • kaggle3/
      • test.csv
      • train.csv
    • kaggle4/
      • WELFake_Dataset.csv

Installation

Before installing the required libraries, It is recommended to create a virtual environment.

The libraries required for the project are listed in the requirements.txt file. To download and install the necessary libraries,

pip install -r requirements.txt

Classification

The example code snippets for classification can be found inside main.py. Especially for multiple data, It would be useful to use the NewsDataset class implemented inside the dataset.py.

For server-side or single-query operations, please check the FakeNewsClassifier class given inside the classify.py. This file includes the production-level code that we used in the final form of our project.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages