Skip to content

Latest commit

 

History

33 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Feriji: A French-Zarma Parallel Corpus, Glossary & Translator

This repository contains Feriji, a work-in-progress French-Zarma parallel corpus curated by Habibatou Abdoulaye Alfari, Elysabhete Amadou Ibrahim, Christopher Homan, and Mamadou K. KEITA. Feriji is a collection of 61,085 aligned machine translation-ready French-Zarma lines curated from various sources. The corpus aims to contribute to the development of machine translation systems and linguistic studies between French and Zarma languages.

Dataset Description

  • Size: 61,085 sentences in Zarma and 42,789 in French.
  • Glossary: 4,062 words.

Dataset Statistics

French Zarma
Number of sentences 42,789 61,085
Glossary entries 4,062 4,062
Unique words 21,592 9,902

Usage

The dataset is intended for academic research and development of Machine Translation systems. You can test the Feriji Translator here.

Acknowledgements

We would like to thank our institutions, especially Ashesi University, and contributors for their support in creating this resource. The Computer Science department of Ashesi University provided financial and cloud resources support, which was crucial for this project.

Citations

If you use this dataset in your research, please cite it as follows:

@dataset{Feriji,
  author       = {Habibatou Abdoulaye Alfari, Elysabhete Amadou Ibrahim, Christopher Homan, and Mamadou K. KEITA},
  title        = {Feriji: A French-Zarma Parallel Corpus, Glossary & Translator},
  year         = 2023,
  publisher    = {GitHub},
  journal      = {GitHub repository},
  howpublished = {\url{https://github.com/27-GROUP/Feriji}}
}

About

A Zarma-French parallel corpus for Machine Translation

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages