Masader

The first online catalogue for Arabic NLP datasets. This catalogue contains 200 datasets with more than 25 metadata annotations for each dataset. You can view the list of all datasets using the link of the webiste https://arbml.github.io/masader/

Title Masader: Metadata Sourcing for Arabic Text and Speech Data Resources
Authors Zaid Alyafeai, Maraim Masoud, Mustafa Ghaleb, Maged S. Al-shaibani
https://arxiv.org/abs/2110.06744

Abstract: The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata annotations that describe their attributes. Not to mention, the absence of a public catalogue that indexes all the publicly available datasets related to specific regions or languages. When we consider low-resource dialectical languages, for example, this issue becomes more prominent. In this paper we create \textit{Masader}, the largest public catalogue for Arabic NLP datasets, which consists of 200 datasets annotated with 25 attributes. Furthermore, We develop a metadata annotation strategy that could be extended to other languages. We also make remarks and highlight some issues about the current status of Arabic NLP datasets and suggest recommendations to address them.*

Metadata

No. dataset number
Name name of the dataset
Subsets subsets of the datasets
Link direct link to the dataset or instructions on how to download it
License license of the dataset
Year year of the publishing the dataset/paper
Language ar or multilingual
Dialect region ar-LEV: (Arabic(Levant)), country ar-EGY: (Arabic (Egypt)) or type ar-MSA: (Arabic (Modern Standard Arabic))
Domain social media, news articles, reviews, commentary, books, transcribed audio or other
Form text, audio or sign language
Collection style crawling, crawling and annotation (translation), crawling and annotation (other), machine translation, human translation, human curation or other
Description short statement describing the dataset
Volume the size of the dataset in numbers
Unit unit of the volume, could be tokens, sentences, documents, MB, GB, TB, hours or other
Provider company or university providing the dataset
Related Datasets any datasets that is related in terms of content to the dataset
Paper Title title of the paper
Paper Link direct link to the paper pdf
Script writing system either Arab, Latn, Arab-Latn or other
Tokenized whether the dataset is segmented using morphology: Yes or No
Host the host website for the data i.e GitHub
Access is the data free, upon-request or with-fee.
Cost cost of the data is with-fee.
Test split does the data contain test split: Yes or No
Tasks the tasks included in the dataset spearated by comma
Evaluation Set is the data included in the evaluation suit by BigScience
Venue Title the venue title i.e ACL
Citations the number of citations
Venue Type conference, workshop, journal or preprint
Venue Name full name of the venue i.e Associations of computation linguistics
authors list of the paper authors separated by comma
affiliations list of the paper authors' affiliations separated by comma
abstract abstract of the paper
Added by name of the person who added the entry
Notes any extra notes on the dataset

Contribution

If you want to add a new dataset feel free to update the sheet. Please follow the instructions there for adding the entry.

Citation

@misc{alyafeai2021masader,
      title={Masader: Metadata Sourcing for Arabic Text and Speech Data Resources}, 
      author={Zaid Alyafeai and Maraim Masoud and Mustafa Ghaleb and Maged S. Al-shaibani},
      year={2021},
      eprint={2110.06744},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

The first online catalogue for Arabic NLP datasets.

Related tags

Overview

Masader

Metadata

Contribution

Citation

Owner

ARBML

Beautiful visualizations of how language differs among document types.

PyTorch implementation of NATSpeech: A Non-Autoregressive Text-to-Speech Framework

easySpeech is an open-source Python wrapper for google speech to text API that doesn't require PyAudio(So you especially windows user don't have to deal with the errors while installing PyAudio) and also works with hugging face transformers

This program do translate english words to portuguese

SIGIR'22 paper: Axiomatically Regularized Pre-training for Ad hoc Search

Prompt tuning toolkit for GPT-2 and GPT-Neo

PyTorch implementation of the paper: Text is no more Enough! A Benchmark for Profile-based Spoken Language Understanding

Code for EMNLP20 paper: "ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training"

An IVR Chatbot which can exponentially reduce the burden of companies as well as can improve the consumer/end user experience.

Silero Models: pre-trained speech-to-text, text-to-speech models and benchmarks made embarrassingly simple

Wind Speed Prediction using LSTMs in PyTorch

PyTorch impelementations of BERT-based Spelling Error Correction Models.

Galois is an auto code completer for code editors (or any text editor) based on OpenAI GPT-2.

Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System

Task-based datasets, preprocessing, and evaluation for sequence models.

Auto_code_complete is a auto word-completetion program which allows you to customize it on your needs

Ukrainian TTS (text-to-speech) using Coqui TTS

A number of methods in order to perform Natural Language Processing on live data derived from Twitter

🤗 Transformers: State-of-the-art Natural Language Processing for Pytorch, TensorFlow, and JAX.

Yet Another Compiler Visualizer