Introduction

Fairseq(-py) is a sequence modeling toolkit that allows researchers and developers to train custom models for translation, summarization, language modeling and other text generation tasks.

What's New:

August 2019: WMT'19 models released
July 2019: fairseq relicensed under MIT license
July 2019: RoBERTa models and code released
June 2019: wav2vec models and code released

Features:

Fairseq provides reference implementations of various sequence-to-sequence models, including:

Convolutional Neural Networks (CNN)
LightConv and DynamicConv models
- Pay Less Attention with Lightweight and Dynamic Convolutions (Wu et al., 2019)
Long Short-Term Memory (LSTM) networks
- Effective Approaches to Attention-based Neural Machine Translation (Luong et al., 2015)
Transformer (self-attention) networks

Additionally:

multi-GPU (distributed) training on one machine or across multiple machines
fast generation on both CPU and GPU with multiple search algorithms implemented:
- beam search
- Diverse Beam Search (Vijayakumar et al., 2016)
- sampling (unconstrained, top-k and top-p/nucleus)
large mini-batch training even on a single GPU via delayed updates
mixed precision training (trains faster with less GPU memory on NVIDIA tensor cores)
extensible: easily register new models, criterions, tasks, optimizers and learning rate schedulers

We also provide pre-trained models for several benchmark translation and language modeling datasets.

Requirements and Installation

PyTorch version >= 1.1.0
Python version >= 3.5
For training new models, you'll also need an NVIDIA GPU and NCCL
For faster training install NVIDIA's apex library with the --cuda_ext option

To install fairseq:

pip install fairseq

On MacOS:

CFLAGS="-stdlib=libc++" pip install fairseq

If you use Docker make sure to increase the shared memory size either with --ipc=host or --shm-size as command line options to nvidia-docker run.

Installing from source

To install fairseq from source and develop locally:

git clone https://github.com/pytorch/fairseq
cd fairseq
pip install --editable .

Getting Started

The full documentation contains instructions for getting started, training new models and extending fairseq with new model types and tasks.

Pre-trained models and examples

We provide pre-trained models and pre-processed, binarized test sets for several tasks listed below, as well as example training and evaluation commands.

Translation: convolutional and transformer models are available
Language Modeling: convolutional and transformer models are available

We also have more detailed READMEs to reproduce results from specific papers:

Join the fairseq community

Facebook page: https://www.facebook.com/groups/fairseq.users
Google group: https://groups.google.com/forum/#!forum/fairseq-users

License

fairseq(-py) is MIT-licensed. The license applies to the pre-trained models as well.

Citation

Please cite as:

@inproceedings{ott2019fairseq,
  title = {fairseq: A Fast, Extensible Toolkit for Sequence Modeling},
  author = {Myle Ott and Sergey Edunov and Alexei Baevski and Angela Fan and Sam Gross and Nathan Ng and David Grangier and Michael Auli},
  booktitle = {Proceedings of NAACL-HLT 2019: Demonstrations},
  year = {2019},
}

Code for the project carried out fulfilling the course requirements for Fall 2021 NLP at NYU

Related tags

Overview

Introduction

What's New:

Features:

Requirements and Installation

Getting Started

Pre-trained models and examples

Join the fairseq community

License

Citation

Owner

Sai Himal Allu

Pre-Training with Whole Word Masking for Chinese BERT

The official implementation of VAENAR-TTS, a VAE based non-autoregressive TTS model.

Predict the spans of toxic posts that were responsible for the toxic label of the posts

Tensorflow Implementation of A Generative Flow for Text-to-Speech via Monotonic Alignment Search

Sentiment Analysis Project using Count Vectorizer and TF-IDF Vectorizer

Tools for curating biomedical training data for large-scale language modeling

Chatbot for the Chatango messaging platform

Natural Language Processing Specialization

NeuralQA: A Usable Library for Question Answering on Large Datasets with BERT

Sequence-to-Sequence learning using PyTorch

DVC-NLP-Simple-usecase

Official implementations for various pre-training models of ERNIE-family, covering topics of Language Understanding & Generation, Multimodal Understanding & Generation, and beyond.

PyTorch Language Model for 1-Billion Word (LM1B / GBW) Dataset

A program that uses real statistics to choose the best times to bet on BloxFlip's crash gamemode

Summarization, translation, sentiment-analysis, text-generation and more at blazing speed using a T5 version implemented in ONNX.

APEACH: Attacking Pejorative Expressions with Analysis on Crowd-generated Hate Speech Evaluation Datasets

A Word Level Transformer layer based on PyTorch and 🤗 Transformers.

ByT5: Towards a token-free future with pre-trained byte-to-byte models

Write Alphabet, Words and Sentences with your eyes.

Ecco is a python library for exploring and explaining Natural Language Processing models using interactive visualizations.