Automatic Idiomatic Expression Detection

Last update: Jun 09, 2022

Related tags

Deep Learning DISC

Overview

IDentifier of Idiomatic Expressions via Semantic Compatibility (DISC)

An Idiomatic identifier that detects the presence and span of idiomatic expression in a given sentence.

Table of Contents

About The Project
- Built With
Getting Started
- Prerequisites
- Installation
Usage

Configuration
Demo
Data Processing
Training and Testing

License
Contact
Acknowledgements

About The Project

This project is a supervised idiomatic expression identification method. Given a sentence that contains a potentially idiomatic expression (PIE), the model identifies the span of the PIE if it is indeed used in an idiomatic sense, otherwise, the model does not identify the PIE. The identification is done via checking the smemantic compatibility. More details will be updated here (Detail description, figures, etc.).

The paper will appear in TACL.

Built With

This model is heavily relying the resources/libraries list as following:

Getting Started

The implementation here includes processed data created for MAGPIE random-split dataset. The model checkpoint that trained with MAGPIE random-split is also provided.

Prerequisites

All the dependencies for this project is listed in requirements.txt. You can install them via a standard command:

pip install -r requirements.txt

It is highly recommanded to start a conda environment with PyTorch properly installed based on your hardward before install the other requirements.

Checkpoint

To run the model with a pre-trained checkpoint, please first create a ./checkpoints folder at root. Then, please download the checkpoint from Google Drive via this Link. Please put the checkpoint in the ./checkpoints folder.

Usage

Configuration

Before running the demo or experiments (training or testing), please see the config.py which sets the configuration of the model. Some parameters there, such as MODE needs to be set appropriately for the model to run correctly. Please see comments for more details.

Demo

To start, please go through the examples provided in demo.ipynb. In there, we process a given input sentence into the model input data and then run model inference to extract the idiomatic expression (if present) from the input sentence (visualized).

Data processing

To process a dataset (such as MAGPIE) for model training and testing, please refer to ./data_processing/MAGPIE/read_comp_data_processing.ipynb. It takes a dataset with sententences and their PIE lcoations as input and generate all the necessary files for model training and inference.

Training and Testing

For training and testing, please refer to train.py and test.py. Note that test.py is used to produce evaluation scores as shown in the paper. inference.py is used to produce prediction for sentences.

License

Distributed under the MIT License. See LICENSE for more information.

Contact

Ziheng Zeng - [email protected]

Project Link: https://github.com/your_username/repo_name

Acknowledgements

[TODO]:

Add the following in README:

Method detail descrption
Method figure
Demo walkthrough
Data processing tips and instructions Add requirements.txt

Automatic Idiomatic Expression Detection

Related tags

Overview

IDentifier of Idiomatic Expressions via Semantic Compatibility (DISC)

About The Project

Built With

Getting Started

Prerequisites

Checkpoint

Usage

Configuration

Demo

Data processing

Training and Testing

License

Contact

Acknowledgements

[TODO]:

Owner

This repository builds a basic vision transformer from scratch so that one beginner can understand the theory of vision transformer.

Official implementation of "Motif-based Graph Self-Supervised Learning forMolecular Property Prediction"

Unofficial implementation of the Involution operation from CVPR 2021

DeepLabv3+：Encoder-Decoder with Atrous Separable Convolution语义分割模型在tensorflow2当中的实现

M2MRF: Many-to-Many Reassembly of Features for Tiny Lesion Segmentation in Fundus Images

Implements pytorch code for the Accelerated SGD algorithm.

Multi-modal Vision Transformers Excel at Class-agnostic Object Detection

(NeurIPS '21 Spotlight) IQ-Learn: Inverse Q-Learning for Imitation

Awesome Remote Sensing Toolkit based on PaddlePaddle.

A Fast and Accurate One-Stage Approach to Visual Grounding, ICCV 2019 (Oral)

This repository contains several image-to-image translation models, whcih were tested for RGB to NIR image generation. The models are Pix2Pix, Pix2PixHD, CycleGAN and PointWise.

Interactive Terraform visualization. State and configuration explorer.

This repo provides code for QB-Norm (Cross Modal Retrieval with Querybank Normalisation)

Trax — Deep Learning with Clear Code and Speed

Pytorch implementation for "Open Compound Domain Adaptation" (CVPR 2020 ORAL)

Official repository for "Intriguing Properties of Vision Transformers" (2021)

Code for Deep Single-image Portrait Image Relighting

K-Nearest Neighbor in Pytorch

"Domain Adaptive Semantic Segmentation without Source Data" (ACM MM 2021)

[NeurIPS'21] "AugMax: Adversarial Composition of Random Augmentations for Robust Training" by Haotao Wang, Chaowei Xiao, Jean Kossaifi, Zhiding Yu, Animashree Anandkumar, and Zhangyang Wang.