jel - Japanese Entity Linker - is Bi-encoder based entity linker for japanese.

Last update: Jan 06, 2023

Overview

jel: Japanese Entity Linker

jel - Japanese Entity Linker - is Bi-encoder based entity linker for japanese.

Usage

Currently, link and question methods are supported.

`el.link`

This returnes named entity and its candidate ones from Wikipedia titles.

from jel import EntityLinker
el = EntityLinker()

el.link('今日は東京都のマックにアップルを買いに行き、スティーブジョブスとドナルドに会い、堀田区に引っ越した。')
>> [
    {
        "text": "東京都",
        "label": "GPE",
        "span": [
            3,
            6
        ],
        "predicted_normalized_entities": [
            [
                "東京都庁",
                0.1084
            ],
            [
                "東京",
                0.0633
            ],
            [
                "国家地方警察東京都本部",
                0.0604
            ],
            [
                "東京都",
                0.0598
            ],
            ...
        ]
    },
    {
        "text": "アップル",
        "label": "ORG",
        "span": [
            11,
            15
        ],
        "predicted_normalized_entities": [
            [
                "アップル",
                0.2986
            ],
            [
                "アップル インコーポレイテッド",
                0.1792
            ],
            …
        ]
    }

`el.question`

This returnes candidate entity for any question from Wikipedia titles.

>>> linker.question('日本の総理大臣は？')
[('菅内閣', 0.05791765857101555), ('枢密院', 0.05592481946602986), ('党', 0.05430194711042564), ('総選挙', 0.052795400668513175)]

Setup

$ pip install jel
$ python -m spacy download ja_core_news_md

Run as API

$ uvicorn jel.api.server:app --reload --port 8000 --host 0.0.0.0 --log-level trace

Example

# link
$ curl localhost:8000/link -X POST -H "Content-Type: application/json" \
    -d '{"sentence": "日本の総理は菅総理だ。"}'

# question
$ curl localhost:8000/question -X POST -H "Content-Type: application/json" \
    -d '{"sentence": "日本で有名な総理は？"}

Test

$ python pytest

Notes

faiss==1.5.3 from pip causes error _swigfaiss.
To solve this, see this issue.

LICENSE

Apache 2.0 License.

CITATION

@INPROCEEDINGS{manabe2019chive,
    author    = {真鍋陽俊, 岡照晃, 海川祥毅, 髙岡一馬, 内田佳孝, 浅原正幸},
    title     = {複数粒度の分割結果に基づく日本語単語分散表現},
    booktitle = "言語処理学会第25回年次大会(NLP2019)",
    year      = "2019",
    pages     = "NLP2019-P8-5",
    publisher = "言語処理学会",
}

Japanese synonym library

chikkarpy chikkarpyはchikkarのPython版です。 chikkarpy is a Python version of chikkar. chikkarpy は Sudachi 同義語辞書を利用し、SudachiPyの出力に同義語展開を追加するために開発されたライブラリです。

48 Dec 14, 2022

AllenNLP integration for Shiba: Japanese CANINE model

Allennlp Integration for Shiba allennlp-shiab-model is a Python library that provides AllenNLP integration for shiba-model. SHIBA is an approximate re

12 Feb 16, 2022

Codes to pre-train Japanese T5 models

t5-japanese Codes to pre-train a T5 (Text-to-Text Transfer Transformer) model pre-trained on Japanese web texts. The model is available at https://hug

37 Dec 25, 2022

Auto translate textbox from Japanese to English or Indonesia

priconne-auto-translate Auto translate textbox from Japanese to English or Indonesia How to use Install python first, Anaconda is recommended Install

5 Aug 25, 2022

Code for evaluating Japanese pretrained models provided by NTT Ltd.

japanese-dialog-transformers 日本語の説明文はこちら This repository provides the information necessary to evaluate the Japanese Transformer Encoder-decoder dialo

216 Dec 22, 2022

Script to download some free japanese lessons in portuguse from NHK

Nihongo_nhk This is a script to download some free japanese lessons in portuguese from NHK. It can be executed by installing the packages with: pip in

2 Jan 6, 2022

An open collection of annotated voices in Japanese language

声庭 (Koniwa): オープンな日本語音声とアノテーションのコレクション Koniwa (声庭): An open collection of annotated voices in Japanese language 概要 Koniwa(声庭)は利用・修正・再配布が自由でオープンな音声とアノテ

32 Dec 14, 2022

Japanese Long-Unit-Word Tokenizer with RemBertTokenizerFast of Transformers

Japanese-LUW-Tokenizer Japanese Long-Unit-Word (国語研長単位) Tokenizer for Transformers based on 青空文庫 Basic Usage from transformers import RemBertToken

3 Dec 22, 2021

aMLP Transformer Model for Japanese

aMLP-japanese Japanese aMLP Pretrained Model aMLPとは、Liu, Daiらが提案する、Transformerモデルです。ざっくりというと、BERTの代わりに使えて、より性能の良いモデルです。詳しい解説は、こちらの記事などを参考にしてください。この

13 Aug 11, 2022

Comments

ModuleNotFoundError

Traceback (most recent call last):
  File "scripts/biencoder_training_check.py", line 1, in <module>
    from jel.biencoder.train import biencoder_training
ModuleNotFoundError: No module named 'jel'

opened by izuna385 1

Separate Estimation Model and DB

Because the inference model and knowledge base are currently loaded together, it takes 30 seconds to load the model. To prevent this, we will separate the DB into a separate container.

opened by izuna385 0

Releases(v0.1.1)

v0.1.1(May 29, 2021)

First release.
Source code(tar.gz)
Source code(zip)

jel - Japanese Entity Linker - is Bi-encoder based entity linker for japanese.

Related tags

Overview

jel: Japanese Entity Linker

Usage

el.link

el.question

Setup

Run as API

Example

Test

Notes

LICENSE

CITATION

You might also like...

Japanese synonym library

AllenNLP integration for Shiba: Japanese CANINE model

Codes to pre-train Japanese T5 models

Auto translate textbox from Japanese to English or Indonesia

Code for evaluating Japanese pretrained models provided by NTT Ltd.

Script to download some free japanese lessons in portuguse from NHK

An open collection of annotated voices in Japanese language

Japanese Long-Unit-Word Tokenizer with RemBertTokenizerFast of Transformers

aMLP Transformer Model for Japanese

Comments

ModuleNotFoundError

Separate Estimation Model and DB

Releases(v0.1.1)

v0.1.1(May 29, 2021)

Owner

izuna385

txtai: Build AI-powered semantic search applications in Go

Deeply Supervised, Layer-wise Prediction-aware (DSLP) Transformer for Non-autoregressive Neural Machine Translation

Implementation of COCO-LM, Correcting and Contrasting Text Sequences for Language Model Pretraining, in Pytorch

NLP and Text Generation Experiments in TensorFlow 2.x / 1.x

Predict an emoji that is associated with a text

An open collection of annotated voices in Japanese language

Quantifiers and Negations in RE Documents

xFormers is a modular and field agnostic library to flexibly generate transformer architectures by interoperable and optimized building blocks.

문장단위로 분절된 나무위키 데이터셋. Releases에서 다운로드 받거나, tfds-korean을 통해 다운로드 받으세요.

Quick insights from Zoom meeting transcripts using Graph + NLP

Perform sentiment analysis on textual data that people generally post on websites like social networks and movie review sites.

Japanese synonym library

Nested Named Entity Recognition for Chinese Biomedical Text

Create a semantic search engine with a neural network (i.e. BERT) whose knowledge base can be updated

File-based TF-IDF: Calculates keywords in a document, using a word corpus.

DLO8012: Natural Language Processing & CSL804: Computational Lab - II

Negative sampling for solving the unlabeled entity problem in NER. ICLR-2021 paper: Empirical Analysis of Unlabeled Entity Problem in Named Entity Recognition.

Knowledge Graph,Question Answering System，基于知识图谱和向量检索的医疗诊断问答系统

A2T: Towards Improving Adversarial Training of NLP Models (EMNLP 2021 Findings)

A collection of scripts to preprocess ASR datasets and finetune language-specific Wav2Vec2 XLSR models

`el.link`

`el.question`