nlpcommon

nlpcommon, Python Text Tool. Python3开发。

Guide

Feature
Install
Usage
Dataset
Contact
Cite
Reference

Feature

nlpcommon is a python Open Source Toolkit for text classification. The goal is to implement text analysis algorithm, so as to achieve the use in the production environment.

nlpcommon has the characteristics of clear algorithm, high performance and customizable corpus.

Functions：

Classifier

Cluster

MiniBatchKmeans

While providing rich functions, nlpcommon internal modules adhere to low coupling, model adherence to inert loading, dictionary publication, and easy to use.

Install

Requirements and Installation

pip3 install nlpcommon

git clone https://github.com/shibing624/nlpcommon.git
cd nlpcommon
python3 setup.py install

Usage

data

Stopwrods

examples/base_demo.py:

import sys

sys.path.append('..')
from nlpcommon import stopwords

if __name__ == '__main__':
    print(len(stopwords), stopwords)

output:

2438 {'．', '大家', '孰知', '至于', './', '知道', '二话没说', '一何', '从宽', 'especially' ... }

Contact

Issue(建议)：
邮件我：xuming: [email protected]
微信我：加我微信号：xuming624, 进Python-NLP交流群，备注：姓名-公司名-NLP

Cite

如果你在研究中使用了nlpcommon，请按如下格式引用：

@software{nlpcommon,
  author = {Xu Ming},
  title = {nlpcommon: A Tool for Text NLP},
  year = {2021},
  url = {https://github.com/shibing624/nlpcommon},
}

License

授权协议为 The Apache License 2.0，可免费用做商业用途。请在产品说明中附加nlpcommon的链接和授权协议。

Contribute

项目代码还很粗糙，如果大家对代码有所改进，欢迎提交回本项目，在提交之前，注意以下两点：

在tests添加相应的单元测试
使用python setup.py test来运行所有单元测试，确保所有单测都是通过的

之后即可提交PR。

Reference

pytextclassifier

nlpcommon is a python Open Source Toolkit for text classification.

Related tags

Overview

nlpcommon

Feature

Classifier

Cluster

Install

Usage

data

Stopwrods

Contact

Cite

License

Contribute

Reference

Owner

xuming

An open source library for deep learning end-to-end dialog systems and chatbots.

PyWorld3 is a Python implementation of the World3 model

pyupbit 라이브러리를 활용하여 upbit에서 비트코인을 자동매매하는 코드입니다. 조코딩 유튜브 채널에서 자세한 강의 영상을 보실 수 있습니다.

Code for hyperboloid embeddings for knowledge graph entities

Code for paper: An Effective, Robust and Fairness-awareHate Speech Detection Framework

ConvBERT-Prod

使用pytorch+transformers复现了SimCSE论文中的有监督训练和无监督训练方法

Official implementation of MLP Singer: Towards Rapid Parallel Korean Singing Voice Synthesis

ACL'2021: Learning Dense Representations of Phrases at Scale

A cross platform OCR Library based on PaddleOCR & OnnxRuntime

Code for our paper "Transfer Learning for Sequence Generation: from Single-source to Multi-source" in ACL 2021.

To classify the News into Real/Fake using Features from the Text Content of the article

Text vectorization tool to outperform TFIDF for classification tasks

CCF BDCI 2020 房产行业聊天问答匹配赛道 A榜47/2985

🌐 Translation microservice powered by AI

Fake Shakespearean Text Generator

A Multilingual Latent Dirichlet Allocation (LDA) Pipeline with Stop Words Removal, n-gram features, and Inverse Stemming, in Python.

MMDA - multimodal document analysis

Code for EMNLP'21 paper "Types of Out-of-Distribution Texts and How to Detect Them"

A retro text-to-speech bot for Discord