nlpcommon

nlpcommon, Python Text Tool. Python3开发。

Guide

Feature
Install
Usage
Dataset
Contact
Cite
Reference

Feature

nlpcommon is a python Open Source Toolkit for text classification. The goal is to implement text analysis algorithm, so as to achieve the use in the production environment.

nlpcommon has the characteristics of clear algorithm, high performance and customizable corpus.

Functions：

Classifier

Cluster

MiniBatchKmeans

While providing rich functions, nlpcommon internal modules adhere to low coupling, model adherence to inert loading, dictionary publication, and easy to use.

Install

Requirements and Installation

pip3 install nlpcommon

git clone https://github.com/shibing624/nlpcommon.git
cd nlpcommon
python3 setup.py install

Usage

data

Stopwrods

examples/base_demo.py:

import sys

sys.path.append('..')
from nlpcommon import stopwords

if __name__ == '__main__':
    print(len(stopwords), stopwords)

output:

2438 {'．', '大家', '孰知', '至于', './', '知道', '二话没说', '一何', '从宽', 'especially' ... }

Contact

Issue(建议)：
邮件我：xuming: [email protected]
微信我：加我微信号：xuming624, 进Python-NLP交流群，备注：姓名-公司名-NLP

Cite

如果你在研究中使用了nlpcommon，请按如下格式引用：

@software{nlpcommon,
  author = {Xu Ming},
  title = {nlpcommon: A Tool for Text NLP},
  year = {2021},
  url = {https://github.com/shibing624/nlpcommon},
}

License

授权协议为 The Apache License 2.0，可免费用做商业用途。请在产品说明中附加nlpcommon的链接和授权协议。

Contribute

项目代码还很粗糙，如果大家对代码有所改进，欢迎提交回本项目，在提交之前，注意以下两点：

在tests添加相应的单元测试
使用python setup.py test来运行所有单元测试，确保所有单测都是通过的

之后即可提交PR。

Reference

pytextclassifier

nlpcommon is a python Open Source Toolkit for text classification.

Related tags

Overview

nlpcommon

Feature

Classifier

Cluster

Install

Usage

data

Stopwrods

Contact

Cite

License

Contribute

Reference

Owner

xuming

glow-speak is a fast, local, neural text to speech system that uses eSpeak-ng as a text/phoneme front-end.

Develop open-source Python Arabic NLP libraries that the Arab world will easily use in all Natural Language Processing applications

L3Cube-MahaCorpus a Marathi monolingual data set scraped from different internet sources.

DomainWordsDict, Chinese words dict that contains more than 68 domains, which can be used as text classification、knowledge enhance task

MHtyper is an end-to-end pipeline for recognized the Forensic microhaplotypes in Nanopore sequencing data.

Rhyme with AI

Transformer Based Korean Sentence Spacing Corrector

AI-Broad-casting - AI Broad casting with python

Code of paper: A Recurrent Vision-and-Language BERT for Navigation

Code for the paper: Sequence-to-Sequence Learning with Latent Neural Grammars

Opal-lang - A WIP programming language based on Python

A Fast Sequence Transducer Implementation with PyTorch Bindings

fastNLP: A Modularized and Extensible NLP Framework. Currently still in incubation.

Crie tokens de autenticação íntegros e seguros com UToken.

CoSENT 比Sentence-BERT更有效的句向量方案

KR-FinBert And KR-FinBert-SC

Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Code for the ACL 2021 paper "Structural Guidance for Transformer Language Models"

The projects lets you extract glossary words and their definitions from a given piece of text automatically using NLP techniques

Simple telegram bot to convert files into direct download link.you can use telegram as a file server 🪁