ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Last update: Mar 30, 2022

Related tags

Overview

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

ANEA is a tool to automatically annotate named entities in unlabeled text based on entity lists for the use as distant supervision.

Distant supervision allows obtaining labeled training corpora for low-resource settings where only limited hand-annotated data exists. However, to be used effectively, the distant supervision must be easy to gather. ANEA is a tool to automatically annotate named entities in texts based on entity lists. It spans the whole pipeline from obtaining the lists to analyzing the errors of the distant supervision. A tuning step allows the user to improve the automatic annotation with their linguistic insights without labelling or checking all tokens manually.

An example of the workflow can be seen in this video. For more details, take a look at our paper (accepted at PML4DC @ ICLR'21). For the additional material of the paper, please check the subdirectory additional of this repository.

Installation

ANEA should run on all major operating systems. We recommend the installation via conda or miniconda:

git clone https://github.com/uds-lsv/anea

conda create -n anea python=3.7
conda activate anea
pip install spacy==2.2.4 Flask==1.1.1 fuzzywuzzy==0.18.0

For tokenizationa and lemmatization, a spacy language pack needs to be installed. Run the following command with the corresponding language code, e.g. en for English. Check https://spacy.io/usage for supported languages

python -m spacy download en

Download the Wikidata JSON dump from https://dumps.wikimedia.org/wikidatawiki/entities/ and extract it to the instance directory (this may take a while).

Running

After the installation, you can run ANEA using the following commands on the command line

conda activate anea
./run.sh

Then open the browser and go to the address http://localhost:5000/ If you run it for the first time, you should configure ANEA at the Settings tab.

The ANEA (server) tool can run on a different machine than the browser of the user. It is just necessary that the user's computer can access the port 5000 on the machine that the ANEA server is running on (e.g. via ssh port forwarding or opening the correspoding port on the firewall).

Support for Other Languages

ANEA uses Spacy for language preprocessing (tokenization and lemmatization). It currently supports English, German, French, Spanish, Portuguese, Italian, Dutch, Greek, Norwegian Bokmål and Lithuanian. For Estonian, EstNLTK, version 1.6, is supported by ANEA. In that case, ANEA needs to be installed with Python 3.6.

Text can also be preprocessed using external tools and then uploaded as whitespace tokenized text or in the CoNLL format (one token per line).

Other external preprocessing libraries can be added directly to ANEA by implementing a new Tokenizer class in autom_labeling_library/preprocessing.py (you can take a look at EstnltkTokenizer as an example) and adding it to the Preprocessing class. If you encounter any issues, just contact us.

Citation

If you use this tool, please cite us:

@article{hedderich21ANEA,
  author    = {Michael A. Hedderich and
               Lukas Lange and
               Dietrich Klakow},
  title     = {{ANEA:} Distant Supervision for Low-Resource Named Entity Recognition},
  journal   = {CoRR},
  volume    = {abs/2102.13129},
  year      = {2021},
  url       = {https://arxiv.org/abs/2102.13129},
  archivePrefix = {arXiv},
  eprint    = {2102.13129},
}

Development, Support & License

If you encounter any issues or problems when using ANEA, feel free to raise an issue on Github or contact us directly (mhedderich [at] lsv.uni-saarland [dot] de). We welcome contributes from other developers.

ANEA is licensed under the Apache License 2.0.

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Related tags

Overview

ANEA: Distant Supervision for Low-Resource Named Entity Recognition

Installation

Running

Support for Other Languages

Citation

Development, Support & License

Owner

Saarland University Spoken Language Systems Group

AMTML-KD: Adaptive Multi-teacher Multi-level Knowledge Distillation

Blind Image Super-resolution with Elaborate Degradation Modeling on Noise and Kernel

MAVE: : A Product Dataset for Multi-source Attribute Value Extraction

TensorFlow implementation for Bayesian Modeling and Uncertainty Quantification for Learning to Optimize: What, Why, and How

PyTorch implementation of the Crafting Better Contrastive Views for Siamese Representation Learning

Spline is a tool that is capable of running locally as well as part of well known pipelines like Jenkins (Jenkinsfile), Travis CI (.travis.yml) or similar ones.

Reference code for the paper "Cross-Camera Convolutional Color Constancy" (ICCV 2021)

Cross-Image Region Mining with Region Prototypical Network for Weakly Supervised Segmentation

Official repository of DeMFI (arXiv.)

PASTRIE: A Corpus of Prepositions Annotated with Supersense Tags in Reddit International English

Turning SymPy expressions into PyTorch modules.

Sign-to-Speech for Sign Language Understanding: A case study of Nigerian Sign Language

OpenPCDet Toolbox for LiDAR-based 3D Object Detection.

Code for "Learning Graph Cellular Automata"

Objective of the repository is to learn and build machine learning models using Pytorch. 30DaysofML Using Pytorch

TiP-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

An image base contains 490 images for learning (400 cars and 90 boats), and another 21 images for testingAn image base contains 490 images for learning (400 cars and 90 boats), and another 21 images for testing

Attentive Implicit Representation Networks (AIR-Nets)

Repository for "Space-Time Correspondence as a Contrastive Random Walk" (NeurIPS 2020)

The fastai deep learning library