[EMNLP 2021] Improving and Simplifying Pattern Exploiting Training

Last update: Dec 26, 2022

Related tags

Computer Vision ADAPET

Overview

ADAPET

This repository contains the official code for the paper: "Improving and Simplifying Pattern Exploiting Training".

The model improves and simplifies PET with a decoupled label objective and label-conditioned MLM objective.

Model

Decoupled Label Loss Label Conditioned Masked Language Modelling

Updates

[November 2021] You can run ADAPET on your own dataset now! See instructions here

Setup

Setup environment by running source bin/init.sh. This will

Download the FewGLUE and SuperGLUE datasets in data/fewglue/{task} and data/superglue/{task} respectively.
Install and setup environment with correct dependencies.

Training

First, create a config JSON file with the necessary hyperparameters. For reference, please see config/BoolQ.json.

Then, to train the model, run the following commands:

sh bin/setup.sh
sh bin/train.sh {config_file}

The output will be in the experiment directory exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/. Once the model has been trained, the following files can be found in the directory:

exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/
    |
    |__ best_model.pt
    |__ dev_scores.json
    |__ config.json
    |__ dev_logits.npy
    |__ src

To aid reproducibility, we provide the JSON files to replicate the paper's results at config/{task_name}.json.

Evaluation

To evaluate the model on the SuperGLUE dev set, run the following command:

sh bin/dev.sh exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/

The dev scores can be found in exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/dev_scores.json.

To evaluate the model on the SuperGLUE test set, run the following command.

sh bin/test.sh exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/

The generated predictions can be found in exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/test.json.

Train your own ADAPET

Setup your dataset in the data folder as

data/{dataset_name}/
    |
    |__ train.jsonl
    |__ val.jsonl
    |__ test.jsonl

Each jsonl file consists of lines of dictionaries. Each dictionaries should have the following format:

{
    "TEXT1": (insert text), 
    "TEXT2": (insert text), 
    "TEXT3": (insert text), 
    ..., 
    "TEXTN": (insert text), 
    "LBL": (insert label)
}

Run the experiment

python cli.py --data_dir data/{dataset_name} \
              --pattern '(INSERT PATTERN)' \
              --dict_verbalizer '{"lbl_1": "verbalizer_1", "lbl_2": "verbalizer_2"}'

Here, INSERT PATTERN consists of [TEXT1], [TEXT2], [TEXT3], ..., [LBL]. For example, if the new dataset had two text inputs and one label, a sample pattern would be [TEXT1] and [TEXT2] imply [LBL].

Fine-tuned Models

Our fine-tuned models can be found in this link.

To evaluate these fine-tuned models for different tasks, run the following command:

python src/run_pretrained.py -m {finetuned_model_dir}/{task_name} -c config/{task_name}.json -k pattern={best_pattern_for_task}

The scores can be found in exp_out/fewglue/{task_name}/albert-xxlarge-v2/{timestamp}/dev_scores.json. Note: The best_pattern_for_task can be found in Table 4 of the paper.

Contact

For any doubts or questions regarding the work, please contact Derek ([email protected]) or Rakesh ([email protected]). For any bug or issues with the code, feel free to open a GitHub issue or pull request.

Citation

Please cite us if ADAPET is useful in your work:

@inproceedings{tam2021improving,
          title={Improving and Simplifying Pattern Exploiting Training},
          author={Tam, Derek and Menon, Rakesh R and Bansal, Mohit and Srivastava, Shashank and Raffel, Colin},
          journal={Empirical Methods in Natural Language Processing (EMNLP)},
          year={2021}
}

[EMNLP 2021] Improving and Simplifying Pattern Exploiting Training

Related tags

Overview

ADAPET

Model

Updates

Setup

Training

Evaluation

Train your own ADAPET

Fine-tuned Models

Contact

Citation

Owner

Rakesh R Menon

Optical character recognition for Japanese text, with the main focus being Japanese manga

かの有名なあの東方二次創作ソング、「bad apple!」のMVをPythonでやってみたって話

Let's explore how we can extract text from forms

An Implementation of the seglink alogrithm in paper Detecting Oriented Text in Natural Images by Linking Segments

Generate a list of papers with publicly available source code in the daily arxiv

Deep learning based page layout analysis

Controlling the computer volume with your hands // OpenCV

天池2021"全球人工智能技术创新大赛"【赛道一】：医学影像报告异常检测 - 第三名解决方案

This can be use to convert text in a file to handwritten text.

TableBank: A Benchmark Dataset for Table Detection and Recognition

Virtualdragdrop - Virtual Drag and Drop Using OpenCV and Arduino

GDB python tool to pretty print and debug c++ xtensor containers

Maze generator and solver with python

Isearch (OSINT) 🔎 Face recognition reverse image search on Instagram profile feed photos.

Rubik's Cube in pygame with OpenGL

Captcha Recognition

In this project we will be using the live feed coming from the webcam to create a virtual mouse with complete functionalities.

Repository for playing the computer vision apps: People analytics on Raspberry Pi.

POT : Python Optimal Transport

A simple QR-Code Reader in Python