Package for controllable summarization

Last update: Dec 07, 2022

Related tags

Overview

summarizers

summarizers is package for controllable summarization based CTRLsum.
currently, we only supports English. It doesn't work in other languages.

Installation

pip install summarizers

Usage

1. Create Summarizers

First at all, create summarizers obejct to summarize your own article.

>>> from summarizers import Summarizers
>>> summ = Summarizers()

You can select type of source article between [normal, paper, patent].
If you don't input any parameter, default type is normal.

>>> from summarizers import Summarizers
>>> summ = Summarizers('normal')  # <-- default.
>>> summ = Summarizers('paper')
>>> summ = Summarizers('patent')

If you want GPU acceleration, set param device='cuda'.

>>> from summarizers import Summarizers
>>> summ = Summarizers('normal', device='cuda')

2. Basic Summarization

If you inputted source article, basic summariztion is conducted.

>>> contents = """
Tunip is the Octonauts' head cook and gardener. 
He is a Vegimal, a half-animal, half-vegetable creature capable of breathing on land as well as underwater. 
Tunip is very childish and innocent, always wanting to help the Octonauts in any way he can. 
He is the smallest main character in the Octonauts crew.
"""

>>> summ(contents)
'Tunip is a Vegimal, a half-animal, half-vegetable creature'

3. Query focused Summarization

If you want to input query together, Query focused summarization conducted.

>>> summ(contents, query="main character of Octonauts")
'Tunip is the smallest main character in the Octonauts crew.'

3. Abstractive QA (Auto Question Detection)

If you inputted question as query, Abstractive QA is conducted.

>>> summ(contents, query="What is Vegimal?")
'Half-animal, half-vegetable'

You can turn off this feature by setting param question_detection=False.

>>> summ(contents, query="SOME_QUERY", question_detection=False)

4. Prompt based Summarization

You can generate summary that begins with some sequence using param prompt.
It works like GPT-3's Prompt based generation. (but It doesn't work very well.)

>>> summ(contents, prompt="Q:Who is Tunip? A:")
"Q:Who is Tunip? A: Tunip is the Octonauts' head"

5. Query focused Summarization with Prompt

You can also input both query and prompt.
In this case, a query focus summary is generated that starts with a prompt.

>>> summ(contents, query="personality of Tunip", prompt="Tunip is very")
"Tunip is very childish and innocent, always wanting to help the Octonauts."

6. Options for Decoding Strategy

For generative models, decoding strategy is very important.
summarizers support variety of options for decoding strategy.

>>> summ(
...     contents=contents,
...     num_beams=10,
...     top_k=30,
...     top_p=0.85,
...     no_repeat_ngram_size=3,                  
... )

License

Copyright 2021 Hyunwoong Ko.

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.

Package for controllable summarization

Related tags

Overview

summarizers

Installation

Usage

1. Create Summarizers

2. Basic Summarization

3. Query focused Summarization

3. Abstractive QA (Auto Question Detection)

4. Prompt based Summarization

5. Query focused Summarization with Prompt

6. Options for Decoding Strategy

License

Owner

Hyunwoong Ko

This library is testing the ethics of language models by using natural adversarial texts.

A framework for cleaning Chinese dialog data

Meta learning algorithms to train cross-lingual NLI (multi-task) models

Semantic search for quotes.

BERT, LDA, and TFIDF based keyword extraction in Python

CJK computer science terms comparison / 中日韓電腦科學術語對照 / 日中韓のコンピュータ科学の用語対照 / 한·중·일 전산학 용어 대조

Bpe algorithm can finetune tokenizer - Bpe algorithm can finetune tokenizer

🤗 The largest hub of ready-to-use NLP datasets for ML models with fast, easy-to-use and efficient data manipulation tools

NewsMTSC: (Multi-)Target-dependent Sentiment Classification in News Articles

天池中药说明书实体识别挑战冠军方案；中文命名实体识别；NER; BERT-CRF & BERT-SPAN & BERT-MRC；Pytorch

Finally decent dictionaries based on Wiktionary for your beloved eBook reader.

Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

This is the source code of RPG (Reward-Randomized Policy Gradient)

Unsupervised Language Modeling at scale for robust sentiment classification

Sentiment Analysis Project using Count Vectorizer and TF-IDF Vectorizer

Conversational text Analysis using various NLP techniques

Smart discord chatbot integrated with Dialogflow to manage different classrooms and assist in teaching!

State of the Art Natural Language Processing

Simple bots or Simbots is a library designed to create simple bots using the power of python. This library utilises Intent, Entity, Relation and Context model to create bots .