AdamW optimizer for bfloat16 models in pytorch.

Last update: Nov 20, 2022

Related tags

Deep Learning adamw_bfloat16

Overview

_{Image source}

AdamW optimizer for bfloat16 models in pytorch.

Bfloat16 is currently an optimal tradeoff between range and relative error for deep networks.
Bfloat16 can be used quite efficiently on Nvidia GPUs with Ampere architecture (A100, A10, A30, RTX3090...)

However, neither AMP in pytorch is ready for bfloat16, nor optimizers.

If you just convert all weights and inputs to bfloat16, you're likely to run into an issue of stale weights: updates are too small to modify bfloat16 weight (see gopher paper, section C2 for a large-scale example).

There are two possible remedies:

keep weights in float32 (precise) and bfloat16 (approximate)
keep weights in bfloat16, and keep correction term in bfloat16

As recent study has shown, both options are completely competitive in quality to float32 training.

Usage

Install:

pip install git+https://github.com/arogozhnikov/adamw_bfloat16.git

Use as a drop-in replacement for pytorch's AdamW:

import torch
from adamw_bfloat16 import LR, AdamW_BF16
model = model.to(torch.bfloat16)

# default preheat and decay
optimizer = AdamW_BF16(model.parameters())

# configure LR schedule. Use built-in scheduling opportunity
optimizer = AdamW_BF16(model.parameters(), lr_function=LR(lr=1e-4, preheat_steps=5000, decay_power=-0.25))

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Initial implementation of AdamW for pytorch supports cuda graphs and has a built-in mechanism for control of learning rate, because external are unlikely to make a friendship with cuda graphs
Source code(tar.gz)
Source code(zip)

AdamW optimizer for bfloat16 models in pytorch.

Related tags

Overview

AdamW optimizer for bfloat16 models in pytorch.

Usage

You might also like...

Storage-optimizer - Identify potintial optimizations on the cloud storage accounts

PyTorch implementation and pretrained models for XCiT models. See XCiT: Cross-Covariance Image Transformer

Objective of the repository is to learn and build machine learning models using Pytorch. 30DaysofML Using Pytorch

Pretrained SOTA Deep Learning models, callbacks and more for research and production with PyTorch Lightning and PyTorch

A bunch of random PyTorch models using PyTorch's C++ frontend

PyTorch-LIT is the Lite Inference Toolkit (LIT) for PyTorch which focuses on easy and fast inference of large models on end-devices.

Pytorch-diffusion - A basic PyTorch implementation of 'Denoising Diffusion Probabilistic Models'

pyhsmm - library for approximate unsupervised inference in Bayesian Hidden Markov Models (HMMs) and explicit-duration Hidden semi-Markov Models (HSMMs), focusing on the Bayesian Nonparametric extensions, the HDP-HMM and HDP-HSMM, mostly with weak-limit approximations.

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Owner

Alex Rogozhnikov

Controlling the MicriSpotAI robot from scratch

Perform zero-order Hankel Transform for an 1D array (float or real valued).

Image-Stitching - Panorama composition using SIFT Features and a custom implementaion of RANSAC algorithm

Implementations of CNNs, RNNs, GANs, etc

Pytorch library for fast transformer implementations

MLSpace: Hassle-free machine learning & deep learning development

PyTorch implementation of the paper: Label Noise Transition Matrix Estimation for Tasks with Lower-Quality Features

1st place solution in CCF BDCI 2021 ULSEG challenge

How to Leverage Multimodal EHR Data for Better Medical Predictions?

UMPNet: Universal Manipulation Policy Network for Articulated Objects

ACL'2021: LM-BFF: Better Few-shot Fine-tuning of Language Models

Towhee is a flexible machine learning framework currently focused on computing deep learning embeddings over unstructured data.

An implementation of an abstract algebra for music tones (pitches).

CvT-ASSD: Convolutional vision-Transformerbased Attentive Single Shot MultiBox Detector (ICTAI 2021 CCF-C 会议)The 33rd IEEE International Conference on Tools with Artificial Intelligence

Code for the prototype tool in our paper "CoProtector: Protect Open-Source Code against Unauthorized Training Usage with Data Poisoning".

CARL provides highly configurable contextual extensions to several well-known RL environments.

library for nonlinear optimization, wrapping many algorithms for global and local, constrained or unconstrained, optimization

This repo holds the code of TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation

[UNMAINTAINED] Automated machine learning for analytics & production

This reposityory contains the PyTorch implementation of our paper "Generative Dynamic Patch Attack".