AdamW optimizer for bfloat16 models in pytorch.

Last update: Nov 20, 2022

Related tags

Deep Learning adamw_bfloat16

Overview

_{Image source}

AdamW optimizer for bfloat16 models in pytorch.

Bfloat16 is currently an optimal tradeoff between range and relative error for deep networks.
Bfloat16 can be used quite efficiently on Nvidia GPUs with Ampere architecture (A100, A10, A30, RTX3090...)

However, neither AMP in pytorch is ready for bfloat16, nor optimizers.

If you just convert all weights and inputs to bfloat16, you're likely to run into an issue of stale weights: updates are too small to modify bfloat16 weight (see gopher paper, section C2 for a large-scale example).

There are two possible remedies:

keep weights in float32 (precise) and bfloat16 (approximate)
keep weights in bfloat16, and keep correction term in bfloat16

As recent study has shown, both options are completely competitive in quality to float32 training.

Usage

Install:

pip install git+https://github.com/arogozhnikov/adamw_bfloat16.git

Use as a drop-in replacement for pytorch's AdamW:

import torch
from adamw_bfloat16 import LR, AdamW_BF16
model = model.to(torch.bfloat16)

# default preheat and decay
optimizer = AdamW_BF16(model.parameters())

# configure LR schedule. Use built-in scheduling opportunity
optimizer = AdamW_BF16(model.parameters(), lr_function=LR(lr=1e-4, preheat_steps=5000, decay_power=-0.25))

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Initial implementation of AdamW for pytorch supports cuda graphs and has a built-in mechanism for control of learning rate, because external are unlikely to make a friendship with cuda graphs
Source code(tar.gz)
Source code(zip)

AdamW optimizer for bfloat16 models in pytorch.

Related tags

Overview

AdamW optimizer for bfloat16 models in pytorch.

Usage

You might also like...

Storage-optimizer - Identify potintial optimizations on the cloud storage accounts

PyTorch implementation and pretrained models for XCiT models. See XCiT: Cross-Covariance Image Transformer

Objective of the repository is to learn and build machine learning models using Pytorch. 30DaysofML Using Pytorch

Pretrained SOTA Deep Learning models, callbacks and more for research and production with PyTorch Lightning and PyTorch

A bunch of random PyTorch models using PyTorch's C++ frontend

PyTorch-LIT is the Lite Inference Toolkit (LIT) for PyTorch which focuses on easy and fast inference of large models on end-devices.

Pytorch-diffusion - A basic PyTorch implementation of 'Denoising Diffusion Probabilistic Models'

pyhsmm - library for approximate unsupervised inference in Bayesian Hidden Markov Models (HMMs) and explicit-duration Hidden semi-Markov Models (HSMMs), focusing on the Bayesian Nonparametric extensions, the HDP-HMM and HDP-HSMM, mostly with weak-limit approximations.

Releases(v0.1.0)

v0.1.0(Dec 14, 2021)

Owner

Alex Rogozhnikov

"Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices", official implementation

Predict bus arrival time using VertexAI and Nvidia's Jetson Nano

Advanced Signal Processing Notebooks and Tutorials

Neural style in TensorFlow! 🎨

Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose Estimation

Off-policy continuous control in PyTorch, with RDPG, RTD3 & RSAC

High performance, easy-to-use, and scalable machine learning (ML) package, including linear model (LR), factorization machines (FM), and field-aware factorization machines (FFM) for Python and CLI interface.

Streamlit App For Product Analysis - Streamlit App For Product Analysis

Rank 3 : Source code for OPPO 6G Data Generation Challenge

It helps user to learn Pick-up lines and share if he has a better one

Implementation of the "PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences" paper.

AITUS - An atomatic notr maker for CYTUS

Alpha-IoU: A Family of Power Intersection over Union Losses for Bounding Box Regression

Simple SN-GAN to generate CryptoPunks

XViT - Space-time Mixing Attention for Video Transformer

Leibniz is a python package which provide facilities to express learnable partial differential equations with PyTorch

Implementation of the ICCV'21 paper Temporally-Coherent Surface Reconstruction via Metric-Consistent Atlases

Efficiently Disentangle Causal Representations

A pre-trained language model for social media text in Spanish

π-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis