Free MLOps course from DataTalks.Club

Last update: Dec 31, 2022

Related tags

Machine Learning mlops-zoomcamp

Overview

MLOps Zoomcamp

Our MLOps Zoomcamp course

Sign up here: https://airtable.com/shrCb8y6eTbPKwSTL (it's not automated, you will not receive an email immediately after filling in the form)
Register in DataTalks.Club's Slack
Join the #course-mlops-zoomcamp channel
Tweet about the course!
Subscribe to the public Google calendar (subscription works from desktop only)
Start watching course videos! Course playlist
Technical FAQ

Overview

Objective

Teach practical aspects of productionizing ML services — from collecting requirements to model deployment and monitoring.

Target audience

Data scientists and ML engineers. Also software engineers and data engineers interested in learning about putting ML in production.

Pre-requisites

Python
Docker
Being comfortable with command line
Prior exposure to machine learning (at work or from other courses, e.g. from ML Zoomcamp)
Prior programming experience (at least 1+ year)

Timeline

Course start: 16 of May

Syllabus

This is a draft and will change.

Module 1: Introduction

What is MLOps
MLOps maturity model
Running example: NY Taxi trips dataset
Why do we need MLOps
Course overview
Environment preparation
Homework

More details

Module 2: Experiment tracking and model management

Experiment tracking intro
Getting started with MLflow
Experiment tracking with MLflow
Saving and loading models with MLflow
Model registry
MLflow in practice
Homework

More details

Module 3: Orchestration and ML Pipelines

ML Pipelines: introduction
Prefect
Turning a notebook into a pipeline
Kubeflow Pipelines
Homework

Module 4: Model Deployment

Batch vs online
For online: web services vs streaming
Serving models in Batch mode
Web services
Streaming (Kinesis/SQS + AWS Lambda)
Homework

Module 5: Model Monitoring

ML monitoring vs software monitoring
Data quality monitoring
Data drift / concept drift
Batch vs real-time monitoring
Tools: Evidently, Prometheus and Grafana
Homework

Module 6: Best Practices

Devops
Virtual environments and Docker
Python: logging, linting
Testing: unit, integration, regression
CI/CD (github actions)
Infrastructure as code (terraform, cloudformation)
Cookiecutter
Makefiles
Homework

Module 7: Processes

CRISP-DM, CRISP-ML
ML Canvas
Data Landscape canvas
MLOps Stack Canvas
Documentation practices in ML projects (Model Cards Toolkit)

Project

End-to-end project with all the things above

Running example

To make it easier to connect different modules together, we’d like to use the same running example throughout the course.

Possible candidates:

https://www1.nyc.gov/site/tlc/about/tlc-trip-record-data.page - predict the ride duration or if the driver is going to be tipped or not

Instructors

Larysa Visengeriyeva
Cristian Martinez
Kevin Kho
Theofilos Papapanagiotou
Alexey Grigorev
Emeli Dral
Sejal Vaidya

Other courses from DataTalks.Club:

FAQ

I want to start preparing for the course. What can I do?

If you haven't used Flask or Docker

Check Module 5 form ML Zoomcamp
The section about Docker from Data Engineering Zoomcamp could also be useful

If you have no previous experience with ML

Check Module 1 from ML Zoomcamp for an overview
Module 3 will also be helpful if you want to learn Scikit-Learn (we'll use it in this course)
We'll also use XGBoost. You don't have to know it well, but if you want to learn more about it, refer to module 6 of ML Zoomcamp

I registered but haven't received an invite link. Is it normal?

Yes, we haven't automated it. You'll get a mail from us eventually, don't worry.

If you want to make sure you don't miss anything:

Register in our Slack and join the #course-mlops-zoomcamp channel
Subscribe to our YouTube channel

Is it going to be live?

No and yes. There will be two parts:

Lectures: Pre-recorded, you can watch them when it's convenient for you.
Office hours: Live on Mondays (17:00 CET), but recorded, so you can watch later.

Supporters and partners

Thanks to the course sponsors for making it possible to create this course

Thanks to our friends for spreading the word about the course

Free MLOps course from DataTalks.Club

Related tags

Overview

MLOps Zoomcamp

Overview

Objective

Target audience

Pre-requisites

Timeline

Syllabus

Module 1: Introduction

Module 2: Experiment tracking and model management

Module 3: Orchestration and ML Pipelines

Module 4: Model Deployment

Module 5: Model Monitoring

Module 6: Best Practices

Module 7: Processes

Project

Running example

Instructors

Other courses from DataTalks.Club:

FAQ

Supporters and partners

Owner

DataTalksClub

A Powerful Serverless Analysis Toolkit That Takes Trial And Error Out of Machine Learning Projects

Dual Adaptive Sampling for Machine Learning Interatomic potential.

A Streamlit demo to interactively visualize Uber pickups in New York City

Test symmetries with sklearn decision tree models

A library of sklearn compatible categorical variable encoders

Data from "Datamodels: Predicting Predictions with Training Data"

Python library which makes it possible to dynamically mask/anonymize data using JSON string or python dict rules in a PySpark environment.

Flask app to predict daily radiation from the time series of Solcast from Islamabad, Pakistan

SynapseML - an open source library to simplify the creation of scalable machine learning pipelines

Getting Profit and Loss Make Easy From Binance

PyNNDescent is a Python nearest neighbor descent for approximate nearest neighbors.

PLUR is a collection of source code datasets suitable for graph-based machine learning.

Open MLOps - A Production-focused Open-Source Machine Learning Framework

Warren - Stock Price Predictor

A comprehensive set of fairness metrics for datasets and machine learning models, explanations for these metrics, and algorithms to mitigate bias in datasets and models.

This is a Machine Learning model which predicts the presence of Diabetes in Patients

NumPy-based implementation of a multilayer perceptron (MLP)

A concept I came up which ditches the idea of "layers" in a neural network.

Required for a machine learning pipeline data preprocessing and variable engineering script needs to be prepared

Compare MLOps Platforms. Breakdowns of SageMaker, VertexAI, AzureML, Dataiku, Databricks, h2o, kubeflow, mlflow...