FastAPI Model Serving for ML | DeployBytes 8 - Line by Line
Dr. Sandeep Grover
0:00 / 0:00
FastAPI Model Serving for ML | DeployBytes 8 - Line by Line
23 просмотра · 13 дней назад
Dr. Sandeep Grover
81 подписчик
23 просмотра · 13 дней назад
Real MLOps, step by step. Part 8 serves the trained 7-encoder model over HTTP with FastAPI: a thin app on Uvicorn with /health and /predict, a lazy-loading singleton so the weights load once, a one-line Prometheus /metrics endpoint, and a container healthcheck and restart policy that keep the serving service self-healing.
Why this series exists:
This video is part of a series created to make central concepts easier to understand and hold onto. It was built with accessibility in mind, particularly for learners with learning differences, so the pace and style are kept simple and supportive. There may be a few mispronunciations in the spoken narration, and occasionally a small mistake in the content. These do not take away from the purpose of the series, which is simply to help the learning stick.
Work with me: / sandeep-grover-b3192a16
Productionizing a real 7-encoder multimodal Rakuten classifier (84,916 products, weighted-F1 0.9147), from git to Kubernetes.
Chapters:
0:00 Why this series exists
0:29 Step 8 - Serving hardening: metrics, healthcheck, sacremoses
2:09 The shape of this step
4:30 What Step 8 is: one Uvicorn line, and where the request goes
6:09 The five endpoints
7:16 From a product title to a prdtypecode: featurize, standardize, 20-head bag
10:10 Provenance vs serving: the MLflow registry and the local file
11:26 Live run
12:38 8a - the metrics dependency: requirements-extra.txt and two new Dockerfile layers
14:00 8b - the missing tokenizer helper: sacremoses into requirements.txt, then rebuild
15:17 8c - /metrics in one line: the Instrumentator in api.py
16:34 8d - self-healing: restart policy and healthcheck on the serving service
18:57 8e - run it: the Uvicorn CMD, then /health and /predict from the host
20:04 8f - commit
20:42 curl the health endpoint on the container port
21:47 The shape of this step
22:51 curl the health endpoint through the proxy
24:23 The shape of this step
25:28 Walkthrough: src/serving/api.py
33:56 Walkthrough: docker-compose.yml
40:52 Walkthrough: docker/serving/Dockerfile
44:43 Walkthrough: docker/serving/requirements-extra.txt
45:53 Walkthrough: docker/serving/requirements.txt
49:15 Reality check: two errors, the slow first request, and /train
51:12 Design decisions to keep
52:58 Takeaway, and on to Step 9
Series stack (all 12 episodes): Python, cookiecutter-data-science, Git/GitHub, Docker, Docker Compose, PostgreSQL, MinIO (S3), MLflow, DVC, FastAPI, Prometheus, pytest, Nginx, Apache Airflow, Kubernetes, Argo Rollouts, PyTorch, Hugging Face Transformers.
#MLOps #FastAPI #ModelServing #ModelDeployment #Prometheus #Docker #MachineLearning #Python #DevOps #MLEngineering