Перейти к содержимому

What Happens When You Call model.fit() in XGBoost? | XGBoost | Machine Learning

Gradient Canvas

0:00 / 0:00

What Happens When You Call model.fit() in XGBoost? | XGBoost | Machine Learning

139 просмотров · 2 месяца назад
Gradient Canvas
74 подписчика
139 просмотров · 2 месяца назад
What does model.fit(X, y) actually do in XGBoost? This XGBoost tutorial opens up that single line of code and shows you every step it triggers — from raw data all the way to a finished ensemble of gradient-boosted decision trees. If you've ever called .fit() and wondered what happens in those few seconds, this is the complete, no-hand-waving walkthrough. We build the whole XGBoost training loop from the ground up: how fit bins your features into a DMatrix for speed, why it starts from a constant base_score, and then the core boosting loop — computing each data point's gradient and hessian, growing one regularized decision tree, and choosing splits with XGBoost's gain formula. You'll see where the closed-form leaf weight w* = −G/(H+λ) comes from, what the learning rate (shrinkage) really does, and how subsample, colsample, and early_stopping_rounds shape training. By the end, every hyperparameter you pass to XGBoost — n_estimators, learning_rate, max_depth, gamma, lambda, min_child_weight — maps onto a concrete step inside fit. What you'll learn • What model.fit(X, y) does internally in XGBoost, step by step • How gradient boosting computes gradients and hessians of the loss each round • The XGBoost regularized objective and second-order Taylor expansion • The split-gain formula and how histogram binning (max_bin) makes it fast • The closed-form leaf weight −G/(H+λ) and how lambda (L2) regularizes it • Pruning with gamma, min_child_weight, and sparsity-aware missing-value handling • Shrinkage / learning_rate, row & column subsampling, and early stopping • How predict reuses the trained trees: base_score + η·Σ trees This XGBoost explained walkthrough is aimed at anyone learning machine learning, gradient boosting, or decision tree ensembles — a beginner-friendly but complete look under the hood of the most popular tabular ML algorithm. If gradient boosting, XGBoost hyperparameter tuning, or the math behind boosted trees has ever felt like a black box, this video turns it into a clear, repeatable mental model. 📌 Chapters 00:00 Intro 00:12 One line: model.fit(X, y) 00:56 Inputs X, y — the goal F(x)≈y 01:38 Step 1: Bin the data (DMatrix) 02:31 Step 2: The base_score 03:12 Gradients & Hessians (g, h) 04:08 The regularized objective 05:03 The split-gain formula 06:01 Pruning, min_child_weight & missing values 06:58 Leaf weight: w* = −G/(H+λ) 07:46 Shrinkage (learning_rate) 08:37 Subsampling (rows & columns) 09:26 The loop & early stopping 10:14 What fit leaves behind + predict 11:08 Recap: the whole call 🔗 More from Gradient Canvas • How XGBoost Works (full intuition → math):    • XGBoost | How it works? | Boosting, Gradie...   • The Bias–Variance Tradeoff:    • Bias vs Variance: Why Models Overfit and U...   • Machine Learning playlist:    • Machine Learning   👍 If this made XGBoost's fit click, like and subscribe for more visual, intuition-first machine learning explainers — new videos on the math behind ML, deep learning, and RL. Keywords: xgboost, model.fit xgboost, xgboost fit explained, xgboost tutorial, gradient boosting explained, xgboost model.fit, how xgboost works, gradient boosted trees, xgboost hyperparameters, learning rate, n_estimators, max_depth, gamma, lambda, min_child_weight, gradient and hessian, second order taylor, split gain, leaf weight, DMatrix, max_bin, early stopping, subsample, colsample, decision tree ensemble, machine learning tutorial, boosting algorithm, tabular machine learning #XGBoost #MachineLearning #GradientBoosting #DataScience #ML