.. DO NOT EDIT. .. THIS FILE WAS AUTOMATICALLY GENERATED BY SPHINX-GALLERY. .. TO MAKE CHANGES, EDIT THE SOURCE PYTHON FILE: .. "auto_examples/pitfalls_and_solutions/plot_skd010_slower_than_baseline.py" .. LINE NUMBERS ARE GIVEN BELOW. .. only:: html .. note:: :class: sphx-glr-download-link-note :ref:`Go to the end ` to download the full example code. .. rst-class:: sphx-glr-example-title .. _sphx_glr_auto_examples_pitfalls_and_solutions_plot_skd010_slower_than_baseline.py: .. _example_skd010_slower_than_baseline: SKD010 - Model slower than baseline =================================== This example walks through mitigations when the check :ref:`SKD010 ` fires. The check compares the user's model to a fast linear baseline (:class:`~sklearn.linear_model.RidgeCV` for regression, wrapped in :func:`~skrub.tabular_pipeline`) and flags a problem only when both of the following hold: - the larger of the fit-time and test predict-time ratios is at least 2 times the baseline, with an absolute gap of at least 1 second on that dimension, and - test scores are not significantly better than the baseline on a majority of default predictive metrics. The issue is paying for latency without a significant quality premium. A clear quality gain over :class:`~sklearn.linear_model.RidgeCV` keeps the check quiet. So does a model that is already about as fast as the baseline. Responses tried below, in order: - check that preprocessing is not the actual bottleneck, - reduce the model's complexity, - switch to the fast linear baseline when quality is sufficient, - profile fit time to understand the dominant cost. We use the employee salaries dataset with a heavy :class:`~sklearn.ensemble.RandomForestRegressor` inside :func:`~skrub.tabular_pipeline`. The goal is either to match :class:`~sklearn.linear_model.RidgeCV` quality at lower cost, or to shrink fit time until the speed gap is no longer unjustified. .. GENERATED FROM PYTHON SOURCE LINES 39-48 Load the employee salaries dataset ================================== :func:`~skrub.datasets.fetch_employee_salaries` returns human-resources records with mixed categorical and numeric fields. A 200-tree :class:`~sklearn.ensemble.RandomForestRegressor` inside :func:`~skrub.tabular_pipeline` trains slowly yet often fails to beat the fast :class:`~sklearn.linear_model.RidgeCV` baseline on test scores. SKD010 is a slow check. .. GENERATED FROM PYTHON SOURCE LINES 48-55 .. code-block:: Python from skrub.datasets import fetch_employee_salaries dataset = fetch_employee_salaries() X = dataset.X y = dataset.y .. GENERATED FROM PYTHON SOURCE LINES 56-58 Inspect column types with :class:`~skrub.TableReport`. Encoding cost is part of total fit time. .. GENERATED FROM PYTHON SOURCE LINES 58-63 .. code-block:: Python from skrub import TableReport TableReport(X) .. raw:: html

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 64-65 Salaries are continuous and right-skewed, a typical regression target. .. GENERATED FROM PYTHON SOURCE LINES 65-68 .. code-block:: Python TableReport(y) .. raw:: html

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 69-73 .. code-block:: Python from skore import TrainTestSplit splitter = TrainTestSplit(test_size=0.2, random_state=42) .. GENERATED FROM PYTHON SOURCE LINES 74-82 Trigger SKD010 with large leaves ================================ One hundred trees that must keep 100 training rows in every leaf barely split. :func:`~skrub.tabular_pipeline` still encodes the high-cardinality strings, so the fit stays much slower than :class:`~sklearn.linear_model.RidgeCV`, and the test scores lose the quality premium a forest can have on this table. .. GENERATED FROM PYTHON SOURCE LINES 82-102 .. code-block:: Python from sklearn.ensemble import RandomForestRegressor from skore import evaluate from skrub import tabular_pipeline report_constrained = evaluate( tabular_pipeline( RandomForestRegressor( n_estimators=100, min_samples_leaf=100, random_state=42, n_jobs=4, ) ), X=X, y=y, splitter=splitter, ) report_constrained .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
                                                                    unknown_value=-1))),
                    ('randomforestregressor',
                     RandomForestRegressor(min_samples_leaf=100, n_jobs=4,
                                           random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 103-106 ``SKD010`` is in the issues below: fit time is several times the :class:`~sklearn.linear_model.RidgeCV` baseline, and the test scores are not significantly better. .. GENERATED FROM PYTHON SOURCE LINES 106-109 .. code-block:: Python report_constrained.checks.summarize() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 110-116 Fit the fast linear baseline ============================ :class:`~sklearn.linear_model.RidgeCV` inside :func:`~skrub.tabular_pipeline` is the reference skore uses for regression speed. This pipeline is the baseline SKD010 compares against, fit on the same split as the forest above. .. GENERATED FROM PYTHON SOURCE LINES 116-127 .. code-block:: Python from sklearn.linear_model import RidgeCV report_linear = evaluate( tabular_pipeline(RidgeCV()), X=X, y=y, splitter=splitter, ) report_linear .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(datetime=DatetimeEncoder(periodic_encoding='spline'))),
                    ('simpleimputer', SimpleImputer(add_indicator=True)),
                    ('squashingscaler', SquashingScaler(max_absolute_value=5)),
                    ('ridgecv', RidgeCV())])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 128-134 .. code-block:: Python from skore import compare compare({"large_leaves": report_constrained, "ridge": report_linear}).metrics.summarize( data_source="test", ).frame() .. raw:: html
estimator large_leaves ridge
metric
r2 0.737815 0.758462
rmse 15024.106340 14420.395938
mae 8599.922385 9111.963211
mape 0.125128 0.131599
fit_time 4.949131 1.048433
predict_time 0.283453 0.267647


.. GENERATED FROM PYTHON SOURCE LINES 135-139 Fit time clears the 2x / 1s gate. R², RMSE, and MAE stay close to :class:`~sklearn.linear_model.RidgeCV`, short of ``max(0.01, 0.05 * |baseline|)`` on a majority of the default predictive metrics. That is the warning. .. GENERATED FROM PYTHON SOURCE LINES 141-147 Grow the trees to full depth ============================ Drop the leaf limit and keep the same 100 trees. The fit gets slower. The test scores move far enough ahead of :class:`~sklearn.linear_model.RidgeCV` that SKD010 passes. .. GENERATED FROM PYTHON SOURCE LINES 147-158 .. code-block:: Python report_full = evaluate( tabular_pipeline( RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4) ), X=X, y=y, splitter=splitter, ) report_full .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
                                                                    unknown_value=-1))),
                    ('randomforestregressor',
                     RandomForestRegressor(n_jobs=4, random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 159-160 ``SKD010`` is among the passed checks. Slowness alone is not the warning. .. GENERATED FROM PYTHON SOURCE LINES 160-163 .. code-block:: Python report_full.checks.summarize() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 164-172 .. code-block:: Python compare( { "large_leaves": report_constrained, "full_depth": report_full, "ridge": report_linear, } ).metrics.summarize(data_source="test").frame() .. raw:: html
estimator large_leaves full_depth ridge
metric
r2 0.737815 0.900830 0.758462
rmse 15024.106340 9240.077297 14420.395938
mae 8599.922385 4332.343421 9111.963211
mape 0.125128 0.058477 0.131599
fit_time 4.949131 11.314932 1.048433
predict_time 0.283453 0.299839 0.267647


.. GENERATED FROM PYTHON SOURCE LINES 173-175 Fit time increased from the constrained forest, and R² cleared the quality bar. The warning is gone. .. GENERATED FROM PYTHON SOURCE LINES 177-188 Use a cheaper encoder ===================== Sometimes you can spend less time before touching the forest. Drop a column you do not need: ``date_first_hired`` repeats ``year_first_hired``, and a :class:`~skrub.StringEncoder` still has to run on its thousands of distinct dates. Or keep the columns and pick a cheaper encoder. A tree can split on category codes, so an :class:`~sklearn.preprocessing.OrdinalEncoder` is enough for the high-cardinality strings that :func:`~skrub.tabular_pipeline` otherwise sends to :class:`~skrub.StringEncoder`. .. GENERATED FROM PYTHON SOURCE LINES 188-210 .. code-block:: Python from sklearn.pipeline import make_pipeline from sklearn.preprocessing import OrdinalEncoder from skrub import TableVectorizer report_ordinal = evaluate( make_pipeline( TableVectorizer( high_cardinality=OrdinalEncoder( handle_unknown="use_encoded_value", unknown_value=-1, encoded_missing_value=-1, ) ), RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4), ), X=X, y=y, splitter=splitter, ) report_ordinal .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(high_cardinality=OrdinalEncoder(encoded_missing_value=-1,
                                                                     handle_unknown='use_encoded_value',
                                                                     unknown_value=-1))),
                    ('randomforestregressor',
                     RandomForestRegressor(n_jobs=4, random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 211-215 .. code-block:: Python compare( {"full_depth": report_full, "ordinal_encoder": report_ordinal} ).metrics.summarize(data_source="test").frame() .. raw:: html
estimator full_depth ordinal_encoder
metric
r2 0.900830 0.804144
rmse 9240.077297 12985.332649
mae 4332.343421 6681.730995
mape 0.058477 0.090946
fit_time 11.314932 2.627551
predict_time 0.299839 0.221712


.. GENERATED FROM PYTHON SOURCE LINES 216-227 The :class:`~sklearn.preprocessing.OrdinalEncoder` cuts fit time and predict time. Test R² shows what that speed costs relative to the :class:`~skrub.StringEncoder`. Reduce model complexity ======================= Keep the 100 trees and the full depth, and stop splits that would leave a leaf with fewer than 20 training rows. ``min_samples_leaf`` shortens each tree, which cuts fit time and predict time on the :class:`~skrub.StringEncoder` pipeline. .. GENERATED FROM PYTHON SOURCE LINES 227-243 .. code-block:: Python report_leaf = evaluate( tabular_pipeline( RandomForestRegressor( n_estimators=100, min_samples_leaf=20, random_state=42, n_jobs=4, ) ), X=X, y=y, splitter=splitter, ) report_leaf .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
                                                                    unknown_value=-1))),
                    ('randomforestregressor',
                     RandomForestRegressor(min_samples_leaf=20, n_jobs=4,
                                           random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 244-245 SKD010 stays passed. The larger leaves shorten fit time on the same 100 trees. .. GENERATED FROM PYTHON SOURCE LINES 245-248 .. code-block:: Python report_leaf.checks.summarize() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 249-254 .. code-block:: Python report_leaf.metrics.summarize( metric=["fit_time", "predict_time"], data_source="test", ).frame() .. rst-class:: sphx-glr-script-out .. code-block:: none metric fit_time 7.048906 predict_time 0.292267 Name: RandomForestRegressor, dtype: float64 .. GENERATED FROM PYTHON SOURCE LINES 255-264 Compare the pipelines ===================== Fit time, test predict time, and test scores for every pipeline above. Large leaves are the case that raises SKD010. Full depth is slower and justified. The :class:`~sklearn.preprocessing.OrdinalEncoder` and a moderate leaf limit spend less time. :class:`~sklearn.linear_model.RidgeCV` is the speed reference: SKD010 asks whether extra seconds buy a significant score gain over that baseline. .. GENERATED FROM PYTHON SOURCE LINES 264-277 .. code-block:: Python comparison = compare( { "large_leaves": report_constrained, "full_depth": report_full, "ordinal_encoder": report_ordinal, "moderate_leaves": report_leaf, "ridge": report_linear, } ) comparison.metrics.summarize(data_source="test").frame() .. raw:: html
estimator large_leaves full_depth ordinal_encoder moderate_leaves ridge
metric
r2 0.737815 0.900830 0.804144 0.835295 0.758462
rmse 15024.106340 9240.077297 12985.332649 11907.957168 14420.395938
mae 8599.922385 4332.343421 6681.730995 5788.395857 9111.963211
mape 0.125128 0.058477 0.090946 0.078879 0.131599
fit_time 4.949131 11.314932 2.627551 7.048906 1.048433
predict_time 0.283453 0.299839 0.221712 0.292267 0.267647


.. GENERATED FROM PYTHON SOURCE LINES 278-288 Conclusion ========== A usual forest can be slower than :class:`~sklearn.linear_model.RidgeCV` and still pass SKD010 when test scores are significantly better. A cheaper high-cardinality encoder, such as :class:`~sklearn.preprocessing.OrdinalEncoder`, or dropping a redundant column such as ``date_first_hired``, cuts that cost. ``min_samples_leaf`` shrinks the trees while keeping the same number of trees. :class:`~sklearn.linear_model.RidgeCV` remains the speed reference when a linear model is enough. .. rst-class:: sphx-glr-timing **Total running time of the script:** (1 minutes 33.161 seconds) .. _sphx_glr_download_auto_examples_pitfalls_and_solutions_plot_skd010_slower_than_baseline.py: .. only:: html .. container:: sphx-glr-footer sphx-glr-footer-example .. container:: sphx-glr-download sphx-glr-download-jupyter :download:`Download Jupyter notebook: plot_skd010_slower_than_baseline.ipynb ` .. container:: sphx-glr-download sphx-glr-download-python :download:`Download Python source code: plot_skd010_slower_than_baseline.py ` .. container:: sphx-glr-download sphx-glr-download-zip :download:`Download zipped: plot_skd010_slower_than_baseline.zip ` .. only:: html .. rst-class:: sphx-glr-signature `Gallery generated by Sphinx-Gallery `_