SKD010 - Model slower than baseline#

This example walks through mitigations when the check SKD010 fires. The check compares the user’s model to a fast linear baseline (RidgeCV for regression, wrapped in tabular_pipeline()) and flags a problem only when both of the following hold:

  • the larger of the fit-time and test predict-time ratios is at least 2 times the baseline, with an absolute gap of at least 1 second on that dimension, and

  • test scores are not significantly better than the baseline on a majority of default predictive metrics.

The issue is paying for latency without a significant quality premium. A clear quality gain over RidgeCV keeps the check quiet. So does a model that is already about as fast as the baseline.

Responses tried below, in order:

  • check that preprocessing is not the actual bottleneck,

  • reduce the model’s complexity,

  • switch to the fast linear baseline when quality is sufficient,

  • profile fit time to understand the dominant cost.

We use the employee salaries dataset with a heavy RandomForestRegressor inside tabular_pipeline(). The goal is either to match RidgeCV quality at lower cost, or to shrink fit time until the speed gap is no longer unjustified.

Load the employee salaries dataset#

fetch_employee_salaries() returns human-resources records with mixed categorical and numeric fields. A 200-tree RandomForestRegressor inside tabular_pipeline() trains slowly yet often fails to beat the fast RidgeCV baseline on test scores. SKD010 is a slow check.

Inspect column types with TableReport. Encoding cost is part of total fit time.

from skrub import TableReport

TableReport(X)

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



Salaries are continuous and right-skewed, a typical regression target.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



from skore import TrainTestSplit

splitter = TrainTestSplit(test_size=0.2, random_state=42)

Trigger SKD010 with large leaves#

One hundred trees that must keep 100 training rows in every leaf barely split. tabular_pipeline() still encodes the high-cardinality strings, so the fit stays much slower than RidgeCV, and the test scores lose the quality premium a forest can have on this table.

from sklearn.ensemble import RandomForestRegressor
from skore import evaluate
from skrub import tabular_pipeline

report_constrained = evaluate(
    tabular_pipeline(
        RandomForestRegressor(
            n_estimators=100,
            min_samples_leaf=100,
            random_state=42,
            n_jobs=4,
        )
    ),
    X=X,
    y=y,
    splitter=splitter,
)
report_constrained
Pipeline(steps=[('tablevectorizer',
                 TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
                                                                unknown_value=-1))),
                ('randomforestregressor',
                 RandomForestRegressor(min_samples_leaf=100, n_jobs=4,
                                       random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



SKD010 is in the issues below: fit time is several times the RidgeCV baseline, and the test scores are not significantly better.

report_constrained.checks.summarize()


Fit the fast linear baseline#

RidgeCV inside tabular_pipeline() is the reference skore uses for regression speed. This pipeline is the baseline SKD010 compares against, fit on the same split as the forest above.

from sklearn.linear_model import RidgeCV

report_linear = evaluate(
    tabular_pipeline(RidgeCV()),
    X=X,
    y=y,
    splitter=splitter,
)
report_linear
Pipeline(steps=[('tablevectorizer',
                 TableVectorizer(datetime=DatetimeEncoder(periodic_encoding='spline'))),
                ('simpleimputer', SimpleImputer(add_indicator=True)),
                ('squashingscaler', SquashingScaler(max_absolute_value=5)),
                ('ridgecv', RidgeCV())])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



from skore import compare

compare({"large_leaves": report_constrained, "ridge": report_linear}).metrics.summarize(
    data_source="test",
).frame()
estimator large_leaves ridge
metric
r2 0.737815 0.758462
rmse 15024.106340 14420.395938
mae 8599.922385 9111.963211
mape 0.125128 0.131599
fit_time 4.949131 1.048433
predict_time 0.283453 0.267647


Fit time clears the 2x / 1s gate. R², RMSE, and MAE stay close to RidgeCV, short of max(0.01, 0.05 * |baseline|) on a majority of the default predictive metrics. That is the warning.

Grow the trees to full depth#

Drop the leaf limit and keep the same 100 trees. The fit gets slower. The test scores move far enough ahead of RidgeCV that SKD010 passes.

report_full = evaluate(
    tabular_pipeline(
        RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4)
    ),
    X=X,
    y=y,
    splitter=splitter,
)
report_full
Pipeline(steps=[('tablevectorizer',
                 TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
                                                                unknown_value=-1))),
                ('randomforestregressor',
                 RandomForestRegressor(n_jobs=4, random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



SKD010 is among the passed checks. Slowness alone is not the warning.

report_full.checks.summarize()


compare(
    {
        "large_leaves": report_constrained,
        "full_depth": report_full,
        "ridge": report_linear,
    }
).metrics.summarize(data_source="test").frame()
estimator large_leaves full_depth ridge
metric
r2 0.737815 0.900830 0.758462
rmse 15024.106340 9240.077297 14420.395938
mae 8599.922385 4332.343421 9111.963211
mape 0.125128 0.058477 0.131599
fit_time 4.949131 11.314932 1.048433
predict_time 0.283453 0.299839 0.267647


Fit time increased from the constrained forest, and R² cleared the quality bar. The warning is gone.

Use a cheaper encoder#

Sometimes you can spend less time before touching the forest. Drop a column you do not need: date_first_hired repeats year_first_hired, and a StringEncoder still has to run on its thousands of distinct dates. Or keep the columns and pick a cheaper encoder. A tree can split on category codes, so an OrdinalEncoder is enough for the high-cardinality strings that tabular_pipeline() otherwise sends to StringEncoder.

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OrdinalEncoder
from skrub import TableVectorizer

report_ordinal = evaluate(
    make_pipeline(
        TableVectorizer(
            high_cardinality=OrdinalEncoder(
                handle_unknown="use_encoded_value",
                unknown_value=-1,
                encoded_missing_value=-1,
            )
        ),
        RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4),
    ),
    X=X,
    y=y,
    splitter=splitter,
)
report_ordinal
Pipeline(steps=[('tablevectorizer',
                 TableVectorizer(high_cardinality=OrdinalEncoder(encoded_missing_value=-1,
                                                                 handle_unknown='use_encoded_value',
                                                                 unknown_value=-1))),
                ('randomforestregressor',
                 RandomForestRegressor(n_jobs=4, random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



compare(
    {"full_depth": report_full, "ordinal_encoder": report_ordinal}
).metrics.summarize(data_source="test").frame()
estimator full_depth ordinal_encoder
metric
r2 0.900830 0.804144
rmse 9240.077297 12985.332649
mae 4332.343421 6681.730995
mape 0.058477 0.090946
fit_time 11.314932 2.627551
predict_time 0.299839 0.221712


The OrdinalEncoder cuts fit time and predict time. Test R² shows what that speed costs relative to the StringEncoder.

Reduce model complexity#

Keep the 100 trees and the full depth, and stop splits that would leave a leaf with fewer than 20 training rows. min_samples_leaf shortens each tree, which cuts fit time and predict time on the StringEncoder pipeline.

report_leaf = evaluate(
    tabular_pipeline(
        RandomForestRegressor(
            n_estimators=100,
            min_samples_leaf=20,
            random_state=42,
            n_jobs=4,
        )
    ),
    X=X,
    y=y,
    splitter=splitter,
)
report_leaf
Pipeline(steps=[('tablevectorizer',
                 TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
                                                                unknown_value=-1))),
                ('randomforestregressor',
                 RandomForestRegressor(min_samples_leaf=20, n_jobs=4,
                                       random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



SKD010 stays passed. The larger leaves shorten fit time on the same 100 trees.

report_leaf.checks.summarize()


report_leaf.metrics.summarize(
    metric=["fit_time", "predict_time"],
    data_source="test",
).frame()
metric
fit_time        7.048906
predict_time    0.292267
Name: RandomForestRegressor, dtype: float64

Compare the pipelines#

Fit time, test predict time, and test scores for every pipeline above. Large leaves are the case that raises SKD010. Full depth is slower and justified. The OrdinalEncoder and a moderate leaf limit spend less time. RidgeCV is the speed reference: SKD010 asks whether extra seconds buy a significant score gain over that baseline.

comparison = compare(
    {
        "large_leaves": report_constrained,
        "full_depth": report_full,
        "ordinal_encoder": report_ordinal,
        "moderate_leaves": report_leaf,
        "ridge": report_linear,
    }
)
comparison.metrics.summarize(data_source="test").frame()
estimator large_leaves full_depth ordinal_encoder moderate_leaves ridge
metric
r2 0.737815 0.900830 0.804144 0.835295 0.758462
rmse 15024.106340 9240.077297 12985.332649 11907.957168 14420.395938
mae 8599.922385 4332.343421 6681.730995 5788.395857 9111.963211
mape 0.125128 0.058477 0.090946 0.078879 0.131599
fit_time 4.949131 11.314932 2.627551 7.048906 1.048433
predict_time 0.283453 0.299839 0.221712 0.292267 0.267647


Conclusion#

A usual forest can be slower than RidgeCV and still pass SKD010 when test scores are significantly better. A cheaper high-cardinality encoder, such as OrdinalEncoder, or dropping a redundant column such as date_first_hired, cuts that cost. min_samples_leaf shrinks the trees while keeping the same number of trees. RidgeCV remains the speed reference when a linear model is enough.

Total running time of the script: (1 minutes 33.161 seconds)

Gallery generated by Sphinx-Gallery