{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "\n\n# SKD010 - Model slower than baseline\n\nThis example walks through mitigations when the check\n`SKD010 <skd010-slower-than-baseline>` fires. The check compares the user's\nmodel to a fast linear baseline\n(:class:`~sklearn.linear_model.RidgeCV` for regression, wrapped in\n:func:`~skrub.tabular_pipeline`) and flags a problem only when both of the\nfollowing hold:\n\n- the larger of the fit-time and test predict-time ratios is at least 2 times\n  the baseline, with an absolute gap of at least 1 second on that dimension,\n  and\n- test scores are not significantly better than the baseline on a majority of\n  default predictive metrics.\n\nThe issue is paying for latency without a significant quality premium. A clear\nquality gain over :class:`~sklearn.linear_model.RidgeCV` keeps the check quiet.\nSo does a model that is already about as fast as the baseline.\n\nResponses tried below, in order:\n\n- check that preprocessing is not the actual bottleneck,\n- reduce the model's complexity,\n- switch to the fast linear baseline when quality is sufficient,\n- profile fit time to understand the dominant cost.\n\nWe use the employee salaries dataset with a heavy\n:class:`~sklearn.ensemble.RandomForestRegressor` inside\n:func:`~skrub.tabular_pipeline`. The goal is either to match\n:class:`~sklearn.linear_model.RidgeCV` quality at lower cost, or to shrink fit\ntime until the speed gap is no longer unjustified.\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Load the employee salaries dataset\n\n:func:`~skrub.datasets.fetch_employee_salaries` returns human-resources\nrecords with mixed categorical and numeric fields. A 200-tree\n:class:`~sklearn.ensemble.RandomForestRegressor` inside\n:func:`~skrub.tabular_pipeline` trains slowly yet often fails to beat\nthe fast :class:`~sklearn.linear_model.RidgeCV` baseline on test scores.\nSKD010 is a slow check.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skrub.datasets import fetch_employee_salaries\n\ndataset = fetch_employee_salaries()\nX = dataset.X\ny = dataset.y"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Inspect column types with :class:`~skrub.TableReport`. Encoding cost is part\nof total fit time.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skrub import TableReport\n\nTableReport(X)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Salaries are continuous and right-skewed, a typical regression target.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "TableReport(y)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skore import TrainTestSplit\n\nsplitter = TrainTestSplit(test_size=0.2, random_state=42)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Trigger SKD010 with large leaves\n\nOne hundred trees that must keep 100 training rows in every leaf barely\nsplit. :func:`~skrub.tabular_pipeline` still encodes the high-cardinality\nstrings, so the fit stays much slower than :class:`~sklearn.linear_model.RidgeCV`,\nand the test scores lose\nthe quality premium a forest can have on this table.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.ensemble import RandomForestRegressor\nfrom skore import evaluate\nfrom skrub import tabular_pipeline\n\nreport_constrained = evaluate(\n    tabular_pipeline(\n        RandomForestRegressor(\n            n_estimators=100,\n            min_samples_leaf=100,\n            random_state=42,\n            n_jobs=4,\n        )\n    ),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport_constrained"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "``SKD010`` is in the issues below: fit time is several times the\n:class:`~sklearn.linear_model.RidgeCV` baseline, and the test scores are not\nsignificantly better.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_constrained.checks.summarize()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Fit the fast linear baseline\n\n:class:`~sklearn.linear_model.RidgeCV` inside :func:`~skrub.tabular_pipeline`\nis the reference skore uses for regression speed. This pipeline is the baseline\nSKD010 compares against, fit on the same split as the forest above.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.linear_model import RidgeCV\n\nreport_linear = evaluate(\n    tabular_pipeline(RidgeCV()),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport_linear"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skore import compare\n\ncompare({\"large_leaves\": report_constrained, \"ridge\": report_linear}).metrics.summarize(\n    data_source=\"test\",\n).frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Fit time clears the 2x / 1s gate. R\u00b2, RMSE, and MAE stay close to\n:class:`~sklearn.linear_model.RidgeCV`,\nshort of ``max(0.01, 0.05 * |baseline|)`` on a majority of the default\npredictive metrics. That is the warning.\n\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Grow the trees to full depth\n\nDrop the leaf limit and keep the same 100 trees. The fit gets slower. The\ntest scores move far enough ahead of :class:`~sklearn.linear_model.RidgeCV`\nthat SKD010 passes.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_full = evaluate(\n    tabular_pipeline(\n        RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4)\n    ),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport_full"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "``SKD010`` is among the passed checks. Slowness alone is not the warning.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_full.checks.summarize()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "compare(\n    {\n        \"large_leaves\": report_constrained,\n        \"full_depth\": report_full,\n        \"ridge\": report_linear,\n    }\n).metrics.summarize(data_source=\"test\").frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Fit time increased from the constrained forest, and R\u00b2 cleared the quality\nbar. The warning is gone.\n\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Use a cheaper encoder\n\nSometimes you can spend less time before touching the forest. Drop a column you do not\nneed: ``date_first_hired`` repeats ``year_first_hired``, and a\n:class:`~skrub.StringEncoder` still has to run on its thousands of distinct\ndates. Or keep the columns and pick a cheaper\nencoder. A tree can split on category codes, so an\n:class:`~sklearn.preprocessing.OrdinalEncoder` is enough for the high-cardinality\nstrings that :func:`~skrub.tabular_pipeline` otherwise sends to\n:class:`~skrub.StringEncoder`.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.pipeline import make_pipeline\nfrom sklearn.preprocessing import OrdinalEncoder\nfrom skrub import TableVectorizer\n\nreport_ordinal = evaluate(\n    make_pipeline(\n        TableVectorizer(\n            high_cardinality=OrdinalEncoder(\n                handle_unknown=\"use_encoded_value\",\n                unknown_value=-1,\n                encoded_missing_value=-1,\n            )\n        ),\n        RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4),\n    ),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport_ordinal"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "compare(\n    {\"full_depth\": report_full, \"ordinal_encoder\": report_ordinal}\n).metrics.summarize(data_source=\"test\").frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "The :class:`~sklearn.preprocessing.OrdinalEncoder` cuts fit time and predict\ntime. Test R\u00b2 shows what that speed costs relative to the\n:class:`~skrub.StringEncoder`.\n\n# Reduce model complexity\n\nKeep the 100 trees and the full depth, and stop splits that would leave a\nleaf with fewer than 20 training rows. ``min_samples_leaf`` shortens each\ntree, which cuts fit time and predict time on the\n:class:`~skrub.StringEncoder` pipeline.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_leaf = evaluate(\n    tabular_pipeline(\n        RandomForestRegressor(\n            n_estimators=100,\n            min_samples_leaf=20,\n            random_state=42,\n            n_jobs=4,\n        )\n    ),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport_leaf"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "SKD010 stays passed. The larger leaves shorten fit time on the same 100 trees.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_leaf.checks.summarize()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_leaf.metrics.summarize(\n    metric=[\"fit_time\", \"predict_time\"],\n    data_source=\"test\",\n).frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Compare the pipelines\n\nFit time, test predict time, and test scores for every pipeline above. Large\nleaves are the case that raises SKD010. Full depth is slower and justified.\nThe :class:`~sklearn.preprocessing.OrdinalEncoder` and a moderate leaf limit\nspend less time. :class:`~sklearn.linear_model.RidgeCV` is the speed\nreference: SKD010 asks whether extra seconds buy a significant score gain\nover that baseline.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "comparison = compare(\n    {\n        \"large_leaves\": report_constrained,\n        \"full_depth\": report_full,\n        \"ordinal_encoder\": report_ordinal,\n        \"moderate_leaves\": report_leaf,\n        \"ridge\": report_linear,\n    }\n)\ncomparison.metrics.summarize(data_source=\"test\").frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Conclusion\n\nA usual forest can be slower than :class:`~sklearn.linear_model.RidgeCV` and\nstill pass SKD010 when test scores are significantly better. A cheaper\nhigh-cardinality encoder, such as :class:`~sklearn.preprocessing.OrdinalEncoder`,\nor dropping a redundant column such as ``date_first_hired``, cuts that cost.\n``min_samples_leaf`` shrinks the trees while keeping the same number of\ntrees. :class:`~sklearn.linear_model.RidgeCV` remains the speed reference when\na linear model is enough.\n\n"
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.14.8"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}