{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "\n\n# SKD006 - Coefficient interpretation\n\n`SKD006 <skd006-unscaled-coefficients>` is a tip about interpretation.\nThis check flags that on mixed-scale features, raw coefficient magnitudes are\nnot directly comparable across columns, and after standardization they are\ncomparable but no longer expressed in the original feature units. Which path\nyou take depends on what you want to read from the coefficients. See also\nscikit-learn's example on [linear model coefficient interpretation](https://scikit-learn.org/stable/auto_examples/inspection/plot_linear_model_coefficient_interpretation.html).\n\nWhen reading coefficients under SKD006, think of:\n\n- standardizing the inputs when you want coefficients that are comparable\n  across features (see [scale matters](https://scikit-learn.org/stable/auto_examples/inspection/plot_linear_model_coefficient_interpretation.html#interpreting-coefficients-scale-matters)),\n- multiplying each coefficient by the feature's standard deviation for an\n  \"effect per one std\" without refitting,\n- relying on a scale-invariant ranking such as permutation importance when you\n  need importance that does not depend on coefficient units.\n\nWe use California housing regression with\n:class:`~sklearn.linear_model.Ridge` on raw numeric columns. The goal is to\ninterpret effect sizes without mistaking scale differences for importance.\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Load the California housing dataset\n\nMedian income is measured in 10k USD blocks, latitude in degrees, and\npopulation in head counts. Fitting Ridge on these raw columns produces\ncoefficients whose magnitudes reflect units as much as predictive strength.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skrub.datasets import fetch_california_housing\n\nhousing = fetch_california_housing()\nX, y = housing.X, housing.y"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        ":class:`~skrub.TableReport` highlights the range differences across columns.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skrub import TableReport\n\nTableReport(X)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "TableReport(y)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "A single :class:`~skore.TrainTestSplit` with a 20 % test set feeds every\nevaluation below.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skore import TrainTestSplit\n\nsplitter = TrainTestSplit(test_size=0.2, random_state=42)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Trigger SKD006 with Ridge on raw mixed-scale features\n\nRidge with a moderate ``alpha`` fits quickly on unscaled inputs. SKD006\nindicates that the features are not on the same scale, so raw coefficient\nmagnitudes should not be read as a ranking of importance.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.linear_model import Ridge\nfrom skore import evaluate\n\nreport = evaluate(\n    Ridge(alpha=1.0),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Find ``SKD006`` in the Tips tab below.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report.checks.summarize()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Let us inspect coefficients with\n:meth:`~skore.EstimatorReport.inspection.coefficients`, and\ncompare each coefficient's magnitude to the feature's typical range (or\nstandard deviation): a large coefficient on a small-scale column is not\nnecessarily more important than a small coefficient on a large-scale column\nsuch as ``Population``.\n\nUnscaled coefficients are still useful on their own: they answer \"if I change\nthis feature by one unit in its original scale, how much does the prediction\nchange in target units?\" They are misleading regarding feature importance.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "coef_display = report.inspection.coefficients()\ncoef_display.frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Side by side, the raw coefficient magnitudes and the feature standard deviations tell\ndifferent stories. For instance, ``Population`` has a tiny coefficient because a\none-person change is negligible on a head-count scale, even though the column varies a\nlot across districts.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "import pandas as pd\n\ncoef = coef_display.frame(include_intercept=False).set_index(\"feature\")\nfeature_std = report.X_train.std().rename(\"feature std. dev. (train)\")\n\ndf = pd.concat([coef, feature_std], axis=1)\naxes = df.plot.barh(\n    subplots=True,\n    layout=(1, 2),\n    legend=False,\n    sharey=True,\n    sharex=False,\n    figsize=(10, 4),\n)\naxes[0, 0].set_xlabel(\"coefficient\")\n_ = axes[0, 1].set_xlabel(\"std. dev.\")"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Scale coefficients by feature standard deviation\n\nRecall the pitfall above: when the model was trained on features with\ndifferent dynamic ranges, you cannot compare features by looking at raw\ncoefficients alone. Multiplying each fitted coefficient by its feature\nstandard deviation gives an \"effect per one standard deviation\" that is\ncomparable across columns without refitting. See scikit-learn's discussion of\n[scale matters](https://scikit-learn.org/stable/auto_examples/inspection/plot_linear_model_coefficient_interpretation.html#interpreting-coefficients-scale-matters).\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "comparable = coef.copy()\ncomparable[\"effect per std. dev.\"] = comparable[\"coefficient\"] * feature_std\n_ = comparable.plot.barh()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "That rescaling erases the surprises from the side-by-side bars above. The\nlarge raw ``AveBedrms`` coefficient shrinks once multiplied by its small\nstandard deviation, so it drops in the ranking while other features become, in\ncomparison, more important.\n\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Standardize inputs in a pipeline\n\n:class:`~sklearn.preprocessing.StandardScaler` inside a\n:class:`~sklearn.pipeline.Pipeline` puts every feature on a common scale at\nfit time, so coefficients become directly comparable as \"effect per one\nstandard deviation.\" That is useful for ranking features, but you lose\nstatements in original units (for example, \"one extra year of age\").\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import StandardScaler\n\nreport_scaled = evaluate(\n    Pipeline([(\"scale\", StandardScaler()), (\"ridge\", Ridge(alpha=1.0))]),\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nreport_scaled"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "``SKD006`` is still reported in the Tips tab but the warning message changed.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "report_scaled.checks.summarize()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "The coefficient table is now on a shared scale. Read magnitudes as relative\nimportance after standardization, not as effects per original unit of income,\nrooms, or population.\n\nThese values are the same quantity as the post-hoc\n``coefficient * feature_std`` column from the previous section: both are an\neffect per one standard deviation. Fitting with ``StandardScaler`` does that\nrescaling inside the pipeline; multiplying raw coefficients by the train-set\nstd does it after the fact. Small numerical differences can remain because the\nscaler's internal mean/std and ``X_train.std()`` are not bit-identical, but the\nranking and interpretation should match.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "pd.concat(\n    [\n        report_scaled.inspection.coefficients()\n        .frame(include_intercept=False)\n        .set_index(\"feature\"),\n        comparable[\"effect per std. dev.\"],\n    ],\n    axis=1,\n)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Permutation importance (scale-invariant)\n\nScaled coefficients are one way to rank features. Another is permutation\nimportance on held-out data: shuffle one column at a time and measure how\nmuch the score drops. That answers questions such as \"which features did the\nmodel rely on most?\" and works for nonlinear models as well. Present it as an\nalternative to scaled coefficients when you care about predictive reliance\nrather than a linear effect size.\n\nCorrelated features can distort both coefficient rankings and permutation\nimportance (credit is shared or shifted between partners). See the\n`SKD008 example <example_skd008_correlated_features>` and scikit-learn's\nnote on [misleading values on strongly correlated features](https://scikit-learn.org/stable/modules/permutation_importance.html#misleading-values-on-strongly-correlated-features).\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "display = report.inspection.permutation_importance(\n    seed=42,\n    n_repeats=5,\n)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "display.frame()"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "display.plot()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Conclusion\n\nSKD006 is a reminder to be careful when interpreting coefficients of linear\nmodels. When features are on different scales, linear coefficients mix\nfeature importance with feature magnitude; when features are standardized,\ncoefficients are comparable but no longer in original units.\n\nUnscaled coefficients remain useful when the question is \"if I change this\nfeature by this much in its own units, how much does the prediction change?\"\nChoose standardization, an effect-per-std rescaling, or a scale-invariant\nmethod such as permutation importance depending on the question you want to\nanswer.\n\n"
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.14.8"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}