{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "\n\n# SKD011 & SKD012 - Golden feature and useless features\n\nThis example walks through mitigations when checks\n`SKD011 <skd011-golden-feature>` and `SKD012 <skd012-useless-features>`\nfire together: a common pattern when a leaky column carries almost all signal\nand remaining inputs look negligible by comparison.\n\nMitigations from the `automated_checks` user guide, in the order we try\nthem here:\n\n**SKD011 - golden feature**\n\n- audit the suspect feature for leakage (is it derived from the target or\n  from data that would not be available at inference time?),\n- decide whether to keep or drop it,\n- collect or engineer additional features so the model is less dependent on\n  a single one.\n\n**SKD012 - useless features**\n\n- review the flagged features and consider dropping them,\n- refit on a reduced feature set and verify performance is preserved,\n- if a flagged feature should matter, investigate encoding (here: feature\n  engineering before pruning).\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Load the medical charge dataset\n\nLet's load the `medical_charge` dataset, which we will use to predict the average\ncost of an inpatient stay from the hospital's location and the type of medical\nprocedure.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skrub.datasets import fetch_medical_charge\n\ndataset = fetch_medical_charge()\nX_full, y_full = dataset.X, dataset.y"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Then we can inspect predictors and target with :class:`~skrub.TableReport`.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skrub import TableReport\n\nTableReport(X_full)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "TableReport(y_full)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Next, we subsample 3,000 rows to make the example run faster.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "X = X_full.sample(3_000, random_state=42).reset_index(drop=True)\ny = y_full.sample(3_000, random_state=42).reset_index(drop=True)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "We now create a splitter, vectorizer, and regressor that we reuse throughout\nthe example. :func:`~skore.evaluate` clones them before fitting.\n\nHigh-cardinality strings are encoded with :class:`~skrub.StringEncoder` (TF-IDF\nthen randomized SVD); we seed it so reruns stay comparable.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.ensemble import HistGradientBoostingRegressor\nfrom sklearn.pipeline import make_pipeline\nfrom skore import TrainTestSplit\nfrom skrub import StringEncoder, TableVectorizer\n\nsplitter = TrainTestSplit(random_state=42, test_size=0.2)\nvectorizer = TableVectorizer(high_cardinality=StringEncoder(random_state=42))\nregressor = HistGradientBoostingRegressor(random_state=42)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Trigger SKD011 and SKD012\n\nLet us train a gradient boosting model, using skrub's `TableVectorizer` to vectorize\nthe data first.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from skore import evaluate\n\nmodel = make_pipeline(vectorizer, regressor)\n\nfirst_report = evaluate(\n    model,\n    X=X,\n    y=y,\n    splitter=splitter,\n)\nfirst_report"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "We can see from the metrics table that the models performs very well on the data,\nwith an R\u00b2 of 0.96.\nHowever, looking at the checks, we can see in the Tips tab that SKD011 triggers on\nthe `Average_Medicare_Payments` column.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "first_report.checks.summarize()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "We should then question whether this column would be available in a real deployment.\nIf it would, then we have found a good proxy for the target and we can move on.\nIf it would not, then we should not use this column to train our model.\n\nIn our case, the target (`Average_Total_Payments`) and the golden feature\n(`Average_Medicare_Payments`) come from the same process (billing aggregates).\nTherefore, `Average_Medicare_Payments` would not be available in a real deployment,\nand we should not use it to train our model. As a matter of fact,\n`Average_Covered_Charges` would also not be able in a deployment setting so we will\nalso drop it.\n\nA final thing to be careful about is that SKD012 flags most provider fields\nand `Total_Discharges` as useless. As we can see in the importance plot\nbelow, every feature looks weak next to the golden billing column, even\nthough some of them may recover once we drop it. We will keep them for now.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "_ = first_report.inspection.permutation_importance().plot()\n\n\n# SKD011 - Inspect after dropping golden feature\n# ==============================================\n#\n# First let's evaluate the same model on the same test set with the payment features\n# removed, and compare it with our original model.\n#\n# Thankfully there is still signal in the remaining features, as we get a decent R\u00b2\n# of 0.89.\n\nfrom skore import compare\n\nX_without_payment = X.drop(\n    columns=[\"Average_Medicare_Payments\", \"Average_Covered_Charges\"]\n)\n\nsecond_report = evaluate(\n    model,\n    X=X_without_payment,\n    y=y,\n    splitter=splitter,\n)\n\ncomparison = compare(\n    {\n        \"with_payment_features\": first_report,\n        \"without_payment_features\": second_report,\n    }\n)\ncomparison.metrics.summarize().frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Let us inspect which features are now the ones our model relies on the most.\nWe can see that `DRG_Definition` is by far most important one, according to\npermutation importance. This feature encodes the medical reason for which patients\nwere treated. It will therefore be available at inference time and we can use it.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "_ = second_report.inspection.permutation_importance().plot()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "SKD012 now flags `Provider_Id` and `Total_Discharges`. Those look genuinely\nweak once billing leakage is gone: an identifier and a discharge count that\nbarely moves the score. Other provider fields recovered some signal, which\nis why we did not treat the leaky SKD012 list as columns to drop.\n\nThe check still only sees original columns. `TableVectorizer` creates extra\ncomponents from high-cardinality strings that can be weak too; we prune\nthose next.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "second_report.checks.summarize()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# SKD012 - prune weak features inside the pipeline\n\nSKD012 flagged `Provider_Id` and `Total_Discharges` on the original columns.\nIt does not see the extra components `TableVectorizer` builds from\nhigh-cardinality fields such as `DRG_Definition` and provider names; some of\nthose add little once the useful ones are present. We therefore select\n*after* vectorizing, with :class:`~sklearn.feature_selection.SelectFromModel`,\nso the choice is fit on the training fold only.\n\nHistogram gradient boosting has no `feature_importances_`, so the selector\nuses :class:`~sklearn.ensemble.GradientBoostingRegressor`. The final\nestimator stays a :class:`~sklearn.ensemble.HistGradientBoostingRegressor` on\nthe reduced matrix.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "from sklearn.ensemble import GradientBoostingRegressor\nfrom sklearn.feature_selection import SelectFromModel\n\nmodel_reduced = make_pipeline(\n    vectorizer,\n    SelectFromModel(\n        GradientBoostingRegressor(random_state=42),\n        threshold=\"median\",\n    ),\n    regressor,\n)\nmodel_reduced"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "Let us evaluate the reduced pipeline on the same split. Test scores are a\non par, and the model is lighter: selection dropped the weak\nvectorized components before fitting histogram boosting.\n\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "collapsed": false
      },
      "outputs": [],
      "source": [
        "third_report = evaluate(\n    model_reduced,\n    X=X_without_payment,\n    y=y,\n    splitter=splitter,\n)\n\ncomparison_reduced = compare(\n    {\n        \"without_payment_features\": second_report,\n        \"without_payment_features_reduced\": third_report,\n    }\n)\ncomparison_reduced.metrics.summarize().frame()"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {},
      "source": [
        "# Conclusion\n\nSKD011 and SKD012 often appear together when one leaky column dominates.\nAudit that column and compare with-and-without it; do not treat SKD012 flags\non a leaky report as a drop list. Once leakage is gone, SKD012 may still\nflag identifiers or weak columns (`Provider_Id`, `Total_Discharges` here)\nwhile `TableVectorizer` also builds weak encoded components. Prune those\ninside the pipeline, for example with\n:class:`~sklearn.feature_selection.SelectFromModel`, and check that test\nperformance is preserved.\n\n"
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.14.8"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}