.. DO NOT EDIT. .. THIS FILE WAS AUTOMATICALLY GENERATED BY SPHINX-GALLERY. .. TO MAKE CHANGES, EDIT THE SOURCE PYTHON FILE: .. "auto_examples/pitfalls_and_solutions/plot_skd011_skd012_golden_useless_features.py" .. LINE NUMBERS ARE GIVEN BELOW. .. only:: html .. note:: :class: sphx-glr-download-link-note :ref:`Go to the end ` to download the full example code. .. rst-class:: sphx-glr-example-title .. _sphx_glr_auto_examples_pitfalls_and_solutions_plot_skd011_skd012_golden_useless_features.py: .. _example_skd011_golden_feature_skd012_useless_features: SKD011 & SKD012 - Golden feature and useless features ===================================================== This example walks through mitigations when checks :ref:`SKD011 ` and :ref:`SKD012 ` fire together: a common pattern when a leaky column carries almost all signal and remaining inputs look negligible by comparison. Mitigations from the :ref:`automated_checks` user guide, in the order we try them here: **SKD011 - golden feature** - audit the suspect feature for leakage (is it derived from the target or from data that would not be available at inference time?), - decide whether to keep or drop it, - collect or engineer additional features so the model is less dependent on a single one. **SKD012 - useless features** - review the flagged features and consider dropping them, - refit on a reduced feature set and verify performance is preserved, - if a flagged feature should matter, investigate encoding (here: feature engineering before pruning). .. GENERATED FROM PYTHON SOURCE LINES 33-39 Load the medical charge dataset =============================== Let's load the `medical_charge` dataset, which we will use to predict the average cost of an inpatient stay from the hospital's location and the type of medical procedure. .. GENERATED FROM PYTHON SOURCE LINES 39-45 .. code-block:: Python from skrub.datasets import fetch_medical_charge dataset = fetch_medical_charge() X_full, y_full = dataset.X, dataset.y .. GENERATED FROM PYTHON SOURCE LINES 46-47 Then we can inspect predictors and target with :class:`~skrub.TableReport`. .. GENERATED FROM PYTHON SOURCE LINES 47-52 .. code-block:: Python from skrub import TableReport TableReport(X_full) .. raw:: html

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 53-55 .. code-block:: Python TableReport(y_full) .. raw:: html

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 56-57 Next, we subsample 3,000 rows to make the example run faster. .. GENERATED FROM PYTHON SOURCE LINES 57-61 .. code-block:: Python X = X_full.sample(3_000, random_state=42).reset_index(drop=True) y = y_full.sample(3_000, random_state=42).reset_index(drop=True) .. GENERATED FROM PYTHON SOURCE LINES 62-67 We now create a splitter, vectorizer, and regressor that we reuse throughout the example. :func:`~skore.evaluate` clones them before fitting. High-cardinality strings are encoded with :class:`~skrub.StringEncoder` (TF-IDF then randomized SVD); we seed it so reruns stay comparable. .. GENERATED FROM PYTHON SOURCE LINES 67-77 .. code-block:: Python from sklearn.ensemble import HistGradientBoostingRegressor from sklearn.pipeline import make_pipeline from skore import TrainTestSplit from skrub import StringEncoder, TableVectorizer splitter = TrainTestSplit(random_state=42, test_size=0.2) vectorizer = TableVectorizer(high_cardinality=StringEncoder(random_state=42)) regressor = HistGradientBoostingRegressor(random_state=42) .. GENERATED FROM PYTHON SOURCE LINES 78-83 Trigger SKD011 and SKD012 ========================= Let us train a gradient boosting model, using skrub's `TableVectorizer` to vectorize the data first. .. GENERATED FROM PYTHON SOURCE LINES 83-96 .. code-block:: Python from skore import evaluate model = make_pipeline(vectorizer, regressor) first_report = evaluate( model, X=X, y=y, splitter=splitter, ) first_report .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(high_cardinality=StringEncoder(random_state=42))),
                    ('histgradientboostingregressor',
                     HistGradientBoostingRegressor(random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 97-101 We can see from the metrics table that the models performs very well on the data, with an R² of 0.96. However, looking at the checks, we can see in the Tips tab that SKD011 triggers on the `Average_Medicare_Payments` column. .. GENERATED FROM PYTHON SOURCE LINES 101-104 .. code-block:: Python first_report.checks.summarize() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 105-120 We should then question whether this column would be available in a real deployment. If it would, then we have found a good proxy for the target and we can move on. If it would not, then we should not use this column to train our model. In our case, the target (`Average_Total_Payments`) and the golden feature (`Average_Medicare_Payments`) come from the same process (billing aggregates). Therefore, `Average_Medicare_Payments` would not be available in a real deployment, and we should not use it to train our model. As a matter of fact, `Average_Covered_Charges` would also not be able in a deployment setting so we will also drop it. A final thing to be careful about is that SKD012 flags most provider fields and `Total_Discharges` as useless. As we can see in the importance plot below, every feature looks weak next to the golden billing column, even though some of them may recover once we drop it. We will keep them for now. .. GENERATED FROM PYTHON SOURCE LINES 120-154 .. code-block:: Python _ = first_report.inspection.permutation_importance().plot() # SKD011 - Inspect after dropping golden feature # ============================================== # # First let's evaluate the same model on the same test set with the payment features # removed, and compare it with our original model. # # Thankfully there is still signal in the remaining features, as we get a decent R² # of 0.89. from skore import compare X_without_payment = X.drop( columns=["Average_Medicare_Payments", "Average_Covered_Charges"] ) second_report = evaluate( model, X=X_without_payment, y=y, splitter=splitter, ) comparison = compare( { "with_payment_features": first_report, "without_payment_features": second_report, } ) comparison.metrics.summarize().frame() .. image-sg:: /auto_examples/pitfalls_and_solutions/images/sphx_glr_plot_skd011_skd012_golden_useless_features_001.png :alt: Permutation importance of HistGradientBoostingRegressor on test set :srcset: /auto_examples/pitfalls_and_solutions/images/sphx_glr_plot_skd011_skd012_golden_useless_features_001.png :class: sphx-glr-single-img .. raw:: html
estimator with_payment_features without_payment_features
metric
r2 0.968625 0.890437
rmse 1236.859022 2311.336441
mae 555.159914 1318.776062
mape 0.052583 0.138348
fit_time 2.457898 2.374167
predict_time 0.232536 0.224389


.. GENERATED FROM PYTHON SOURCE LINES 155-159 Let us inspect which features are now the ones our model relies on the most. We can see that `DRG_Definition` is by far most important one, according to permutation importance. This feature encodes the medical reason for which patients were treated. It will therefore be available at inference time and we can use it. .. GENERATED FROM PYTHON SOURCE LINES 159-162 .. code-block:: Python _ = second_report.inspection.permutation_importance().plot() .. image-sg:: /auto_examples/pitfalls_and_solutions/images/sphx_glr_plot_skd011_skd012_golden_useless_features_002.png :alt: Permutation importance of HistGradientBoostingRegressor on test set :srcset: /auto_examples/pitfalls_and_solutions/images/sphx_glr_plot_skd011_skd012_golden_useless_features_002.png :class: sphx-glr-single-img .. GENERATED FROM PYTHON SOURCE LINES 163-171 SKD012 now flags `Provider_Id` and `Total_Discharges`. Those look genuinely weak once billing leakage is gone: an identifier and a discharge count that barely moves the score. Other provider fields recovered some signal, which is why we did not treat the leaky SKD012 list as columns to drop. The check still only sees original columns. `TableVectorizer` creates extra components from high-cardinality strings that can be weak too; we prune those next. .. GENERATED FROM PYTHON SOURCE LINES 171-174 .. code-block:: Python second_report.checks.summarize() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 175-189 SKD012 - prune weak features inside the pipeline ================================================ SKD012 flagged `Provider_Id` and `Total_Discharges` on the original columns. It does not see the extra components `TableVectorizer` builds from high-cardinality fields such as `DRG_Definition` and provider names; some of those add little once the useful ones are present. We therefore select *after* vectorizing, with :class:`~sklearn.feature_selection.SelectFromModel`, so the choice is fit on the training fold only. Histogram gradient boosting has no `feature_importances_`, so the selector uses :class:`~sklearn.ensemble.GradientBoostingRegressor`. The final estimator stays a :class:`~sklearn.ensemble.HistGradientBoostingRegressor` on the reduced matrix. .. GENERATED FROM PYTHON SOURCE LINES 189-203 .. code-block:: Python from sklearn.ensemble import GradientBoostingRegressor from sklearn.feature_selection import SelectFromModel model_reduced = make_pipeline( vectorizer, SelectFromModel( GradientBoostingRegressor(random_state=42), threshold="median", ), regressor, ) model_reduced .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(high_cardinality=StringEncoder(random_state=42))),
                    ('selectfrommodel',
                     SelectFromModel(estimator=GradientBoostingRegressor(random_state=42),
                                     threshold='median')),
                    ('histgradientboostingregressor',
                     HistGradientBoostingRegressor(random_state=42))])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.


.. GENERATED FROM PYTHON SOURCE LINES 204-207 Let us evaluate the reduced pipeline on the same split. Test scores are a on par, and the model is lighter: selection dropped the weak vectorized components before fitting histogram boosting. .. GENERATED FROM PYTHON SOURCE LINES 207-223 .. code-block:: Python third_report = evaluate( model_reduced, X=X_without_payment, y=y, splitter=splitter, ) comparison_reduced = compare( { "without_payment_features": second_report, "without_payment_features_reduced": third_report, } ) comparison_reduced.metrics.summarize().frame() .. raw:: html
estimator without_payment_features without_payment_features_reduced
metric
r2 0.890437 0.890613
rmse 2311.336441 2309.483190
mae 1318.776062 1316.978826
mape 0.138348 0.138346
fit_time 2.374167 7.935268
predict_time 0.224389 0.134846


.. GENERATED FROM PYTHON SOURCE LINES 224-235 Conclusion ========== SKD011 and SKD012 often appear together when one leaky column dominates. Audit that column and compare with-and-without it; do not treat SKD012 flags on a leaky report as a drop list. Once leakage is gone, SKD012 may still flag identifiers or weak columns (`Provider_Id`, `Total_Discharges` here) while `TableVectorizer` also builds weak encoded components. Prune those inside the pipeline, for example with :class:`~sklearn.feature_selection.SelectFromModel`, and check that test performance is preserved. .. rst-class:: sphx-glr-timing **Total running time of the script:** (2 minutes 17.182 seconds) .. _sphx_glr_download_auto_examples_pitfalls_and_solutions_plot_skd011_skd012_golden_useless_features.py: .. only:: html .. container:: sphx-glr-footer sphx-glr-footer-example .. container:: sphx-glr-download sphx-glr-download-jupyter :download:`Download Jupyter notebook: plot_skd011_skd012_golden_useless_features.ipynb ` .. container:: sphx-glr-download sphx-glr-download-python :download:`Download Python source code: plot_skd011_skd012_golden_useless_features.py ` .. container:: sphx-glr-download sphx-glr-download-zip :download:`Download zipped: plot_skd011_skd012_golden_useless_features.zip ` .. only:: html .. rst-class:: sphx-glr-signature `Gallery generated by Sphinx-Gallery `_