Note
Go to the end to download the full example code.
SKD006 - Coefficient interpretation#
SKD006 is a tip about interpretation. This check flags that on mixed-scale features, raw coefficient magnitudes are not directly comparable across columns, and after standardization they are comparable but no longer expressed in the original feature units. Which path you take depends on what you want to read from the coefficients. See also scikit-learn’s example on linear model coefficient interpretation.
When reading coefficients under SKD006, think of:
standardizing the inputs when you want coefficients that are comparable across features (see scale matters),
multiplying each coefficient by the feature’s standard deviation for an “effect per one std” without refitting,
relying on a scale-invariant ranking such as permutation importance when you need importance that does not depend on coefficient units.
We use California housing regression with
Ridge on raw numeric columns. The goal is to
interpret effect sizes without mistaking scale differences for importance.
Load the California housing dataset#
Median income is measured in 10k USD blocks, latitude in degrees, and population in head counts. Fitting Ridge on these raw columns produces coefficients whose magnitudes reflect units as much as predictive strength.
from skrub.datasets import fetch_california_housing
housing = fetch_california_housing()
X, y = housing.X, housing.y
TableReport highlights the range differences across columns.
from skrub import TableReport
TableReport(X)
| MedInc | HouseAge | AveRooms | AveBedrms | Population | AveOccup | Latitude | Longitude | |
|---|---|---|---|---|---|---|---|---|
| 0 | 8.33 | 41.0 | 6.98 | 1.02 | 322. | 2.56 | 37.9 | -122. |
| 1 | 8.30 | 21.0 | 6.24 | 0.972 | 2.40e+03 | 2.11 | 37.9 | -122. |
| 2 | 7.26 | 52.0 | 8.29 | 1.07 | 496. | 2.80 | 37.9 | -122. |
| 3 | 5.64 | 52.0 | 5.82 | 1.07 | 558. | 2.55 | 37.9 | -122. |
| 4 | 3.85 | 52.0 | 6.28 | 1.08 | 565. | 2.18 | 37.9 | -122. |
| 20,635 | 1.56 | 25.0 | 5.05 | 1.13 | 845. | 2.56 | 39.5 | -121. |
| 20,636 | 2.56 | 18.0 | 6.11 | 1.32 | 356. | 3.12 | 39.5 | -121. |
| 20,637 | 1.70 | 17.0 | 5.21 | 1.12 | 1.01e+03 | 2.33 | 39.4 | -121. |
| 20,638 | 1.87 | 18.0 | 5.33 | 1.17 | 741. | 2.12 | 39.4 | -121. |
| 20,639 | 2.39 | 16.0 | 5.25 | 1.16 | 1.39e+03 | 2.62 | 39.4 | -121. |
MedInc
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
12,928 (62.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.87 ± 1.90
- Median ± IQR
- 3.53 ± 2.18
- Min | Max
- 0.500 | 15.0
HouseAge
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
52 (0.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 28.6 ± 12.6
- Median ± IQR
- 29.0 ± 19.0
- Min | Max
- 1.00 | 52.0
AveRooms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
19,392 (94.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 5.43 ± 2.47
- Median ± IQR
- 5.23 ± 1.61
- Min | Max
- 0.846 | 142.
AveBedrms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
14,233 (69.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.10 ± 0.474
- Median ± IQR
- 1.05 ± 0.0934
- Min | Max
- 0.333 | 34.1
Population
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,888 (18.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.43e+03 ± 1.13e+03
- Median ± IQR
- 1.17e+03 ± 938.
- Min | Max
- 3.00 | 3.57e+04
AveOccup
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
18,841 (91.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.07 ± 10.4
- Median ± IQR
- 2.82 ± 0.852
- Min | Max
- 0.692 | 1.24e+03
Latitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
862 (4.2%)
This column has a high cardinality (> 40).
- Mean ± Std
- 35.6 ± 2.14
- Median ± IQR
- 34.3 ± 3.78
- Min | Max
- 32.5 | 42.0
Longitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
844 (4.1%)
This column has a high cardinality (> 40).
- Mean ± Std
- -120. ± 2.00
- Median ± IQR
- -118. ± 3.79
- Min | Max
- -124. | -114.
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | MedInc | Float64DType | False | 0 (0.0%) | 12928 (62.6%) | 3.87 | 1.90 | 0.500 | 3.53 | 15.0 |
| 1 | HouseAge | Float64DType | False | 0 (0.0%) | 52 (0.3%) | 28.6 | 12.6 | 1.00 | 29.0 | 52.0 |
| 2 | AveRooms | Float64DType | False | 0 (0.0%) | 19392 (94.0%) | 5.43 | 2.47 | 0.846 | 5.23 | 142. |
| 3 | AveBedrms | Float64DType | False | 0 (0.0%) | 14233 (69.0%) | 1.10 | 0.474 | 0.333 | 1.05 | 34.1 |
| 4 | Population | Float64DType | False | 0 (0.0%) | 3888 (18.8%) | 1.43e+03 | 1.13e+03 | 3.00 | 1.17e+03 | 3.57e+04 |
| 5 | AveOccup | Float64DType | False | 0 (0.0%) | 18841 (91.3%) | 3.07 | 10.4 | 0.692 | 2.82 | 1.24e+03 |
| 6 | Latitude | Float64DType | False | 0 (0.0%) | 862 (4.2%) | 35.6 | 2.14 | 32.5 | 34.3 | 42.0 |
| 7 | Longitude | Float64DType | False | 0 (0.0%) | 844 (4.1%) | -120. | 2.00 | -124. | -118. | -114. |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
MedInc
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
12,928 (62.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.87 ± 1.90
- Median ± IQR
- 3.53 ± 2.18
- Min | Max
- 0.500 | 15.0
HouseAge
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
52 (0.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 28.6 ± 12.6
- Median ± IQR
- 29.0 ± 19.0
- Min | Max
- 1.00 | 52.0
AveRooms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
19,392 (94.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 5.43 ± 2.47
- Median ± IQR
- 5.23 ± 1.61
- Min | Max
- 0.846 | 142.
AveBedrms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
14,233 (69.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.10 ± 0.474
- Median ± IQR
- 1.05 ± 0.0934
- Min | Max
- 0.333 | 34.1
Population
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,888 (18.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.43e+03 ± 1.13e+03
- Median ± IQR
- 1.17e+03 ± 938.
- Min | Max
- 3.00 | 3.57e+04
AveOccup
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
18,841 (91.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.07 ± 10.4
- Median ± IQR
- 2.82 ± 0.852
- Min | Max
- 0.692 | 1.24e+03
Latitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
862 (4.2%)
This column has a high cardinality (> 40).
- Mean ± Std
- 35.6 ± 2.14
- Median ± IQR
- 34.3 ± 3.78
- Min | Max
- 32.5 | 42.0
Longitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
844 (4.1%)
This column has a high cardinality (> 40).
- Mean ± Std
- -120. ± 2.00
- Median ± IQR
- -118. ± 3.79
- Min | Max
- -124. | -114.
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
| Column 1 | Column 2 | Cramér's V | Pearson's Correlation |
|---|---|---|---|
| AveRooms | AveBedrms | 0.769 | 0.811 |
| Latitude | Longitude | 0.503 | -0.925 |
| Population | AveOccup | 0.316 | 0.111 |
| MedInc | AveOccup | 0.267 | 0.0593 |
| MedInc | AveRooms | 0.194 | 0.354 |
| HouseAge | Longitude | 0.173 | -0.0958 |
| AveBedrms | Longitude | 0.134 | 0.0123 |
| HouseAge | Latitude | 0.128 | -0.000739 |
| HouseAge | Population | 0.116 | -0.270 |
| AveRooms | Longitude | 0.0954 | -0.0316 |
| MedInc | Latitude | 0.0944 | -0.0783 |
| Population | Longitude | 0.0876 | 0.0923 |
| MedInc | Longitude | 0.0874 | -0.0110 |
| HouseAge | AveRooms | 0.0858 | -0.186 |
| AveBedrms | Latitude | 0.0819 | 0.0769 |
| MedInc | HouseAge | 0.0771 | -0.121 |
| AveRooms | Latitude | 0.0757 | 0.117 |
| HouseAge | AveBedrms | 0.0745 | -0.117 |
| MedInc | Population | 0.0642 | 0.0112 |
| HouseAge | AveOccup | 0.0566 | 0.0222 |
| Population | Latitude | 0.0491 | -0.0977 |
| AveRooms | Population | 0.0401 | -0.0969 |
| AveOccup | Latitude | 0.0375 | 0.0173 |
| AveOccup | Longitude | 0.0375 | -0.0159 |
| MedInc | AveBedrms | 0.0349 | -0.0779 |
| AveBedrms | Population | 0.0233 | -0.0957 |
| AveRooms | AveOccup | 0.00558 | -0.0214 |
| AveBedrms | AveOccup | 0.00261 | -0.0147 |
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
| MedHouseVal | |
|---|---|
| 0 | 4.53 |
| 1 | 3.58 |
| 2 | 3.52 |
| 3 | 3.41 |
| 4 | 3.42 |
| 20,635 | 0.781 |
| 20,636 | 0.771 |
| 20,637 | 0.923 |
| 20,638 | 0.847 |
| 20,639 | 0.894 |
MedHouseVal
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,842 (18.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.07 ± 1.15
- Median ± IQR
- 1.80 ± 1.45
- Min | Max
- 0.150 | 5.00
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | MedHouseVal | Float64DType | False | 0 (0.0%) | 3842 (18.6%) | 2.07 | 1.15 | 0.150 | 1.80 | 5.00 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
MedHouseVal
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,842 (18.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.07 ± 1.15
- Median ± IQR
- 1.80 ± 1.45
- Min | Max
- 0.150 | 5.00
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
A single TrainTestSplit with a 20 % test set feeds every
evaluation below.
from skore import TrainTestSplit
splitter = TrainTestSplit(test_size=0.2, random_state=42)
Trigger SKD006 with Ridge on raw mixed-scale features#
Ridge with a moderate alpha fits quickly on unscaled inputs. SKD006
indicates that the features are not on the same scale, so raw coefficient
magnitudes should not be read as a ranking of importance.
| Metric | Ridge |
|---|---|
| R² | 0.575855 |
| RMSE | 0.745522 |
| MAE | 0.533204 |
| MAPE | 0.319523 |
| Fit time (s) | 0.003538 |
| Predict time (s) | 0.001179 |
No issues were detected in your report.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Ridge.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Ridge.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Ridge()In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
| MedInc | HouseAge | AveRooms | AveBedrms | Population | AveOccup | Latitude | Longitude | MedHouseVal | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 3.26 | 33.0 | 5.02 | 1.01 | 2.30e+03 | 3.69 | 32.7 | -117. | 1.03 |
| 1 | 3.81 | 49.0 | 4.47 | 1.04 | 1.31e+03 | 1.74 | 33.8 | -118. | 3.82 |
| 2 | 4.16 | 4.00 | 5.65 | 0.985 | 915. | 2.72 | 34.7 | -120. | 1.73 |
| 3 | 1.94 | 36.0 | 4.00 | 1.03 | 1.42e+03 | 3.99 | 32.7 | -117. | 0.934 |
| 4 | 3.55 | 43.0 | 6.27 | 1.13 | 874. | 2.30 | 36.8 | -120. | 0.965 |
| 20,635 | 4.61 | 16.0 | 7.00 | 1.07 | 1.35e+03 | 2.99 | 33.4 | -117. | 2.63 |
| 20,636 | 2.73 | 28.0 | 6.13 | 1.26 | 1.65e+03 | 2.34 | 35.4 | -121. | 2.67 |
| 20,637 | 9.23 | 25.0 | 7.24 | 0.947 | 1.58e+03 | 2.79 | 37.3 | -122. | 5.00 |
| 20,638 | 2.79 | 36.0 | 5.29 | 0.983 | 1.23e+03 | 2.59 | 36.8 | -120. | 0.723 |
| 20,639 | 3.55 | 17.0 | 3.99 | 1.03 | 1.67e+03 | 3.73 | 34.2 | -118. | 1.51 |
MedInc
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
12,928 (62.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.87 ± 1.90
- Median ± IQR
- 3.53 ± 2.18
- Min | Max
- 0.500 | 15.0
HouseAge
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
52 (0.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 28.6 ± 12.6
- Median ± IQR
- 29.0 ± 19.0
- Min | Max
- 1.00 | 52.0
AveRooms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
19,392 (94.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 5.43 ± 2.47
- Median ± IQR
- 5.23 ± 1.61
- Min | Max
- 0.846 | 142.
AveBedrms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
14,233 (69.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.10 ± 0.474
- Median ± IQR
- 1.05 ± 0.0934
- Min | Max
- 0.333 | 34.1
Population
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,888 (18.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.43e+03 ± 1.13e+03
- Median ± IQR
- 1.17e+03 ± 938.
- Min | Max
- 3.00 | 3.57e+04
AveOccup
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
18,841 (91.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.07 ± 10.4
- Median ± IQR
- 2.82 ± 0.852
- Min | Max
- 0.692 | 1.24e+03
Latitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
862 (4.2%)
This column has a high cardinality (> 40).
- Mean ± Std
- 35.6 ± 2.14
- Median ± IQR
- 34.3 ± 3.78
- Min | Max
- 32.5 | 42.0
Longitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
844 (4.1%)
This column has a high cardinality (> 40).
- Mean ± Std
- -120. ± 2.00
- Median ± IQR
- -118. ± 3.79
- Min | Max
- -124. | -114.
MedHouseVal
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,842 (18.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.07 ± 1.15
- Median ± IQR
- 1.80 ± 1.45
- Min | Max
- 0.150 | 5.00
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | MedInc | Float64DType | False | 0 (0.0%) | 12928 (62.6%) | 3.87 | 1.90 | 0.500 | 3.53 | 15.0 |
| 1 | HouseAge | Float64DType | False | 0 (0.0%) | 52 (0.3%) | 28.6 | 12.6 | 1.00 | 29.0 | 52.0 |
| 2 | AveRooms | Float64DType | False | 0 (0.0%) | 19392 (94.0%) | 5.43 | 2.47 | 0.846 | 5.23 | 142. |
| 3 | AveBedrms | Float64DType | False | 0 (0.0%) | 14233 (69.0%) | 1.10 | 0.474 | 0.333 | 1.05 | 34.1 |
| 4 | Population | Float64DType | False | 0 (0.0%) | 3888 (18.8%) | 1.43e+03 | 1.13e+03 | 3.00 | 1.17e+03 | 3.57e+04 |
| 5 | AveOccup | Float64DType | False | 0 (0.0%) | 18841 (91.3%) | 3.07 | 10.4 | 0.692 | 2.82 | 1.24e+03 |
| 6 | Latitude | Float64DType | False | 0 (0.0%) | 862 (4.2%) | 35.6 | 2.14 | 32.5 | 34.3 | 42.0 |
| 7 | Longitude | Float64DType | False | 0 (0.0%) | 844 (4.1%) | -120. | 2.00 | -124. | -118. | -114. |
| 8 | MedHouseVal | Float64DType | False | 0 (0.0%) | 3842 (18.6%) | 2.07 | 1.15 | 0.150 | 1.80 | 5.00 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
Find SKD006 in the Tips tab below.
report.checks.summarize()
No issues were detected in your report.
- [SKD006] Coefficient interpretation. Features are not on the same scale: coefficient magnitudes are not directly comparable as feature importance.
- [SKD009] Model performance vs. HistGradientBoosting baseline. Test scores are significantly worse than a HistGradientBoosting baseline for 4/4 default predictive metrics. Baseline performance on the test set: MAE=0.31, MAPE=0.18, RMSE=0.464, R²=0.835.
- [SKD012] Useless features. Feature(s) ['Population'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['alpha'] for Ridge.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Ridge.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Ridge.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Let us inspect coefficients with
coefficients(), and
compare each coefficient’s magnitude to the feature’s typical range (or
standard deviation): a large coefficient on a small-scale column is not
necessarily more important than a small coefficient on a large-scale column
such as Population.
Unscaled coefficients are still useful on their own: they answer “if I change this feature by one unit in its original scale, how much does the prediction change in target units?” They are misleading regarding feature importance.
coef_display = report.inspection.coefficients()
coef_display.frame()
| feature | coefficient | |
|---|---|---|
| 0 | Intercept | -37.019420 |
| 1 | MedInc | 0.448511 |
| 2 | HouseAge | 0.009726 |
| 3 | AveRooms | -0.123014 |
| 4 | AveBedrms | 0.781417 |
| 5 | Population | -0.000002 |
| 6 | AveOccup | -0.003526 |
| 7 | Latitude | -0.419787 |
| 8 | Longitude | -0.433681 |
Side by side, the raw coefficient magnitudes and the feature standard deviations tell
different stories. For instance, Population has a tiny coefficient because a
one-person change is negligible on a head-count scale, even though the column varies a
lot across districts.
import pandas as pd
coef = coef_display.frame(include_intercept=False).set_index("feature")
feature_std = report.X_train.std().rename("feature std. dev. (train)")
df = pd.concat([coef, feature_std], axis=1)
axes = df.plot.barh(
subplots=True,
layout=(1, 2),
legend=False,
sharey=True,
sharex=False,
figsize=(10, 4),
)
axes[0, 0].set_xlabel("coefficient")
_ = axes[0, 1].set_xlabel("std. dev.")
Scale coefficients by feature standard deviation#
Recall the pitfall above: when the model was trained on features with different dynamic ranges, you cannot compare features by looking at raw coefficients alone. Multiplying each fitted coefficient by its feature standard deviation gives an “effect per one standard deviation” that is comparable across columns without refitting. See scikit-learn’s discussion of scale matters.
comparable = coef.copy()
comparable["effect per std. dev."] = comparable["coefficient"] * feature_std
_ = comparable.plot.barh()

That rescaling erases the surprises from the side-by-side bars above. The
large raw AveBedrms coefficient shrinks once multiplied by its small
standard deviation, so it drops in the ranking while other features become, in
comparison, more important.
Standardize inputs in a pipeline#
StandardScaler inside a
Pipeline puts every feature on a common scale at
fit time, so coefficients become directly comparable as “effect per one
standard deviation.” That is useful for ranking features, but you lose
statements in original units (for example, “one extra year of age”).
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
report_scaled = evaluate(
Pipeline([("scale", StandardScaler()), ("ridge", Ridge(alpha=1.0))]),
X=X,
y=y,
splitter=splitter,
)
report_scaled
| Metric | Ridge |
|---|---|
| R² | 0.575816 |
| RMSE | 0.745557 |
| MAE | 0.533193 |
| MAPE | 0.319512 |
| Fit time (s) | 0.005895 |
| Predict time (s) | 0.001638 |
No issues were detected in your report.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('scale', StandardScaler()), ('ridge', Ridge())])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
| Name | Type | Value |
|---|---|---|
|
feature_names_in_
feature_names_in_: ndarray of shape (`n_features_in_`,) Names of features seen during :term:`fit`. Only defined if the underlying estimator exposes such an attribute when fit. .. versionadded:: 1.0 |
ndarray[object](8,) | ['MedInc','HouseAge','AveRooms',...,'AveOccup','Latitude','Longitude'] |
|
n_features_in_
n_features_in_: int Number of features seen during :term:`fit`. Only defined if the underlying first estimator in `steps` exposes such an attribute when fit. .. versionadded:: 0.24 |
int | 8 |
Parameters
Fitted attributes
8 features
| MedInc |
| HouseAge |
| AveRooms |
| AveBedrms |
| Population |
| AveOccup |
| Latitude |
| Longitude |
Parameters
Fitted attributes
| MedInc | HouseAge | AveRooms | AveBedrms | Population | AveOccup | Latitude | Longitude | MedHouseVal | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | 3.26 | 33.0 | 5.02 | 1.01 | 2.30e+03 | 3.69 | 32.7 | -117. | 1.03 |
| 1 | 3.81 | 49.0 | 4.47 | 1.04 | 1.31e+03 | 1.74 | 33.8 | -118. | 3.82 |
| 2 | 4.16 | 4.00 | 5.65 | 0.985 | 915. | 2.72 | 34.7 | -120. | 1.73 |
| 3 | 1.94 | 36.0 | 4.00 | 1.03 | 1.42e+03 | 3.99 | 32.7 | -117. | 0.934 |
| 4 | 3.55 | 43.0 | 6.27 | 1.13 | 874. | 2.30 | 36.8 | -120. | 0.965 |
| 20,635 | 4.61 | 16.0 | 7.00 | 1.07 | 1.35e+03 | 2.99 | 33.4 | -117. | 2.63 |
| 20,636 | 2.73 | 28.0 | 6.13 | 1.26 | 1.65e+03 | 2.34 | 35.4 | -121. | 2.67 |
| 20,637 | 9.23 | 25.0 | 7.24 | 0.947 | 1.58e+03 | 2.79 | 37.3 | -122. | 5.00 |
| 20,638 | 2.79 | 36.0 | 5.29 | 0.983 | 1.23e+03 | 2.59 | 36.8 | -120. | 0.723 |
| 20,639 | 3.55 | 17.0 | 3.99 | 1.03 | 1.67e+03 | 3.73 | 34.2 | -118. | 1.51 |
MedInc
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
12,928 (62.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.87 ± 1.90
- Median ± IQR
- 3.53 ± 2.18
- Min | Max
- 0.500 | 15.0
HouseAge
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
52 (0.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 28.6 ± 12.6
- Median ± IQR
- 29.0 ± 19.0
- Min | Max
- 1.00 | 52.0
AveRooms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
19,392 (94.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 5.43 ± 2.47
- Median ± IQR
- 5.23 ± 1.61
- Min | Max
- 0.846 | 142.
AveBedrms
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
14,233 (69.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.10 ± 0.474
- Median ± IQR
- 1.05 ± 0.0934
- Min | Max
- 0.333 | 34.1
Population
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,888 (18.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 1.43e+03 ± 1.13e+03
- Median ± IQR
- 1.17e+03 ± 938.
- Min | Max
- 3.00 | 3.57e+04
AveOccup
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
18,841 (91.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.07 ± 10.4
- Median ± IQR
- 2.82 ± 0.852
- Min | Max
- 0.692 | 1.24e+03
Latitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
862 (4.2%)
This column has a high cardinality (> 40).
- Mean ± Std
- 35.6 ± 2.14
- Median ± IQR
- 34.3 ± 3.78
- Min | Max
- 32.5 | 42.0
Longitude
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
844 (4.1%)
This column has a high cardinality (> 40).
- Mean ± Std
- -120. ± 2.00
- Median ± IQR
- -118. ± 3.79
- Min | Max
- -124. | -114.
MedHouseVal
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,842 (18.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.07 ± 1.15
- Median ± IQR
- 1.80 ± 1.45
- Min | Max
- 0.150 | 5.00
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | MedInc | Float64DType | False | 0 (0.0%) | 12928 (62.6%) | 3.87 | 1.90 | 0.500 | 3.53 | 15.0 |
| 1 | HouseAge | Float64DType | False | 0 (0.0%) | 52 (0.3%) | 28.6 | 12.6 | 1.00 | 29.0 | 52.0 |
| 2 | AveRooms | Float64DType | False | 0 (0.0%) | 19392 (94.0%) | 5.43 | 2.47 | 0.846 | 5.23 | 142. |
| 3 | AveBedrms | Float64DType | False | 0 (0.0%) | 14233 (69.0%) | 1.10 | 0.474 | 0.333 | 1.05 | 34.1 |
| 4 | Population | Float64DType | False | 0 (0.0%) | 3888 (18.8%) | 1.43e+03 | 1.13e+03 | 3.00 | 1.17e+03 | 3.57e+04 |
| 5 | AveOccup | Float64DType | False | 0 (0.0%) | 18841 (91.3%) | 3.07 | 10.4 | 0.692 | 2.82 | 1.24e+03 |
| 6 | Latitude | Float64DType | False | 0 (0.0%) | 862 (4.2%) | 35.6 | 2.14 | 32.5 | 34.3 | 42.0 |
| 7 | Longitude | Float64DType | False | 0 (0.0%) | 844 (4.1%) | -120. | 2.00 | -124. | -118. | -114. |
| 8 | MedHouseVal | Float64DType | False | 0 (0.0%) | 3842 (18.6%) | 2.07 | 1.15 | 0.150 | 1.80 | 5.00 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
SKD006 is still reported in the Tips tab but the warning message changed.
report_scaled.checks.summarize()
No issues were detected in your report.
- [SKD006] Coefficient interpretation. Features appear to be standardized: coefficients are comparable but no longer interpretable in the original feature units.
- [SKD009] Model performance vs. HistGradientBoosting baseline. Test scores are significantly worse than a HistGradientBoosting baseline for 4/4 default predictive metrics. Baseline performance on the test set: MAE=0.31, MAPE=0.18, RMSE=0.464, R²=0.836.
- [SKD012] Useless features. Feature(s) ['Population'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['alpha'] for Ridge.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
The coefficient table is now on a shared scale. Read magnitudes as relative importance after standardization, not as effects per original unit of income, rooms, or population.
These values are the same quantity as the post-hoc
coefficient * feature_std column from the previous section: both are an
effect per one standard deviation. Fitting with StandardScaler does that
rescaling inside the pipeline; multiplying raw coefficients by the train-set
std does it after the fact. Small numerical differences can remain because the
scaler’s internal mean/std and X_train.std() are not bit-identical, but the
ranking and interpretation should match.
pd.concat(
[
report_scaled.inspection.coefficients()
.frame(include_intercept=False)
.set_index("feature"),
comparable["effect per std. dev."],
],
axis=1,
)
| coefficient | effect per std. dev. | |
|---|---|---|
| feature | ||
| MedInc | 0.854327 | 0.854097 |
| HouseAge | 0.122624 | 0.122571 |
| AveRooms | -0.294210 | -0.293681 |
| AveBedrms | 0.339008 | 0.338521 |
| Population | -0.002282 | -0.002303 |
| AveOccup | -0.040833 | -0.040825 |
| Latitude | -0.896168 | -0.896944 |
| Longitude | -0.869071 | -0.869813 |
Permutation importance (scale-invariant)#
Scaled coefficients are one way to rank features. Another is permutation importance on held-out data: shuffle one column at a time and measure how much the score drops. That answers questions such as “which features did the model rely on most?” and works for nonlinear models as well. Present it as an alternative to scaled coefficients when you care about predictive reliance rather than a linear effect size.
Correlated features can distort both coefficient rankings and permutation importance (credit is shared or shifted between partners). See the SKD008 example and scikit-learn’s note on misleading values on strongly correlated features.
| data_source | metric | feature | value_mean | value_std | |
|---|---|---|---|---|---|
| 0 | test | r2 | MedInc | 1.038369 | 0.025849 |
| 1 | test | r2 | HouseAge | 0.022603 | 0.002785 |
| 2 | test | r2 | AveRooms | 0.236201 | 0.007711 |
| 3 | test | r2 | AveBedrms | 0.257610 | 0.007691 |
| 4 | test | r2 | Population | 0.000002 | 0.000039 |
| 5 | test | r2 | AveOccup | 0.001057 | 0.000244 |
| 6 | test | r2 | Latitude | 1.207360 | 0.028557 |
| 7 | test | r2 | Longitude | 1.147209 | 0.020633 |

<Figure size 611.111x600 with 1 Axes>
Conclusion#
SKD006 is a reminder to be careful when interpreting coefficients of linear models. When features are on different scales, linear coefficients mix feature importance with feature magnitude; when features are standardized, coefficients are comparable but no longer in original units.
Unscaled coefficients remain useful when the question is “if I change this feature by this much in its own units, how much does the prediction change?” Choose standardization, an effect-per-std rescaling, or a scale-invariant method such as permutation importance depending on the question you want to answer.
Total running time of the script: (0 minutes 4.060 seconds)