Note
Go to the end to download the full example code.
SKD011 & SKD012 - Golden feature and useless features#
This example walks through mitigations when checks SKD011 and SKD012 fire together: a common pattern when a leaky column carries almost all signal and remaining inputs look negligible by comparison.
Mitigations from the Automated checks user guide, in the order we try them here:
SKD011 - golden feature
audit the suspect feature for leakage (is it derived from the target or from data that would not be available at inference time?),
decide whether to keep or drop it,
collect or engineer additional features so the model is less dependent on a single one.
SKD012 - useless features
review the flagged features and consider dropping them,
refit on a reduced feature set and verify performance is preserved,
if a flagged feature should matter, investigate encoding (here: feature engineering before pruning).
Load the medical charge dataset#
Let’s load the medical_charge dataset, which we will use to predict the average
cost of an inpatient stay from the hospital’s location and the type of medical
procedure.
from skrub.datasets import fetch_medical_charge
dataset = fetch_medical_charge()
X_full, y_full = dataset.X, dataset.y
Then we can inspect predictors and target with TableReport.
from skrub import TableReport
TableReport(X_full)
| DRG_Definition | Provider_Id | Provider_Name | Provider_Street_Address | Provider_City | Provider_State | Provider_Zip_Code | Hospital_Referral_Region_(HRR)_Description | Total_Discharges | Average_Covered_Charges | Average_Medicare_Payments | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 039 - EXTRACRANIAL PROCEDURES W/O CC/MCC | 10,001 | SOUTHEAST ALABAMA MEDICAL CENTER | 1108 ROSS CLARK CIRCLE | DOTHAN | AL | 36,301 | AL - Dothan | 91 | 3.30e+04 | 4.76e+03 |
| 1 | 039 - EXTRACRANIAL PROCEDURES W/O CC/MCC | 10,005 | MARSHALL MEDICAL CENTER SOUTH | 2505 U S HIGHWAY 431 NORTH | BOAZ | AL | 35,957 | AL - Birmingham | 14 | 1.51e+04 | 4.98e+03 |
| 2 | 039 - EXTRACRANIAL PROCEDURES W/O CC/MCC | 10,006 | ELIZA COFFEE MEMORIAL HOSPITAL | 205 MARENGO STREET | FLORENCE | AL | 35,631 | AL - Birmingham | 24 | 3.76e+04 | 4.45e+03 |
| 3 | 039 - EXTRACRANIAL PROCEDURES W/O CC/MCC | 10,011 | ST VINCENT'S EAST | 50 MEDICAL PARK EAST DRIVE | BIRMINGHAM | AL | 35,235 | AL - Birmingham | 25 | 1.40e+04 | 4.13e+03 |
| 4 | 039 - EXTRACRANIAL PROCEDURES W/O CC/MCC | 10,016 | SHELBY BAPTIST MEDICAL CENTER | 1000 FIRST STREET NORTH | ALABASTER | AL | 35,007 | AL - Birmingham | 18 | 3.16e+04 | 4.85e+03 |
| 163,060 | 948 - SIGNS & SYMPTOMS W/O MCC | 670,041 | SETON MEDICAL CENTER WILLIAMSON | 201 SETON PARKWAY | ROUND ROCK | TX | 78,664 | TX - Austin | 23 | 2.63e+04 | 3.07e+03 |
| 163,061 | 948 - SIGNS & SYMPTOMS W/O MCC | 670,055 | METHODIST STONE OAK HOSPITAL | 1139 E SONTERRA BLVD | SAN ANTONIO | TX | 78,258 | TX - San Antonio | 11 | 2.17e+04 | 2.65e+03 |
| 163,062 | 948 - SIGNS & SYMPTOMS W/O MCC | 670,056 | SETON MEDICAL CENTER HAYS | 6001 KYLE PKWY | KYLE | TX | 78,640 | TX - Austin | 19 | 3.91e+04 | 4.06e+03 |
| 163,063 | 948 - SIGNS & SYMPTOMS W/O MCC | 670,060 | TEXAS REGIONAL MEDICAL CENTER AT SUNNYVALE | 231 SOUTH COLLINS ROAD | SUNNYVALE | TX | 75,182 | TX - Dallas | 11 | 2.89e+04 | 6.85e+03 |
| 163,064 | 948 - SIGNS & SYMPTOMS W/O MCC | 670,068 | TEXAS HEALTH PRESBYTERIAN HOSPITAL FLOWER MOUND | 4400 LONG PRAIRIE ROAD | FLOWER MOUND | TX | 75,028 | TX - Dallas | 12 | 1.50e+04 | 2.89e+03 |
DRG_Definition
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
100 (< 0.1%)
This column has a high cardinality (> 40).
Most frequent values
194 - SIMPLE PNEUMONIA & PLEURISY W CC
690 - KIDNEY & URINARY TRACT INFECTIONS W/O MCC
292 - HEART FAILURE & SHOCK W CC
392 - ESOPHAGITIS, GASTROENT & MISC DIGEST DISORDERS W/O MCC
641 - MISC DISORDERS OF NUTRITION,METABOLISM,FLUIDS/ELECTROLYTES W/O MCC
871 - SEPTICEMIA OR SEVERE SEPSIS W/O MV 96+ HOURS W MCC
603 - CELLULITIS W/O MCC
470 - MAJOR JOINT REPLACEMENT OR REATTACHMENT OF LOWER EXTREMITY W/O MCC
191 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W CC
List:190 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W MCC
['194 - SIMPLE PNEUMONIA & PLEURISY W CC', '690 - KIDNEY & URINARY TRACT INFECTIONS W/O MCC', '292 - HEART FAILURE & SHOCK W CC', '392 - ESOPHAGITIS, GASTROENT & MISC DIGEST DISORDERS W/O MCC', '641 - MISC DISORDERS OF NUTRITION,METABOLISM,FLUIDS/ELECTROLYTES W/O MCC', '871 - SEPTICEMIA OR SEVERE SEPSIS W/O MV 96+ HOURS W MCC', '603 - CELLULITIS W/O MCC', '470 - MAJOR JOINT REPLACEMENT OR REATTACHMENT OF LOWER EXTREMITY W/O MCC', '191 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W CC', '190 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W MCC']
Provider_Id
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
3,337 (2.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.56e+05 ± 1.52e+05
- Median ± IQR
- 250,007 ± 269,983
- Min | Max
- 10,001 | 670,077
Provider_Name
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
3,201 (2.0%)
This column has a high cardinality (> 40).
Most frequent values
GOOD SAMARITAN HOSPITAL
ST JOSEPH MEDICAL CENTER
MERCY MEDICAL CENTER
MERCY HOSPITAL
ST JOSEPH HOSPITAL
ST FRANCIS MEDICAL CENTER
ST MARY MEDICAL CENTER
ST LUKES HOSPITAL
ST FRANCIS HOSPITAL
List:JEFFERSON REGIONAL MEDICAL CENTER
['GOOD SAMARITAN HOSPITAL', 'ST JOSEPH MEDICAL CENTER', 'MERCY MEDICAL CENTER', 'MERCY HOSPITAL', 'ST JOSEPH HOSPITAL', 'ST FRANCIS MEDICAL CENTER', 'ST MARY MEDICAL CENTER', 'ST LUKES HOSPITAL', 'ST FRANCIS HOSPITAL', 'JEFFERSON REGIONAL MEDICAL CENTER']
Provider_Street_Address
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
3,326 (2.0%)
This column has a high cardinality (> 40).
Most frequent values
100 MEDICAL CENTER DRIVE
800 WASHINGTON STREET
1 MEDICAL CENTER DRIVE
100 HOSPITAL DRIVE
809 UNIVERSITY BOULEVARD EAST
9601 INTERSTATE 630, EXIT 7
1303 E HERNDON AVE
8700 BEVERLY BLVD
20 YORK ST
List:4755 OGLETOWN-STANTON ROAD
['100 MEDICAL CENTER DRIVE', '800 WASHINGTON STREET', '1 MEDICAL CENTER DRIVE', '100 HOSPITAL DRIVE', '809 UNIVERSITY BOULEVARD EAST', '9601 INTERSTATE 630, EXIT 7', '1303 E HERNDON AVE', '8700 BEVERLY BLVD', '20 YORK ST', '4755 OGLETOWN-STANTON ROAD']
Provider_City
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
1,977 (1.2%)
This column has a high cardinality (> 40).
Most frequent values
CHICAGO
BALTIMORE
HOUSTON
PHILADELPHIA
BROOKLYN
SPRINGFIELD
COLUMBUS
LOS ANGELES
NEW YORK
List:DALLAS
['CHICAGO', 'BALTIMORE', 'HOUSTON', 'PHILADELPHIA', 'BROOKLYN', 'SPRINGFIELD', 'COLUMBUS', 'LOS ANGELES', 'NEW YORK', 'DALLAS']
Provider_State
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
51 (< 0.1%)
This column has a high cardinality (> 40).
Most frequent values
CA
TX
FL
NY
IL
PA
OH
MI
NC
List:GA
['CA', 'TX', 'FL', 'NY', 'IL', 'PA', 'OH', 'MI', 'NC', 'GA']
Provider_Zip_Code
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
3,053 (1.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 4.79e+04 ± 2.79e+04
- Median ± IQR
- 44,309 ± 45,640
- Min | Max
- 1,040 | 99,835
Hospital_Referral_Region_(HRR)_Description
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
306 (0.2%)
This column has a high cardinality (> 40).
Most frequent values
CA - Los Angeles
MA - Boston
GA - Atlanta
TX - Houston
PA - Philadelphia
TX - Dallas
MO - St. Louis
NY - East Long Island
FL - Orlando
List:TN - Nashville
['CA - Los Angeles', 'MA - Boston', 'GA - Atlanta', 'TX - Houston', 'PA - Philadelphia', 'TX - Dallas', 'MO - St. Louis', 'NY - East Long Island', 'FL - Orlando', 'TN - Nashville']
Total_Discharges
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
642 (0.4%)
This column has a high cardinality (> 40).
- Mean ± Std
- 42.8 ± 51.1
- Median ± IQR
- 27 ± 32
- Min | Max
- 11 | 3,383
Average_Covered_Charges
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
161,985 (99.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.61e+04 ± 3.51e+04
- Median ± IQR
- 2.52e+04 ± 2.73e+04
- Min | Max
- 2.46e+03 | 9.29e+05
Average_Medicare_Payments
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
157,817 (96.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 8.49e+03 ± 7.31e+03
- Median ± IQR
- 6.16e+03 ± 5.86e+03
- Min | Max
- 1.15e+03 | 1.55e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | DRG_Definition | StringDtype | True | 0 (0.0%) | 100 (< 0.1%) | |||||
| 1 | Provider_Id | Int64DType | False | 0 (0.0%) | 3337 (2.0%) | 2.56e+05 | 1.52e+05 | 10,001 | 250,007 | 670,077 |
| 2 | Provider_Name | StringDtype | False | 0 (0.0%) | 3201 (2.0%) | |||||
| 3 | Provider_Street_Address | StringDtype | False | 0 (0.0%) | 3326 (2.0%) | |||||
| 4 | Provider_City | StringDtype | False | 0 (0.0%) | 1977 (1.2%) | |||||
| 5 | Provider_State | StringDtype | False | 0 (0.0%) | 51 (< 0.1%) | |||||
| 6 | Provider_Zip_Code | Int64DType | False | 0 (0.0%) | 3053 (1.9%) | 4.79e+04 | 2.79e+04 | 1,040 | 44,309 | 99,835 |
| 7 | Hospital_Referral_Region_(HRR)_Description | StringDtype | False | 0 (0.0%) | 306 (0.2%) | |||||
| 8 | Total_Discharges | Int64DType | False | 0 (0.0%) | 642 (0.4%) | 42.8 | 51.1 | 11 | 27 | 3,383 |
| 9 | Average_Covered_Charges | Float64DType | False | 0 (0.0%) | 161985 (99.3%) | 3.61e+04 | 3.51e+04 | 2.46e+03 | 2.52e+04 | 9.29e+05 |
| 10 | Average_Medicare_Payments | Float64DType | False | 0 (0.0%) | 157817 (96.8%) | 8.49e+03 | 7.31e+03 | 1.15e+03 | 6.16e+03 | 1.55e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
DRG_Definition
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
100 (< 0.1%)
This column has a high cardinality (> 40).
Most frequent values
194 - SIMPLE PNEUMONIA & PLEURISY W CC
690 - KIDNEY & URINARY TRACT INFECTIONS W/O MCC
292 - HEART FAILURE & SHOCK W CC
392 - ESOPHAGITIS, GASTROENT & MISC DIGEST DISORDERS W/O MCC
641 - MISC DISORDERS OF NUTRITION,METABOLISM,FLUIDS/ELECTROLYTES W/O MCC
871 - SEPTICEMIA OR SEVERE SEPSIS W/O MV 96+ HOURS W MCC
603 - CELLULITIS W/O MCC
470 - MAJOR JOINT REPLACEMENT OR REATTACHMENT OF LOWER EXTREMITY W/O MCC
191 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W CC
List:190 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W MCC
['194 - SIMPLE PNEUMONIA & PLEURISY W CC', '690 - KIDNEY & URINARY TRACT INFECTIONS W/O MCC', '292 - HEART FAILURE & SHOCK W CC', '392 - ESOPHAGITIS, GASTROENT & MISC DIGEST DISORDERS W/O MCC', '641 - MISC DISORDERS OF NUTRITION,METABOLISM,FLUIDS/ELECTROLYTES W/O MCC', '871 - SEPTICEMIA OR SEVERE SEPSIS W/O MV 96+ HOURS W MCC', '603 - CELLULITIS W/O MCC', '470 - MAJOR JOINT REPLACEMENT OR REATTACHMENT OF LOWER EXTREMITY W/O MCC', '191 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W CC', '190 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W MCC']
Provider_Id
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
3,337 (2.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.56e+05 ± 1.52e+05
- Median ± IQR
- 250,007 ± 269,983
- Min | Max
- 10,001 | 670,077
Provider_Name
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
3,201 (2.0%)
This column has a high cardinality (> 40).
Most frequent values
GOOD SAMARITAN HOSPITAL
ST JOSEPH MEDICAL CENTER
MERCY MEDICAL CENTER
MERCY HOSPITAL
ST JOSEPH HOSPITAL
ST FRANCIS MEDICAL CENTER
ST MARY MEDICAL CENTER
ST LUKES HOSPITAL
ST FRANCIS HOSPITAL
List:JEFFERSON REGIONAL MEDICAL CENTER
['GOOD SAMARITAN HOSPITAL', 'ST JOSEPH MEDICAL CENTER', 'MERCY MEDICAL CENTER', 'MERCY HOSPITAL', 'ST JOSEPH HOSPITAL', 'ST FRANCIS MEDICAL CENTER', 'ST MARY MEDICAL CENTER', 'ST LUKES HOSPITAL', 'ST FRANCIS HOSPITAL', 'JEFFERSON REGIONAL MEDICAL CENTER']
Provider_Street_Address
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
3,326 (2.0%)
This column has a high cardinality (> 40).
Most frequent values
100 MEDICAL CENTER DRIVE
800 WASHINGTON STREET
1 MEDICAL CENTER DRIVE
100 HOSPITAL DRIVE
809 UNIVERSITY BOULEVARD EAST
9601 INTERSTATE 630, EXIT 7
1303 E HERNDON AVE
8700 BEVERLY BLVD
20 YORK ST
List:4755 OGLETOWN-STANTON ROAD
['100 MEDICAL CENTER DRIVE', '800 WASHINGTON STREET', '1 MEDICAL CENTER DRIVE', '100 HOSPITAL DRIVE', '809 UNIVERSITY BOULEVARD EAST', '9601 INTERSTATE 630, EXIT 7', '1303 E HERNDON AVE', '8700 BEVERLY BLVD', '20 YORK ST', '4755 OGLETOWN-STANTON ROAD']
Provider_City
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
1,977 (1.2%)
This column has a high cardinality (> 40).
Most frequent values
CHICAGO
BALTIMORE
HOUSTON
PHILADELPHIA
BROOKLYN
SPRINGFIELD
COLUMBUS
LOS ANGELES
NEW YORK
List:DALLAS
['CHICAGO', 'BALTIMORE', 'HOUSTON', 'PHILADELPHIA', 'BROOKLYN', 'SPRINGFIELD', 'COLUMBUS', 'LOS ANGELES', 'NEW YORK', 'DALLAS']
Provider_State
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
51 (< 0.1%)
This column has a high cardinality (> 40).
Most frequent values
CA
TX
FL
NY
IL
PA
OH
MI
NC
List:GA
['CA', 'TX', 'FL', 'NY', 'IL', 'PA', 'OH', 'MI', 'NC', 'GA']
Provider_Zip_Code
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
3,053 (1.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 4.79e+04 ± 2.79e+04
- Median ± IQR
- 44,309 ± 45,640
- Min | Max
- 1,040 | 99,835
Hospital_Referral_Region_(HRR)_Description
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
306 (0.2%)
This column has a high cardinality (> 40).
Most frequent values
CA - Los Angeles
MA - Boston
GA - Atlanta
TX - Houston
PA - Philadelphia
TX - Dallas
MO - St. Louis
NY - East Long Island
FL - Orlando
List:TN - Nashville
['CA - Los Angeles', 'MA - Boston', 'GA - Atlanta', 'TX - Houston', 'PA - Philadelphia', 'TX - Dallas', 'MO - St. Louis', 'NY - East Long Island', 'FL - Orlando', 'TN - Nashville']
Total_Discharges
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
642 (0.4%)
This column has a high cardinality (> 40).
- Mean ± Std
- 42.8 ± 51.1
- Median ± IQR
- 27 ± 32
- Min | Max
- 11 | 3,383
Average_Covered_Charges
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
161,985 (99.3%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.61e+04 ± 3.51e+04
- Median ± IQR
- 2.52e+04 ± 2.73e+04
- Min | Max
- 2.46e+03 | 9.29e+05
Average_Medicare_Payments
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
157,817 (96.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 8.49e+03 ± 7.31e+03
- Median ± IQR
- 6.16e+03 ± 5.86e+03
- Min | Max
- 1.15e+03 | 1.55e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
| Column 1 | Column 2 | Cramér's V | Pearson's Correlation |
|---|---|---|---|
| Provider_State | Provider_Zip_Code | 0.620 | |
| Provider_Id | Provider_State | 0.574 | |
| Provider_Name | Provider_Street_Address | 0.570 | |
| Average_Covered_Charges | Average_Medicare_Payments | 0.427 | 0.781 |
| Provider_Id | Provider_Zip_Code | 0.405 | -0.180 |
| Provider_Zip_Code | Hospital_Referral_Region_(HRR)_Description | 0.367 | |
| Provider_State | Hospital_Referral_Region_(HRR)_Description | 0.366 | |
| Provider_Id | Hospital_Referral_Region_(HRR)_Description | 0.318 | |
| Provider_Name | Provider_City | 0.265 | |
| Provider_City | Hospital_Referral_Region_(HRR)_Description | 0.250 | |
| Provider_City | Provider_State | 0.245 | |
| Provider_City | Provider_Zip_Code | 0.209 | |
| Provider_Street_Address | Provider_City | 0.192 | |
| Provider_Street_Address | Hospital_Referral_Region_(HRR)_Description | 0.182 | |
| Provider_Id | Provider_City | 0.181 | |
| Provider_Street_Address | Provider_State | 0.154 | |
| Provider_Name | Hospital_Referral_Region_(HRR)_Description | 0.149 | |
| Provider_Street_Address | Provider_Zip_Code | 0.136 | |
| DRG_Definition | Total_Discharges | 0.128 | |
| Provider_State | Average_Covered_Charges | 0.128 | |
| Provider_Street_Address | Average_Covered_Charges | 0.127 | |
| Provider_Id | Provider_Street_Address | 0.120 | |
| Provider_Name | Provider_Zip_Code | 0.114 | |
| Provider_Zip_Code | Average_Covered_Charges | 0.101 | 0.156 |
| Provider_Id | Provider_Name | 0.100 | |
| Provider_Name | Provider_State | 0.100 | |
| Provider_City | Total_Discharges | 0.0966 | |
| Hospital_Referral_Region_(HRR)_Description | Average_Covered_Charges | 0.0951 | |
| DRG_Definition | Average_Medicare_Payments | 0.0935 | |
| Provider_Zip_Code | Average_Medicare_Payments | 0.0912 | 0.0452 |
| Provider_Street_Address | Average_Medicare_Payments | 0.0862 | |
| Provider_State | Average_Medicare_Payments | 0.0850 | |
| Provider_Id | Average_Covered_Charges | 0.0810 | -0.0975 |
| DRG_Definition | Average_Covered_Charges | 0.0761 | |
| Hospital_Referral_Region_(HRR)_Description | Average_Medicare_Payments | 0.0652 | |
| DRG_Definition | Provider_Id | 0.0640 | |
| Provider_Id | Average_Medicare_Payments | 0.0639 | -0.0388 |
| DRG_Definition | Provider_Zip_Code | 0.0582 | |
| DRG_Definition | Provider_Street_Address | 0.0561 | |
| Provider_City | Average_Medicare_Payments | 0.0555 | |
| Provider_Name | Average_Covered_Charges | 0.0555 | |
| Provider_State | Total_Discharges | 0.0555 | |
| DRG_Definition | Provider_Name | 0.0553 | |
| DRG_Definition | Hospital_Referral_Region_(HRR)_Description | 0.0548 | |
| Provider_Zip_Code | Total_Discharges | 0.0540 | -0.0608 |
| Provider_City | Average_Covered_Charges | 0.0532 | |
| Provider_Id | Total_Discharges | 0.0529 | 0.00671 |
| DRG_Definition | Provider_City | 0.0524 | |
| DRG_Definition | Provider_State | 0.0503 | |
| Total_Discharges | Average_Medicare_Payments | 0.0423 | -0.0157 |
| Provider_Name | Average_Medicare_Payments | 0.0378 | |
| Hospital_Referral_Region_(HRR)_Description | Total_Discharges | 0.0300 | |
| Provider_Street_Address | Total_Discharges | 0.0264 | |
| Total_Discharges | Average_Covered_Charges | 0.0227 | -0.0293 |
| Provider_Name | Total_Discharges | 0.0177 |
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
| Average_Total_Payments | |
|---|---|
| 0 | 5.78e+03 |
| 1 | 5.79e+03 |
| 2 | 5.43e+03 |
| 3 | 5.42e+03 |
| 4 | 5.66e+03 |
| 163,060 | 3.81e+03 |
| 163,061 | 4.03e+03 |
| 163,062 | 5.70e+03 |
| 163,063 | 7.66e+03 |
| 163,064 | 3.54e+03 |
Average_Total_Payments
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
154,891 (95.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 9.71e+03 ± 7.66e+03
- Median ± IQR
- 7.21e+03 ± 6.05e+03
- Min | Max
- 2.67e+03 | 1.56e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | Average_Total_Payments | Float64DType | False | 0 (0.0%) | 154891 (95.0%) | 9.71e+03 | 7.66e+03 | 2.67e+03 | 7.21e+03 | 1.56e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Average_Total_Payments
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
154,891 (95.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 9.71e+03 ± 7.66e+03
- Median ± IQR
- 7.21e+03 ± 6.05e+03
- Min | Max
- 2.67e+03 | 1.56e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
Next, we subsample 3,000 rows to make the example run faster.
X = X_full.sample(3_000, random_state=42).reset_index(drop=True)
y = y_full.sample(3_000, random_state=42).reset_index(drop=True)
We now create a splitter, vectorizer, and regressor that we reuse throughout
the example. evaluate() clones them before fitting.
High-cardinality strings are encoded with StringEncoder (TF-IDF
then randomized SVD); we seed it so reruns stay comparable.
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.pipeline import make_pipeline
from skore import TrainTestSplit
from skrub import StringEncoder, TableVectorizer
splitter = TrainTestSplit(random_state=42, test_size=0.2)
vectorizer = TableVectorizer(high_cardinality=StringEncoder(random_state=42))
regressor = HistGradientBoostingRegressor(random_state=42)
Trigger SKD011 and SKD012#
Let us train a gradient boosting model, using skrub’s TableVectorizer to vectorize
the data first.
from skore import evaluate
model = make_pipeline(vectorizer, regressor)
first_report = evaluate(
model,
X=X,
y=y,
splitter=splitter,
)
first_report
| Metric | HistGradientBoostingRegressor |
|---|---|
| R² | 0.968625 |
| RMSE | 1236.859022 |
| MAE | 555.159914 |
| MAPE | 0.052583 |
| Fit time (s) | 2.457898 |
| Predict time (s) | 0.232536 |
- [SKD001] Potential overfitting. Significant train/test gaps were found for 3/4 default predictive metrics.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['learning_rate', 'max_leaf_nodes'] for HistGradientBoostingRegressor.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('tablevectorizer',
TableVectorizer(high_cardinality=StringEncoder(random_state=42))),
('histgradientboostingregressor',
HistGradientBoostingRegressor(random_state=42))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
Parameters
Fitted attributes
['Provider_Id', 'Provider_Zip_Code', 'Total_Discharges', 'Average_Covered_Charges', 'Average_Medicare_Payments']
Parameters
Parameters
Parameters
['DRG_Definition', 'Provider_Name', 'Provider_Street_Address', 'Provider_City', 'Provider_State', 'Hospital_Referral_Region_(HRR)_Description']
Parameters
100 of 185 features
| DRG_Definition_00 |
| DRG_Definition_01 |
| DRG_Definition_02 |
| DRG_Definition_03 |
| DRG_Definition_04 |
| DRG_Definition_05 |
| DRG_Definition_06 |
| DRG_Definition_07 |
| DRG_Definition_08 |
| DRG_Definition_09 |
| DRG_Definition_10 |
| DRG_Definition_11 |
| DRG_Definition_12 |
| DRG_Definition_13 |
| DRG_Definition_14 |
| DRG_Definition_15 |
| DRG_Definition_16 |
| DRG_Definition_17 |
| DRG_Definition_18 |
| DRG_Definition_19 |
| DRG_Definition_20 |
| DRG_Definition_21 |
| DRG_Definition_22 |
| DRG_Definition_23 |
| DRG_Definition_24 |
| DRG_Definition_25 |
| DRG_Definition_26 |
| DRG_Definition_27 |
| DRG_Definition_28 |
| DRG_Definition_29 |
| Provider_Id |
| Provider_Name_00 |
| Provider_Name_01 |
| Provider_Name_02 |
| Provider_Name_03 |
| Provider_Name_04 |
| Provider_Name_05 |
| Provider_Name_06 |
| Provider_Name_07 |
| Provider_Name_08 |
| Provider_Name_09 |
| Provider_Name_10 |
| Provider_Name_11 |
| Provider_Name_12 |
| Provider_Name_13 |
| Provider_Name_14 |
| Provider_Name_15 |
| Provider_Name_16 |
| Provider_Name_17 |
| Provider_Name_18 |
| Provider_Name_19 |
| Provider_Name_20 |
| Provider_Name_21 |
| Provider_Name_22 |
| Provider_Name_23 |
| Provider_Name_24 |
| Provider_Name_25 |
| Provider_Name_26 |
| Provider_Name_27 |
| Provider_Name_28 |
| Provider_Name_29 |
| Provider_Street_Address_00 |
| Provider_Street_Address_01 |
| Provider_Street_Address_02 |
| Provider_Street_Address_03 |
| Provider_Street_Address_04 |
| Provider_Street_Address_05 |
| Provider_Street_Address_06 |
| Provider_Street_Address_07 |
| Provider_Street_Address_08 |
| Provider_Street_Address_09 |
| Provider_Street_Address_10 |
| Provider_Street_Address_11 |
| Provider_Street_Address_12 |
| Provider_Street_Address_13 |
| Provider_Street_Address_14 |
| Provider_Street_Address_15 |
| Provider_Street_Address_16 |
| Provider_Street_Address_17 |
| Provider_Street_Address_18 |
| Provider_Street_Address_19 |
| Provider_Street_Address_20 |
| Provider_Street_Address_21 |
| Provider_Street_Address_22 |
| Provider_Street_Address_23 |
| Provider_Street_Address_24 |
| Provider_Street_Address_25 |
| Provider_Street_Address_26 |
| Provider_Street_Address_27 |
| Provider_Street_Address_28 |
| Provider_Street_Address_29 |
| Provider_City_00 |
| Provider_City_01 |
| Provider_City_02 |
| Provider_City_03 |
| Provider_City_04 |
| Provider_City_05 |
| Provider_City_06 |
| Provider_City_07 |
| Provider_City_08 |
Parameters
Fitted attributes
| DRG_Definition | Provider_Id | Provider_Name | Provider_Street_Address | Provider_City | Provider_State | Provider_Zip_Code | Hospital_Referral_Region_(HRR)_Description | Total_Discharges | Average_Covered_Charges | Average_Medicare_Payments | Average_Total_Payments | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 191 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W CC | 390,065 | GETTYSBURG HOSPITAL | 147 GETTYS STREET | GETTYSBURG | PA | 17,325 | PA - Harrisburg | 30 | 1.13e+04 | 4.98e+03 | 5.69e+03 |
| 1 | 243 - PERMANENT CARDIAC PACEMAKER IMPLANT W CC | 310,014 | COOPER UNIVERSITY HOSPITAL | 1 COOPER PLAZA | CAMDEN | NJ | 8,103 | NJ - Camden | 20 | 1.51e+05 | 2.29e+04 | 2.37e+04 |
| 2 | 191 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W CC | 360,091 | MEDINA HOSPITAL | 1000 EAST WASHINGTON STREET | MEDINA | OH | 44,256 | OH - Cleveland | 41 | 1.75e+04 | 4.31e+03 | 4.95e+03 |
| 3 | 699 - OTHER KIDNEY & URINARY TRACT DIAGNOSES W CC | 70,020 | MIDDLESEX HOSPITAL | 28 CRESCENT ST | MIDDLETOWN | CT | 6,457 | CT - Hartford | 19 | 2.44e+04 | 5.93e+03 | 6.76e+03 |
| 4 | 470 - MAJOR JOINT REPLACEMENT OR REATTACHMENT OF LOWER EXTREMITY W/O MCC | 50,334 | SALINAS VALLEY MEMORIAL HOSPITAL | 450 EAST ROMIE LANE | SALINAS | CA | 93,901 | CA - Salinas | 204 | 7.06e+04 | 1.60e+04 | 2.04e+04 |
| 2,995 | 066 - INTRACRANIAL HEMORRHAGE OR CEREBRAL INFARCTION W/O CC/MCC | 50,438 | HUNTINGTON MEMORIAL HOSPITAL | 100 W CALIFORNIA BLVD | PASADENA | CA | 91,109 | CA - Los Angeles | 27 | 4.49e+04 | 4.99e+03 | 6.38e+03 |
| 2,996 | 917 - POISONING & TOXIC EFFECTS OF DRUGS W MCC | 450,809 | NORTH AUSTIN MEDICAL CENTER | 12221 MOPAC EXPRESSWAY NORTH | AUSTIN | TX | 78,758 | TX - Austin | 13 | 6.08e+04 | 8.77e+03 | 1.10e+04 |
| 2,997 | 313 - CHEST PAIN | 490,009 | UNIVERSITY OF VIRGINIA MEDICAL CENTER | JEFFERSON PARK AVE | CHARLOTTESVILLE | VA | 22,908 | VA - Charlottesville | 23 | 1.77e+04 | 4.85e+03 | 5.87e+03 |
| 2,998 | 190 - CHRONIC OBSTRUCTIVE PULMONARY DISEASE W MCC | 310,064 | ATLANTICARE REGIONAL MEDICAL CENTER - CITY DIV | 1925 PACIFIC AVE | ATLANTIC CITY | NJ | 8,401 | NJ - Camden | 128 | 7.00e+04 | 7.46e+03 | 8.20e+03 |
| 2,999 | 699 - OTHER KIDNEY & URINARY TRACT DIAGNOSES W CC | 210,051 | DOCTORS' COMMUNITY HOSPITAL | 8118 GOOD LUCK ROAD | LANHAM | MD | 20,706 | MD - Takoma Park | 14 | 9.33e+03 | 7.72e+03 | 8.77e+03 |
DRG_Definition
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
100 (3.3%)
This column has a high cardinality (> 40).
Provider_Id
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
1,754 (58.5%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.54e+05 ± 1.51e+05
- Median ± IQR
- 250,001 ± 260,183
- Min | Max
- 10,007 | 670,060
Provider_Name
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
1,690 (56.3%)
This column has a high cardinality (> 40).
Provider_Street_Address
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
1,750 (58.3%)
This column has a high cardinality (> 40).
Provider_City
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
1,181 (39.4%)
This column has a high cardinality (> 40).
Provider_State
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
51 (1.7%)
This column has a high cardinality (> 40).
Provider_Zip_Code
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
1,669 (55.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 4.80e+04 ± 2.82e+04
- Median ± IQR
- 44,266 ± 47,299
- Min | Max
- 1,040 | 99,801
Hospital_Referral_Region_(HRR)_Description
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
295 (9.8%)
This column has a high cardinality (> 40).
Total_Discharges
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
211 (7.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 43.4 ± 52.0
- Median ± IQR
- 27 ± 33
- Min | Max
- 11 | 778
Average_Covered_Charges
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,000 (100.0%)
This column has a high cardinality (> 40).
- Mean ± Std
- 3.62e+04 ± 3.29e+04
- Median ± IQR
- 2.57e+04 ± 2.67e+04
- Min | Max
- 4.84e+03 | 3.07e+05
Average_Medicare_Payments
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
2,998 (99.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 8.33e+03 ± 6.86e+03
- Median ± IQR
- 6.04e+03 ± 5.52e+03
- Min | Max
- 1.75e+03 | 7.25e+04
Average_Total_Payments
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
2,993 (99.8%)
This column has a high cardinality (> 40).
- Mean ± Std
- 9.53e+03 ± 7.17e+03
- Median ± IQR
- 7.06e+03 ± 5.79e+03
- Min | Max
- 2.72e+03 | 7.32e+04
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | DRG_Definition | StringDtype | False | 0 (0.0%) | 100 (3.3%) | |||||
| 1 | Provider_Id | Int64DType | False | 0 (0.0%) | 1754 (58.5%) | 2.54e+05 | 1.51e+05 | 10,007 | 250,001 | 670,060 |
| 2 | Provider_Name | StringDtype | False | 0 (0.0%) | 1690 (56.3%) | |||||
| 3 | Provider_Street_Address | StringDtype | False | 0 (0.0%) | 1750 (58.3%) | |||||
| 4 | Provider_City | StringDtype | False | 0 (0.0%) | 1181 (39.4%) | |||||
| 5 | Provider_State | StringDtype | False | 0 (0.0%) | 51 (1.7%) | |||||
| 6 | Provider_Zip_Code | Int64DType | False | 0 (0.0%) | 1669 (55.6%) | 4.80e+04 | 2.82e+04 | 1,040 | 44,266 | 99,801 |
| 7 | Hospital_Referral_Region_(HRR)_Description | StringDtype | False | 0 (0.0%) | 295 (9.8%) | |||||
| 8 | Total_Discharges | Int64DType | False | 0 (0.0%) | 211 (7.0%) | 43.4 | 52.0 | 11 | 27 | 778 |
| 9 | Average_Covered_Charges | Float64DType | False | 0 (0.0%) | 3000 (100.0%) | 3.62e+04 | 3.29e+04 | 4.84e+03 | 2.57e+04 | 3.07e+05 |
| 10 | Average_Medicare_Payments | Float64DType | False | 0 (0.0%) | 2998 (99.9%) | 8.33e+03 | 6.86e+03 | 1.75e+03 | 6.04e+03 | 7.25e+04 |
| 11 | Average_Total_Payments | Float64DType | False | 0 (0.0%) | 2993 (99.8%) | 9.53e+03 | 7.17e+03 | 2.72e+03 | 7.06e+03 | 7.32e+04 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
We can see from the metrics table that the models performs very well on the data,
with an R² of 0.96.
However, looking at the checks, we can see in the Tips tab that SKD011 triggers on
the Average_Medicare_Payments column.
first_report.checks.summarize()
- [SKD001] Potential overfitting. Significant train/test gaps were found for 3/4 default predictive metrics.
- [SKD009] Model performance vs. HistGradientBoosting baseline. Your model is on par with or better than a HistGradientBoosting baseline. Baseline performance on the test set, for reference: MAE=563, MAPE=0.0514, RMSE=1.31e+03, R²=0.965.
- [SKD011] Golden feature. A model trained on feature(s) ['Average_Medicare_Payments'] alone has similar performance to a model trained on all the features, on the default predictive metrics. This may signal data leakage or excessive reliance on a single feature.
- [SKD012] Useless features. Feature(s) ['Hospital_Referral_Region_(HRR)_Description', 'Provider_City', 'Provider_Id', 'Provider_State', 'Provider_Street_Address', 'Provider_Zip_Code', 'Total_Discharges'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['learning_rate', 'max_leaf_nodes'] for HistGradientBoostingRegressor.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
We should then question whether this column would be available in a real deployment. If it would, then we have found a good proxy for the target and we can move on. If it would not, then we should not use this column to train our model.
In our case, the target (Average_Total_Payments) and the golden feature
(Average_Medicare_Payments) come from the same process (billing aggregates).
Therefore, Average_Medicare_Payments would not be available in a real deployment,
and we should not use it to train our model. As a matter of fact,
Average_Covered_Charges would also not be able in a deployment setting so we will
also drop it.
A final thing to be careful about is that SKD012 flags most provider fields
and Total_Discharges as useless. As we can see in the importance plot
below, every feature looks weak next to the golden billing column, even
though some of them may recover once we drop it. We will keep them for now.
_ = first_report.inspection.permutation_importance().plot()
# SKD011 - Inspect after dropping golden feature
# ==============================================
#
# First let's evaluate the same model on the same test set with the payment features
# removed, and compare it with our original model.
#
# Thankfully there is still signal in the remaining features, as we get a decent R²
# of 0.89.
from skore import compare
X_without_payment = X.drop(
columns=["Average_Medicare_Payments", "Average_Covered_Charges"]
)
second_report = evaluate(
model,
X=X_without_payment,
y=y,
splitter=splitter,
)
comparison = compare(
{
"with_payment_features": first_report,
"without_payment_features": second_report,
}
)
comparison.metrics.summarize().frame()

| estimator | with_payment_features | without_payment_features |
|---|---|---|
| metric | ||
| r2 | 0.968625 | 0.890437 |
| rmse | 1236.859022 | 2311.336441 |
| mae | 555.159914 | 1318.776062 |
| mape | 0.052583 | 0.138348 |
| fit_time | 2.457898 | 2.374167 |
| predict_time | 0.232536 | 0.224389 |
Let us inspect which features are now the ones our model relies on the most.
We can see that DRG_Definition is by far most important one, according to
permutation importance. This feature encodes the medical reason for which patients
were treated. It will therefore be available at inference time and we can use it.
_ = second_report.inspection.permutation_importance().plot()

SKD012 now flags Provider_Id and Total_Discharges. Those look genuinely
weak once billing leakage is gone: an identifier and a discharge count that
barely moves the score. Other provider fields recovered some signal, which
is why we did not treat the leaky SKD012 list as columns to drop.
The check still only sees original columns. TableVectorizer creates extra
components from high-cardinality strings that can be weak too; we prune
those next.
second_report.checks.summarize()
- [SKD001] Potential overfitting. Significant train/test gaps were found for 4/4 default predictive metrics.
- [SKD009] Model performance vs. HistGradientBoosting baseline. Your model is on par with or better than a HistGradientBoosting baseline. Baseline performance on the test set, for reference: MAE=1.32e+03, MAPE=0.14, RMSE=2.23e+03, R²=0.898.
- [SKD012] Useless features. Feature(s) ['Provider_Id', 'Total_Discharges'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['learning_rate', 'max_leaf_nodes'] for HistGradientBoostingRegressor.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
SKD012 - prune weak features inside the pipeline#
SKD012 flagged Provider_Id and Total_Discharges on the original columns.
It does not see the extra components TableVectorizer builds from
high-cardinality fields such as DRG_Definition and provider names; some of
those add little once the useful ones are present. We therefore select
after vectorizing, with SelectFromModel,
so the choice is fit on the training fold only.
Histogram gradient boosting has no feature_importances_, so the selector
uses GradientBoostingRegressor. The final
estimator stays a HistGradientBoostingRegressor on
the reduced matrix.
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.feature_selection import SelectFromModel
model_reduced = make_pipeline(
vectorizer,
SelectFromModel(
GradientBoostingRegressor(random_state=42),
threshold="median",
),
regressor,
)
model_reduced
Pipeline(steps=[('tablevectorizer',
TableVectorizer(high_cardinality=StringEncoder(random_state=42))),
('selectfrommodel',
SelectFromModel(estimator=GradientBoostingRegressor(random_state=42),
threshold='median')),
('histgradientboostingregressor',
HistGradientBoostingRegressor(random_state=42))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Parameters
Parameters
Parameters
Parameters
Parameters
Parameters
GradientBoostingRegressor(random_state=42)
Parameters
Parameters
Let us evaluate the reduced pipeline on the same split. Test scores are a on par, and the model is lighter: selection dropped the weak vectorized components before fitting histogram boosting.
third_report = evaluate(
model_reduced,
X=X_without_payment,
y=y,
splitter=splitter,
)
comparison_reduced = compare(
{
"without_payment_features": second_report,
"without_payment_features_reduced": third_report,
}
)
comparison_reduced.metrics.summarize().frame()
| estimator | without_payment_features | without_payment_features_reduced |
|---|---|---|
| metric | ||
| r2 | 0.890437 | 0.890613 |
| rmse | 2311.336441 | 2309.483190 |
| mae | 1318.776062 | 1316.978826 |
| mape | 0.138348 | 0.138346 |
| fit_time | 2.374167 | 7.935268 |
| predict_time | 0.224389 | 0.134846 |
Conclusion#
SKD011 and SKD012 often appear together when one leaky column dominates.
Audit that column and compare with-and-without it; do not treat SKD012 flags
on a leaky report as a drop list. Once leakage is gone, SKD012 may still
flag identifiers or weak columns (Provider_Id, Total_Discharges here)
while TableVectorizer also builds weak encoded components. Prune those
inside the pipeline, for example with
SelectFromModel, and check that test
performance is preserved.
Total running time of the script: (2 minutes 17.182 seconds)