Note
Go to the end to download the full example code.
SKD010 - Model slower than baseline#
This example walks through mitigations when the check
SKD010 fires. The check compares the user’s
model to a fast linear baseline
(RidgeCV for regression, wrapped in
tabular_pipeline()) and flags a problem only when both of the
following hold:
the larger of the fit-time and test predict-time ratios is at least 2 times the baseline, with an absolute gap of at least 1 second on that dimension, and
test scores are not significantly better than the baseline on a majority of default predictive metrics.
The issue is paying for latency without a significant quality premium. A clear
quality gain over RidgeCV keeps the check quiet.
So does a model that is already about as fast as the baseline.
Responses tried below, in order:
check that preprocessing is not the actual bottleneck,
reduce the model’s complexity,
switch to the fast linear baseline when quality is sufficient,
profile fit time to understand the dominant cost.
We use the employee salaries dataset with a heavy
RandomForestRegressor inside
tabular_pipeline(). The goal is either to match
RidgeCV quality at lower cost, or to shrink fit
time until the speed gap is no longer unjustified.
Load the employee salaries dataset#
fetch_employee_salaries() returns human-resources
records with mixed categorical and numeric fields. A 200-tree
RandomForestRegressor inside
tabular_pipeline() trains slowly yet often fails to beat
the fast RidgeCV baseline on test scores.
SKD010 is a slow check.
from skrub.datasets import fetch_employee_salaries
dataset = fetch_employee_salaries()
X = dataset.X
y = dataset.y
Inspect column types with TableReport. Encoding cost is part
of total fit time.
from skrub import TableReport
TableReport(X)
| gender | department | department_name | division | assignment_category | employee_position_title | date_first_hired | year_first_hired | |
|---|---|---|---|---|---|---|---|---|
| 0 | F | POL | Department of Police | MSB Information Mgmt and Tech Division Records Management Section | Fulltime-Regular | Office Services Coordinator | 09/22/1986 | 1,986 |
| 1 | M | POL | Department of Police | ISB Major Crimes Division Fugitive Section | Fulltime-Regular | Master Police Officer | 09/12/1988 | 1,988 |
| 2 | F | HHS | Department of Health and Human Services | Adult Protective and Case Management Services | Fulltime-Regular | Social Worker IV | 11/19/1989 | 1,989 |
| 3 | M | COR | Correction and Rehabilitation | PRRS Facility and Security | Fulltime-Regular | Resident Supervisor II | 05/05/2014 | 2,014 |
| 4 | M | HCA | Department of Housing and Community Affairs | Affordable Housing Programs | Fulltime-Regular | Planning Specialist III | 03/05/2007 | 2,007 |
| 9,223 | F | HHS | Department of Health and Human Services | School Based Health Centers | Fulltime-Regular | Community Health Nurse II | 11/03/2015 | 2,015 |
| 9,224 | F | FRS | Fire and Rescue Services | Human Resources Division | Fulltime-Regular | Fire/Rescue Division Chief | 11/28/1988 | 1,988 |
| 9,225 | M | HHS | Department of Health and Human Services | Child and Adolescent Mental Health Clinic Services | Parttime-Regular | Medical Doctor IV - Psychiatrist | 04/30/2001 | 2,001 |
| 9,226 | M | CCL | County Council | Council Central Staff | Fulltime-Regular | Manager II | 09/05/2006 | 2,006 |
| 9,227 | M | DLC | Department of Liquor Control | Licensure, Regulation and Education | Fulltime-Regular | Alcohol/Tobacco Enforcement Specialist II | 01/30/2012 | 2,012 |
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
Most frequent values
M
F
['M', 'F']
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
Most frequent values
POL
HHS
FRS
DOT
COR
DLC
DGS
LIB
DPS
SHF
['POL', 'HHS', 'FRS', 'DOT', 'COR', 'DLC', 'DGS', 'LIB', 'DPS', 'SHF']
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
Most frequent values
Department of Police
Department of Health and Human Services
Fire and Rescue Services
Department of Transportation
Correction and Rehabilitation
Department of Liquor Control
Department of General Services
Department of Public Libraries
Department of Permitting Services
Sheriff's Office
['Department of Police', 'Department of Health and Human Services', 'Fire and Rescue Services', 'Department of Transportation', 'Correction and Rehabilitation', 'Department of Liquor Control', 'Department of General Services', 'Department of Public Libraries', 'Department of Permitting Services', "Sheriff's Office"]
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
Most frequent values
School Health Services
Transit Silver Spring Ride On
Transit Gaithersburg Ride On
Highway Services
Child Welfare Services
FSB Traffic Division School Safety Section
Income Supports
PSB 3rd District Patrol
PSB 4th District Patrol
List:Transit Nicholson Ride On
['School Health Services', 'Transit Silver Spring Ride On', 'Transit Gaithersburg Ride On', 'Highway Services', 'Child Welfare Services', 'FSB Traffic Division School Safety Section', 'Income Supports', 'PSB 3rd District Patrol', 'PSB 4th District Patrol', 'Transit Nicholson Ride On']
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
Most frequent values
Fulltime-Regular
Parttime-Regular
['Fulltime-Regular', 'Parttime-Regular']
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
Most frequent values
Bus Operator
Police Officer III
Firefighter/Rescuer III
Manager III
Firefighter/Rescuer II
Master Firefighter/Rescuer
Office Services Coordinator
School Health Room Technician I
Police Officer II
List:Community Health Nurse II
['Bus Operator', 'Police Officer III', 'Firefighter/Rescuer III', 'Manager III', 'Firefighter/Rescuer II', 'Master Firefighter/Rescuer', 'Office Services Coordinator', 'School Health Room Technician I', 'Police Officer II', 'Community Health Nurse II']
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
Most frequent values
12/12/2016
01/14/2013
02/24/2014
03/10/2014
08/12/2013
10/06/2014
09/22/2014
03/19/2007
07/16/2012
List:07/29/2013
['12/12/2016', '01/14/2013', '02/24/2014', '03/10/2014', '08/12/2013', '10/06/2014', '09/22/2014', '03/19/2007', '07/16/2012', '07/29/2013']
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | gender | StringDtype | False | 17 (0.2%) | 2 (< 0.1%) | |||||
| 1 | department | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 2 | department_name | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 3 | division | StringDtype | False | 0 (0.0%) | 694 (7.5%) | |||||
| 4 | assignment_category | StringDtype | False | 0 (0.0%) | 2 (< 0.1%) | |||||
| 5 | employee_position_title | StringDtype | False | 0 (0.0%) | 443 (4.8%) | |||||
| 6 | date_first_hired | StringDtype | False | 0 (0.0%) | 2264 (24.5%) | |||||
| 7 | year_first_hired | Int64DType | False | 0 (0.0%) | 51 (0.6%) | 2.00e+03 | 9.33 | 1,965 | 2,005 | 2,016 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
Most frequent values
M
F
['M', 'F']
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
Most frequent values
POL
HHS
FRS
DOT
COR
DLC
DGS
LIB
DPS
SHF
['POL', 'HHS', 'FRS', 'DOT', 'COR', 'DLC', 'DGS', 'LIB', 'DPS', 'SHF']
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
Most frequent values
Department of Police
Department of Health and Human Services
Fire and Rescue Services
Department of Transportation
Correction and Rehabilitation
Department of Liquor Control
Department of General Services
Department of Public Libraries
Department of Permitting Services
Sheriff's Office
['Department of Police', 'Department of Health and Human Services', 'Fire and Rescue Services', 'Department of Transportation', 'Correction and Rehabilitation', 'Department of Liquor Control', 'Department of General Services', 'Department of Public Libraries', 'Department of Permitting Services', "Sheriff's Office"]
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
Most frequent values
School Health Services
Transit Silver Spring Ride On
Transit Gaithersburg Ride On
Highway Services
Child Welfare Services
FSB Traffic Division School Safety Section
Income Supports
PSB 3rd District Patrol
PSB 4th District Patrol
List:Transit Nicholson Ride On
['School Health Services', 'Transit Silver Spring Ride On', 'Transit Gaithersburg Ride On', 'Highway Services', 'Child Welfare Services', 'FSB Traffic Division School Safety Section', 'Income Supports', 'PSB 3rd District Patrol', 'PSB 4th District Patrol', 'Transit Nicholson Ride On']
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
Most frequent values
Fulltime-Regular
Parttime-Regular
['Fulltime-Regular', 'Parttime-Regular']
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
Most frequent values
Bus Operator
Police Officer III
Firefighter/Rescuer III
Manager III
Firefighter/Rescuer II
Master Firefighter/Rescuer
Office Services Coordinator
School Health Room Technician I
Police Officer II
List:Community Health Nurse II
['Bus Operator', 'Police Officer III', 'Firefighter/Rescuer III', 'Manager III', 'Firefighter/Rescuer II', 'Master Firefighter/Rescuer', 'Office Services Coordinator', 'School Health Room Technician I', 'Police Officer II', 'Community Health Nurse II']
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
Most frequent values
12/12/2016
01/14/2013
02/24/2014
03/10/2014
08/12/2013
10/06/2014
09/22/2014
03/19/2007
07/16/2012
List:07/29/2013
['12/12/2016', '01/14/2013', '02/24/2014', '03/10/2014', '08/12/2013', '10/06/2014', '09/22/2014', '03/19/2007', '07/16/2012', '07/29/2013']
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
| Column 1 | Column 2 | Cramér's V | Pearson's Correlation |
|---|---|---|---|
| department | department_name | 1.00 | |
| division | assignment_category | 0.593 | |
| assignment_category | employee_position_title | 0.497 | |
| department_name | assignment_category | 0.422 | |
| department | assignment_category | 0.422 | |
| department | employee_position_title | 0.413 | |
| department_name | employee_position_title | 0.413 | |
| division | employee_position_title | 0.410 | |
| department | division | 0.381 | |
| department_name | division | 0.381 | |
| gender | department | 0.380 | |
| gender | department_name | 0.380 | |
| gender | assignment_category | 0.294 | |
| gender | employee_position_title | 0.275 | |
| gender | division | 0.265 | |
| employee_position_title | date_first_hired | 0.179 | |
| date_first_hired | year_first_hired | 0.151 | |
| department | date_first_hired | 0.150 | |
| department_name | date_first_hired | 0.150 | |
| employee_position_title | year_first_hired | 0.131 | |
| gender | date_first_hired | 0.104 | |
| division | year_first_hired | 0.0862 | |
| department | year_first_hired | 0.0811 | |
| department_name | year_first_hired | 0.0811 | |
| assignment_category | date_first_hired | 0.0756 | |
| division | date_first_hired | 0.0728 | |
| gender | year_first_hired | 0.0641 | |
| assignment_category | year_first_hired | 0.0519 |
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
Salaries are continuous and right-skewed, a typical regression target.
| current_annual_salary | |
|---|---|
| 0 | 6.92e+04 |
| 1 | 9.74e+04 |
| 2 | 1.05e+05 |
| 3 | 5.27e+04 |
| 4 | 9.34e+04 |
| 9,223 | 7.21e+04 |
| 9,224 | 1.70e+05 |
| 9,225 | 1.03e+05 |
| 9,226 | 1.54e+05 |
| 9,227 | 7.55e+04 |
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | current_annual_salary | Float64DType | False | 0 (0.0%) | 3403 (36.9%) | 7.34e+04 | 2.91e+04 | 9.20e+03 | 6.94e+04 | 3.03e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
from skore import TrainTestSplit
splitter = TrainTestSplit(test_size=0.2, random_state=42)
Trigger SKD010 with large leaves#
One hundred trees that must keep 100 training rows in every leaf barely
split. tabular_pipeline() still encodes the high-cardinality
strings, so the fit stays much slower than RidgeCV,
and the test scores lose
the quality premium a forest can have on this table.
from sklearn.ensemble import RandomForestRegressor
from skore import evaluate
from skrub import tabular_pipeline
report_constrained = evaluate(
tabular_pipeline(
RandomForestRegressor(
n_estimators=100,
min_samples_leaf=100,
random_state=42,
n_jobs=4,
)
),
X=X,
y=y,
splitter=splitter,
)
report_constrained
| Metric | RandomForestRegressor |
|---|---|
| R² | 0.737815 |
| RMSE | 15024.106340 |
| MAE | 8599.922385 |
| MAPE | 0.125128 |
| Fit time (s) | 4.949131 |
| Predict time (s) | 0.283453 |
- [SKD008] Highly correlated input features. 3 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
No tips were emitted for your report.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('tablevectorizer',
TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
unknown_value=-1))),
('randomforestregressor',
RandomForestRegressor(min_samples_leaf=100, n_jobs=4,
random_state=42))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
Parameters
Fitted attributes
['year_first_hired']
Parameters
['date_first_hired']
Parameters
['gender', 'department', 'department_name', 'assignment_category']
Parameters
['division', 'employee_position_title']
Parameters
69 features
| gender |
| department |
| department_name |
| division_00 |
| division_01 |
| division_02 |
| division_03 |
| division_04 |
| division_05 |
| division_06 |
| division_07 |
| division_08 |
| division_09 |
| division_10 |
| division_11 |
| division_12 |
| division_13 |
| division_14 |
| division_15 |
| division_16 |
| division_17 |
| division_18 |
| division_19 |
| division_20 |
| division_21 |
| division_22 |
| division_23 |
| division_24 |
| division_25 |
| division_26 |
| division_27 |
| division_28 |
| division_29 |
| assignment_category |
| employee_position_title_00 |
| employee_position_title_01 |
| employee_position_title_02 |
| employee_position_title_03 |
| employee_position_title_04 |
| employee_position_title_05 |
| employee_position_title_06 |
| employee_position_title_07 |
| employee_position_title_08 |
| employee_position_title_09 |
| employee_position_title_10 |
| employee_position_title_11 |
| employee_position_title_12 |
| employee_position_title_13 |
| employee_position_title_14 |
| employee_position_title_15 |
| employee_position_title_16 |
| employee_position_title_17 |
| employee_position_title_18 |
| employee_position_title_19 |
| employee_position_title_20 |
| employee_position_title_21 |
| employee_position_title_22 |
| employee_position_title_23 |
| employee_position_title_24 |
| employee_position_title_25 |
| employee_position_title_26 |
| employee_position_title_27 |
| employee_position_title_28 |
| employee_position_title_29 |
| date_first_hired_year |
| date_first_hired_month |
| date_first_hired_day |
| date_first_hired_total_seconds |
| year_first_hired |
Parameters
Fitted attributes
| gender | department | department_name | division | assignment_category | employee_position_title | date_first_hired | year_first_hired | current_annual_salary | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | M | DOT | Department of Transportation | Transit Silver Spring Ride On | Fulltime-Regular | Bus Operator | 09/04/2006 | 2,006 | 5.07e+04 |
| 1 | M | COR | Correction and Rehabilitation | DS MCCF Unit 2 Security | Fulltime-Regular | Correctional Officer III (Corporal) | 11/14/2005 | 2,005 | 6.74e+04 |
| 2 | M | FRS | Fire and Rescue Services | Recruit Training | Fulltime-Regular | Firefighter/Rescuer I (Recruit) | 12/12/2016 | 2,016 | 4.53e+04 |
| 3 | F | HHS | Department of Health and Human Services | Child Welfare Services | Fulltime-Regular | Social Worker II | 06/06/2011 | 2,011 | 6.47e+04 |
| 4 | M | DGS | Department of General Services | Facilities Maintenance | Fulltime-Regular | HVAC Mechanic II | 12/10/2000 | 2,000 | 7.73e+04 |
| 9,223 | M | FRS | Fire and Rescue Services | Station 15 | Fulltime-Regular | Firefighter/Rescuer II | 09/22/2014 | 2,014 | 5.09e+04 |
| 9,224 | F | HHS | Department of Health and Human Services | Care Coordination | Fulltime-Regular | Community Health Nurse II | 10/16/2006 | 2,006 | 9.23e+04 |
| 9,225 | M | DOT | Department of Transportation | Traffic Engineering Studies Section | Fulltime-Regular | Engineer Technician II | 06/29/2003 | 2,003 | 6.46e+04 |
| 9,226 | M | HHS | Department of Health and Human Services | Adult Protective and Case Management Services | Parttime-Regular | Office Clerk | 04/27/2010 | 2,010 | 1.73e+04 |
| 9,227 | F | HHS | Department of Health and Human Services | School Health Services | Parttime-Regular | School Health Room Technician I | 03/05/2001 | 2,001 | 4.98e+04 |
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | gender | StringDtype | False | 17 (0.2%) | 2 (< 0.1%) | |||||
| 1 | department | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 2 | department_name | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 3 | division | StringDtype | False | 0 (0.0%) | 694 (7.5%) | |||||
| 4 | assignment_category | StringDtype | False | 0 (0.0%) | 2 (< 0.1%) | |||||
| 5 | employee_position_title | StringDtype | False | 0 (0.0%) | 443 (4.8%) | |||||
| 6 | date_first_hired | StringDtype | False | 0 (0.0%) | 2264 (24.5%) | |||||
| 7 | year_first_hired | Int64DType | False | 0 (0.0%) | 51 (0.6%) | 2.00e+03 | 9.33 | 1,965 | 2,005 | 2,016 |
| 8 | current_annual_salary | Float64DType | False | 0 (0.0%) | 3403 (36.9%) | 7.34e+04 | 2.91e+04 | 9.20e+03 | 6.94e+04 | 3.03e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
SKD010 is in the issues below: fit time is several times the
RidgeCV baseline, and the test scores are not
significantly better.
report_constrained.checks.summarize()
- [SKD008] Highly correlated input features. 3 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
- [SKD010] Model slower than baseline. Fit time is ~4.7x slower than a fast linear baseline without significantly better test scores (3/4 default predictive metrics).
- [SKD009] Model performance vs. HistGradientBoosting baseline. Test scores are significantly worse than a HistGradientBoosting baseline for 4/4 default predictive metrics. Baseline performance on the test set: MAE=4.58e+03, MAPE=0.0622, RMSE=8.91e+03, R²=0.908.
- [SKD012] Useless features. Feature(s) ['department_name', 'gender', 'year_first_hired'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD011] Golden feature. Failed to create report from single feature.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Fit the fast linear baseline#
RidgeCV inside tabular_pipeline()
is the reference skore uses for regression speed. This pipeline is the baseline
SKD010 compares against, fit on the same split as the forest above.
from sklearn.linear_model import RidgeCV
report_linear = evaluate(
tabular_pipeline(RidgeCV()),
X=X,
y=y,
splitter=splitter,
)
report_linear
| Metric | RidgeCV |
|---|---|
| R² | 0.758462 |
| RMSE | 14420.395938 |
| MAE | 9111.963211 |
| MAPE | 0.131599 |
| Fit time (s) | 1.048433 |
| Predict time (s) | 0.267647 |
- [SKD008] Highly correlated input features. 43 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
- [SKD006] Coefficient interpretation. Features are not on the same scale: coefficient magnitudes are not directly comparable as feature importance.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD007] MDI biased for high-cardinality features. Estimator is not a tree-based model: it does not have a `feature_importances_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD016] Estimator not tuned. No parameter to recommend.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('tablevectorizer',
TableVectorizer(datetime=DatetimeEncoder(periodic_encoding='spline'))),
('simpleimputer', SimpleImputer(add_indicator=True)),
('squashingscaler', SquashingScaler(max_absolute_value=5)),
('ridgecv', RidgeCV())])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
Parameters
Fitted attributes
['year_first_hired']
Parameters
['date_first_hired']
Parameters
['gender', 'department', 'department_name', 'assignment_category']
Parameters
['division', 'employee_position_title']
Parameters
100 of 157 features
| gender_F |
| gender_M |
| gender_nan |
| department_BOA |
| department_BOE |
| department_CAT |
| department_CCL |
| department_CEC |
| department_CEX |
| department_COR |
| department_CUS |
| department_DEP |
| department_DGS |
| department_DHS |
| department_DLC |
| department_DOT |
| department_DPS |
| department_DTS |
| department_ECM |
| department_FIN |
| department_FRS |
| department_HCA |
| department_HHS |
| department_HRC |
| department_IGR |
| department_LIB |
| department_MPB |
| department_NDA |
| department_OAG |
| department_OCP |
| department_OHR |
| department_OIG |
| department_OLO |
| department_OMB |
| department_PIO |
| department_POL |
| department_PRO |
| department_REC |
| department_SHF |
| department_ZAH |
| department_name_Board of Appeals Department |
| department_name_Board of Elections |
| department_name_Community Engagement Cluster |
| department_name_Community Use of Public Facilities |
| department_name_Correction and Rehabilitation |
| department_name_County Attorney's Office |
| department_name_County Council |
| department_name_Department of Environmental Protection |
| department_name_Department of Finance |
| department_name_Department of General Services |
| department_name_Department of Health and Human Services |
| department_name_Department of Housing and Community Affairs |
| department_name_Department of Liquor Control |
| department_name_Department of Permitting Services |
| department_name_Department of Police |
| department_name_Department of Public Libraries |
| department_name_Department of Recreation |
| department_name_Department of Technology Services |
| department_name_Department of Transportation |
| department_name_Ethics Commission |
| department_name_Fire and Rescue Services |
| department_name_Merit System Protection Board Department |
| department_name_Non-Departmental Account |
| department_name_Office of Agriculture |
| department_name_Office of Consumer Protection |
| department_name_Office of Emergency Management and Homeland Security |
| department_name_Office of Human Resources |
| department_name_Office of Human Rights |
| department_name_Office of Intergovernmental Relations Department |
| department_name_Office of Legislative Oversight |
| department_name_Office of Management and Budget |
| department_name_Office of Procurement |
| department_name_Office of Public Information |
| department_name_Office of Zoning and Administrative Hearings |
| department_name_Office of the Inspector General |
| department_name_Offices of the County Executive |
| department_name_Sheriff's Office |
| division_00 |
| division_01 |
| division_02 |
| division_03 |
| division_04 |
| division_05 |
| division_06 |
| division_07 |
| division_08 |
| division_09 |
| division_10 |
| division_11 |
| division_12 |
| division_13 |
| division_14 |
| division_15 |
| division_16 |
| division_17 |
| division_18 |
| division_19 |
| division_20 |
| division_21 |
| division_22 |
Parameters
Fitted attributes
| Name | Type | Value |
|---|---|---|
|
feature_names_in_
feature_names_in_: ndarray of shape (`n_features_in_`,) Names of features seen during :term:`fit`. Defined only when `X` has feature names that are all strings. .. versionadded:: 1.0 |
ndarray[object](157,) | ['gender_F','gender_M','gender_nan',...,'date_first_hired_day_spline_2', 'date_first_hired_day_spline_3','year_first_hired'] |
|
indicator_
indicator_: :class:`~sklearn.impute.MissingIndicator` Indicator used to add binary indicators for missing values. `None` if `add_indicator=False`. |
MissingIndicator | MissingIndica..._on_new=False) |
|
n_features_in_
n_features_in_: int Number of features seen during :term:`fit`. .. versionadded:: 0.24 |
int | 157 |
|
statistics_
statistics_: array of shape (n_features,) The imputation fill value for each feature. Computing statistics can result in `np.nan` values. During :meth:`transform`, features corresponding to `np.nan` statistics will be discarded. |
ndarray[float64](157,) | [ 0.4 , 0.6 , 0. ,..., 0.26, 0.26,2003.61] |
100 of 157 features
| gender_F |
| gender_M |
| gender_nan |
| department_BOA |
| department_BOE |
| department_CAT |
| department_CCL |
| department_CEC |
| department_CEX |
| department_COR |
| department_CUS |
| department_DEP |
| department_DGS |
| department_DHS |
| department_DLC |
| department_DOT |
| department_DPS |
| department_DTS |
| department_ECM |
| department_FIN |
| department_FRS |
| department_HCA |
| department_HHS |
| department_HRC |
| department_IGR |
| department_LIB |
| department_MPB |
| department_NDA |
| department_OAG |
| department_OCP |
| department_OHR |
| department_OIG |
| department_OLO |
| department_OMB |
| department_PIO |
| department_POL |
| department_PRO |
| department_REC |
| department_SHF |
| department_ZAH |
| department_name_Board of Appeals Department |
| department_name_Board of Elections |
| department_name_Community Engagement Cluster |
| department_name_Community Use of Public Facilities |
| department_name_Correction and Rehabilitation |
| department_name_County Attorney's Office |
| department_name_County Council |
| department_name_Department of Environmental Protection |
| department_name_Department of Finance |
| department_name_Department of General Services |
| department_name_Department of Health and Human Services |
| department_name_Department of Housing and Community Affairs |
| department_name_Department of Liquor Control |
| department_name_Department of Permitting Services |
| department_name_Department of Police |
| department_name_Department of Public Libraries |
| department_name_Department of Recreation |
| department_name_Department of Technology Services |
| department_name_Department of Transportation |
| department_name_Ethics Commission |
| department_name_Fire and Rescue Services |
| department_name_Merit System Protection Board Department |
| department_name_Non-Departmental Account |
| department_name_Office of Agriculture |
| department_name_Office of Consumer Protection |
| department_name_Office of Emergency Management and Homeland Security |
| department_name_Office of Human Resources |
| department_name_Office of Human Rights |
| department_name_Office of Intergovernmental Relations Department |
| department_name_Office of Legislative Oversight |
| department_name_Office of Management and Budget |
| department_name_Office of Procurement |
| department_name_Office of Public Information |
| department_name_Office of Zoning and Administrative Hearings |
| department_name_Office of the Inspector General |
| department_name_Offices of the County Executive |
| department_name_Sheriff's Office |
| division_00 |
| division_01 |
| division_02 |
| division_03 |
| division_04 |
| division_05 |
| division_06 |
| division_07 |
| division_08 |
| division_09 |
| division_10 |
| division_11 |
| division_12 |
| division_13 |
| division_14 |
| division_15 |
| division_16 |
| division_17 |
| division_18 |
| division_19 |
| division_20 |
| division_21 |
| division_22 |
Parameters
Fitted attributes
| Name | Type | Value |
|---|---|---|
| minmax_cols_ | ndarray[bool](157,) | [False,False, True,...,False,False,False] |
| minmax_scaler_ | _MinMaxScaler | _MinMaxScaler() |
| n_features_in_ | int | 157 |
| robust_cols_ | ndarray[bool](157,) | [ True, True,False,..., True, True, True] |
| robust_scaler_ | RobustScaler | RobustScaler() |
| zero_cols_ | ndarray[bool](157,) | [False,False,False,...,False,False,False] |
100 of 157 features
| x0 |
| x1 |
| x2 |
| x3 |
| x4 |
| x5 |
| x6 |
| x7 |
| x8 |
| x9 |
| x10 |
| x11 |
| x12 |
| x13 |
| x14 |
| x15 |
| x16 |
| x17 |
| x18 |
| x19 |
| x20 |
| x21 |
| x22 |
| x23 |
| x24 |
| x25 |
| x26 |
| x27 |
| x28 |
| x29 |
| x30 |
| x31 |
| x32 |
| x33 |
| x34 |
| x35 |
| x36 |
| x37 |
| x38 |
| x39 |
| x40 |
| x41 |
| x42 |
| x43 |
| x44 |
| x45 |
| x46 |
| x47 |
| x48 |
| x49 |
| x50 |
| x51 |
| x52 |
| x53 |
| x54 |
| x55 |
| x56 |
| x57 |
| x58 |
| x59 |
| x60 |
| x61 |
| x62 |
| x63 |
| x64 |
| x65 |
| x66 |
| x67 |
| x68 |
| x69 |
| x70 |
| x71 |
| x72 |
| x73 |
| x74 |
| x75 |
| x76 |
| x77 |
| x78 |
| x79 |
| x80 |
| x81 |
| x82 |
| x83 |
| x84 |
| x85 |
| x86 |
| x87 |
| x88 |
| x89 |
| x90 |
| x91 |
| x92 |
| x93 |
| x94 |
| x95 |
| x96 |
| x97 |
| x98 |
| x99 |
Parameters
Fitted attributes
| gender | department | department_name | division | assignment_category | employee_position_title | date_first_hired | year_first_hired | current_annual_salary | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | M | DOT | Department of Transportation | Transit Silver Spring Ride On | Fulltime-Regular | Bus Operator | 09/04/2006 | 2,006 | 5.07e+04 |
| 1 | M | COR | Correction and Rehabilitation | DS MCCF Unit 2 Security | Fulltime-Regular | Correctional Officer III (Corporal) | 11/14/2005 | 2,005 | 6.74e+04 |
| 2 | M | FRS | Fire and Rescue Services | Recruit Training | Fulltime-Regular | Firefighter/Rescuer I (Recruit) | 12/12/2016 | 2,016 | 4.53e+04 |
| 3 | F | HHS | Department of Health and Human Services | Child Welfare Services | Fulltime-Regular | Social Worker II | 06/06/2011 | 2,011 | 6.47e+04 |
| 4 | M | DGS | Department of General Services | Facilities Maintenance | Fulltime-Regular | HVAC Mechanic II | 12/10/2000 | 2,000 | 7.73e+04 |
| 9,223 | M | FRS | Fire and Rescue Services | Station 15 | Fulltime-Regular | Firefighter/Rescuer II | 09/22/2014 | 2,014 | 5.09e+04 |
| 9,224 | F | HHS | Department of Health and Human Services | Care Coordination | Fulltime-Regular | Community Health Nurse II | 10/16/2006 | 2,006 | 9.23e+04 |
| 9,225 | M | DOT | Department of Transportation | Traffic Engineering Studies Section | Fulltime-Regular | Engineer Technician II | 06/29/2003 | 2,003 | 6.46e+04 |
| 9,226 | M | HHS | Department of Health and Human Services | Adult Protective and Case Management Services | Parttime-Regular | Office Clerk | 04/27/2010 | 2,010 | 1.73e+04 |
| 9,227 | F | HHS | Department of Health and Human Services | School Health Services | Parttime-Regular | School Health Room Technician I | 03/05/2001 | 2,001 | 4.98e+04 |
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | gender | StringDtype | False | 17 (0.2%) | 2 (< 0.1%) | |||||
| 1 | department | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 2 | department_name | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 3 | division | StringDtype | False | 0 (0.0%) | 694 (7.5%) | |||||
| 4 | assignment_category | StringDtype | False | 0 (0.0%) | 2 (< 0.1%) | |||||
| 5 | employee_position_title | StringDtype | False | 0 (0.0%) | 443 (4.8%) | |||||
| 6 | date_first_hired | StringDtype | False | 0 (0.0%) | 2264 (24.5%) | |||||
| 7 | year_first_hired | Int64DType | False | 0 (0.0%) | 51 (0.6%) | 2.00e+03 | 9.33 | 1,965 | 2,005 | 2,016 |
| 8 | current_annual_salary | Float64DType | False | 0 (0.0%) | 3403 (36.9%) | 7.34e+04 | 2.91e+04 | 9.20e+03 | 6.94e+04 | 3.03e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
from skore import compare
compare({"large_leaves": report_constrained, "ridge": report_linear}).metrics.summarize(
data_source="test",
).frame()
| estimator | large_leaves | ridge |
|---|---|---|
| metric | ||
| r2 | 0.737815 | 0.758462 |
| rmse | 15024.106340 | 14420.395938 |
| mae | 8599.922385 | 9111.963211 |
| mape | 0.125128 | 0.131599 |
| fit_time | 4.949131 | 1.048433 |
| predict_time | 0.283453 | 0.267647 |
Fit time clears the 2x / 1s gate. R², RMSE, and MAE stay close to
RidgeCV,
short of max(0.01, 0.05 * |baseline|) on a majority of the default
predictive metrics. That is the warning.
Grow the trees to full depth#
Drop the leaf limit and keep the same 100 trees. The fit gets slower. The
test scores move far enough ahead of RidgeCV
that SKD010 passes.
report_full = evaluate(
tabular_pipeline(
RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4)
),
X=X,
y=y,
splitter=splitter,
)
report_full
| Metric | RandomForestRegressor |
|---|---|
| R² | 0.900830 |
| RMSE | 9240.077297 |
| MAE | 4332.343421 |
| MAPE | 0.058477 |
| Fit time (s) | 11.314932 |
| Predict time (s) | 0.299839 |
- [SKD001] Potential overfitting. Significant train/test gaps were found for 3/4 default predictive metrics.
- [SKD008] Highly correlated input features. 3 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['max_features', 'min_samples_leaf'] for RandomForestRegressor.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('tablevectorizer',
TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
unknown_value=-1))),
('randomforestregressor',
RandomForestRegressor(n_jobs=4, random_state=42))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
Parameters
Fitted attributes
['year_first_hired']
Parameters
['date_first_hired']
Parameters
['gender', 'department', 'department_name', 'assignment_category']
Parameters
['division', 'employee_position_title']
Parameters
69 features
| gender |
| department |
| department_name |
| division_00 |
| division_01 |
| division_02 |
| division_03 |
| division_04 |
| division_05 |
| division_06 |
| division_07 |
| division_08 |
| division_09 |
| division_10 |
| division_11 |
| division_12 |
| division_13 |
| division_14 |
| division_15 |
| division_16 |
| division_17 |
| division_18 |
| division_19 |
| division_20 |
| division_21 |
| division_22 |
| division_23 |
| division_24 |
| division_25 |
| division_26 |
| division_27 |
| division_28 |
| division_29 |
| assignment_category |
| employee_position_title_00 |
| employee_position_title_01 |
| employee_position_title_02 |
| employee_position_title_03 |
| employee_position_title_04 |
| employee_position_title_05 |
| employee_position_title_06 |
| employee_position_title_07 |
| employee_position_title_08 |
| employee_position_title_09 |
| employee_position_title_10 |
| employee_position_title_11 |
| employee_position_title_12 |
| employee_position_title_13 |
| employee_position_title_14 |
| employee_position_title_15 |
| employee_position_title_16 |
| employee_position_title_17 |
| employee_position_title_18 |
| employee_position_title_19 |
| employee_position_title_20 |
| employee_position_title_21 |
| employee_position_title_22 |
| employee_position_title_23 |
| employee_position_title_24 |
| employee_position_title_25 |
| employee_position_title_26 |
| employee_position_title_27 |
| employee_position_title_28 |
| employee_position_title_29 |
| date_first_hired_year |
| date_first_hired_month |
| date_first_hired_day |
| date_first_hired_total_seconds |
| year_first_hired |
Parameters
Fitted attributes
| gender | department | department_name | division | assignment_category | employee_position_title | date_first_hired | year_first_hired | current_annual_salary | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | M | DOT | Department of Transportation | Transit Silver Spring Ride On | Fulltime-Regular | Bus Operator | 09/04/2006 | 2,006 | 5.07e+04 |
| 1 | M | COR | Correction and Rehabilitation | DS MCCF Unit 2 Security | Fulltime-Regular | Correctional Officer III (Corporal) | 11/14/2005 | 2,005 | 6.74e+04 |
| 2 | M | FRS | Fire and Rescue Services | Recruit Training | Fulltime-Regular | Firefighter/Rescuer I (Recruit) | 12/12/2016 | 2,016 | 4.53e+04 |
| 3 | F | HHS | Department of Health and Human Services | Child Welfare Services | Fulltime-Regular | Social Worker II | 06/06/2011 | 2,011 | 6.47e+04 |
| 4 | M | DGS | Department of General Services | Facilities Maintenance | Fulltime-Regular | HVAC Mechanic II | 12/10/2000 | 2,000 | 7.73e+04 |
| 9,223 | M | FRS | Fire and Rescue Services | Station 15 | Fulltime-Regular | Firefighter/Rescuer II | 09/22/2014 | 2,014 | 5.09e+04 |
| 9,224 | F | HHS | Department of Health and Human Services | Care Coordination | Fulltime-Regular | Community Health Nurse II | 10/16/2006 | 2,006 | 9.23e+04 |
| 9,225 | M | DOT | Department of Transportation | Traffic Engineering Studies Section | Fulltime-Regular | Engineer Technician II | 06/29/2003 | 2,003 | 6.46e+04 |
| 9,226 | M | HHS | Department of Health and Human Services | Adult Protective and Case Management Services | Parttime-Regular | Office Clerk | 04/27/2010 | 2,010 | 1.73e+04 |
| 9,227 | F | HHS | Department of Health and Human Services | School Health Services | Parttime-Regular | School Health Room Technician I | 03/05/2001 | 2,001 | 4.98e+04 |
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | gender | StringDtype | False | 17 (0.2%) | 2 (< 0.1%) | |||||
| 1 | department | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 2 | department_name | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 3 | division | StringDtype | False | 0 (0.0%) | 694 (7.5%) | |||||
| 4 | assignment_category | StringDtype | False | 0 (0.0%) | 2 (< 0.1%) | |||||
| 5 | employee_position_title | StringDtype | False | 0 (0.0%) | 443 (4.8%) | |||||
| 6 | date_first_hired | StringDtype | False | 0 (0.0%) | 2264 (24.5%) | |||||
| 7 | year_first_hired | Int64DType | False | 0 (0.0%) | 51 (0.6%) | 2.00e+03 | 9.33 | 1,965 | 2,005 | 2,016 |
| 8 | current_annual_salary | Float64DType | False | 0 (0.0%) | 3403 (36.9%) | 7.34e+04 | 2.91e+04 | 9.20e+03 | 6.94e+04 | 3.03e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
SKD010 is among the passed checks. Slowness alone is not the warning.
report_full.checks.summarize()
- [SKD001] Potential overfitting. Significant train/test gaps were found for 3/4 default predictive metrics.
- [SKD008] Highly correlated input features. 3 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
- [SKD009] Model performance vs. HistGradientBoosting baseline. Your model is on par with or better than a HistGradientBoosting baseline. Baseline performance on the test set, for reference: MAE=4.54e+03, MAPE=0.062, RMSE=8.72e+03, R²=0.912.
- [SKD012] Useless features. Feature(s) ['gender'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['max_features', 'min_samples_leaf'] for RandomForestRegressor.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD011] Golden feature. Failed to create report from single feature.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
compare(
{
"large_leaves": report_constrained,
"full_depth": report_full,
"ridge": report_linear,
}
).metrics.summarize(data_source="test").frame()
| estimator | large_leaves | full_depth | ridge |
|---|---|---|---|
| metric | |||
| r2 | 0.737815 | 0.900830 | 0.758462 |
| rmse | 15024.106340 | 9240.077297 | 14420.395938 |
| mae | 8599.922385 | 4332.343421 | 9111.963211 |
| mape | 0.125128 | 0.058477 | 0.131599 |
| fit_time | 4.949131 | 11.314932 | 1.048433 |
| predict_time | 0.283453 | 0.299839 | 0.267647 |
Fit time increased from the constrained forest, and R² cleared the quality bar. The warning is gone.
Use a cheaper encoder#
Sometimes you can spend less time before touching the forest. Drop a column you do not
need: date_first_hired repeats year_first_hired, and a
StringEncoder still has to run on its thousands of distinct
dates. Or keep the columns and pick a cheaper
encoder. A tree can split on category codes, so an
OrdinalEncoder is enough for the high-cardinality
strings that tabular_pipeline() otherwise sends to
StringEncoder.
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import OrdinalEncoder
from skrub import TableVectorizer
report_ordinal = evaluate(
make_pipeline(
TableVectorizer(
high_cardinality=OrdinalEncoder(
handle_unknown="use_encoded_value",
unknown_value=-1,
encoded_missing_value=-1,
)
),
RandomForestRegressor(n_estimators=100, random_state=42, n_jobs=4),
),
X=X,
y=y,
splitter=splitter,
)
report_ordinal
| Metric | RandomForestRegressor |
|---|---|
| R² | 0.804144 |
| RMSE | 12985.332649 |
| MAE | 6681.730995 |
| MAPE | 0.090946 |
| Fit time (s) | 2.627551 |
| Predict time (s) | 0.221712 |
- [SKD001] Potential overfitting. Significant train/test gaps were found for 4/4 default predictive metrics.
- [SKD008] Highly correlated input features. 41 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
- [SKD016] Estimator not tuned. Estimator(s) left at default settings; consider tuning: ['max_features', 'min_samples_leaf'] for RandomForestRegressor.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('tablevectorizer',
TableVectorizer(high_cardinality=OrdinalEncoder(encoded_missing_value=-1,
handle_unknown='use_encoded_value',
unknown_value=-1))),
('randomforestregressor',
RandomForestRegressor(n_jobs=4, random_state=42))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
Parameters
Fitted attributes
['year_first_hired']
Parameters
['date_first_hired']
Parameters
['gender', 'department', 'department_name', 'assignment_category']
Parameters
['division', 'employee_position_title']
Parameters
85 features
| gender_F |
| gender_M |
| gender_nan |
| department_BOA |
| department_BOE |
| department_CAT |
| department_CCL |
| department_CEC |
| department_CEX |
| department_COR |
| department_CUS |
| department_DEP |
| department_DGS |
| department_DHS |
| department_DLC |
| department_DOT |
| department_DPS |
| department_DTS |
| department_ECM |
| department_FIN |
| department_FRS |
| department_HCA |
| department_HHS |
| department_HRC |
| department_IGR |
| department_LIB |
| department_MPB |
| department_NDA |
| department_OAG |
| department_OCP |
| department_OHR |
| department_OIG |
| department_OLO |
| department_OMB |
| department_PIO |
| department_POL |
| department_PRO |
| department_REC |
| department_SHF |
| department_ZAH |
| department_name_Board of Appeals Department |
| department_name_Board of Elections |
| department_name_Community Engagement Cluster |
| department_name_Community Use of Public Facilities |
| department_name_Correction and Rehabilitation |
| department_name_County Attorney's Office |
| department_name_County Council |
| department_name_Department of Environmental Protection |
| department_name_Department of Finance |
| department_name_Department of General Services |
| department_name_Department of Health and Human Services |
| department_name_Department of Housing and Community Affairs |
| department_name_Department of Liquor Control |
| department_name_Department of Permitting Services |
| department_name_Department of Police |
| department_name_Department of Public Libraries |
| department_name_Department of Recreation |
| department_name_Department of Technology Services |
| department_name_Department of Transportation |
| department_name_Ethics Commission |
| department_name_Fire and Rescue Services |
| department_name_Merit System Protection Board Department |
| department_name_Non-Departmental Account |
| department_name_Office of Agriculture |
| department_name_Office of Consumer Protection |
| department_name_Office of Emergency Management and Homeland Security |
| department_name_Office of Human Resources |
| department_name_Office of Human Rights |
| department_name_Office of Intergovernmental Relations Department |
| department_name_Office of Legislative Oversight |
| department_name_Office of Management and Budget |
| department_name_Office of Procurement |
| department_name_Office of Public Information |
| department_name_Office of Zoning and Administrative Hearings |
| department_name_Office of the Inspector General |
| department_name_Offices of the County Executive |
| department_name_Sheriff's Office |
| division |
| assignment_category_Parttime-Regular |
| employee_position_title |
| date_first_hired_year |
| date_first_hired_month |
| date_first_hired_day |
| date_first_hired_total_seconds |
| year_first_hired |
Parameters
Fitted attributes
| gender | department | department_name | division | assignment_category | employee_position_title | date_first_hired | year_first_hired | current_annual_salary | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | M | DOT | Department of Transportation | Transit Silver Spring Ride On | Fulltime-Regular | Bus Operator | 09/04/2006 | 2,006 | 5.07e+04 |
| 1 | M | COR | Correction and Rehabilitation | DS MCCF Unit 2 Security | Fulltime-Regular | Correctional Officer III (Corporal) | 11/14/2005 | 2,005 | 6.74e+04 |
| 2 | M | FRS | Fire and Rescue Services | Recruit Training | Fulltime-Regular | Firefighter/Rescuer I (Recruit) | 12/12/2016 | 2,016 | 4.53e+04 |
| 3 | F | HHS | Department of Health and Human Services | Child Welfare Services | Fulltime-Regular | Social Worker II | 06/06/2011 | 2,011 | 6.47e+04 |
| 4 | M | DGS | Department of General Services | Facilities Maintenance | Fulltime-Regular | HVAC Mechanic II | 12/10/2000 | 2,000 | 7.73e+04 |
| 9,223 | M | FRS | Fire and Rescue Services | Station 15 | Fulltime-Regular | Firefighter/Rescuer II | 09/22/2014 | 2,014 | 5.09e+04 |
| 9,224 | F | HHS | Department of Health and Human Services | Care Coordination | Fulltime-Regular | Community Health Nurse II | 10/16/2006 | 2,006 | 9.23e+04 |
| 9,225 | M | DOT | Department of Transportation | Traffic Engineering Studies Section | Fulltime-Regular | Engineer Technician II | 06/29/2003 | 2,003 | 6.46e+04 |
| 9,226 | M | HHS | Department of Health and Human Services | Adult Protective and Case Management Services | Parttime-Regular | Office Clerk | 04/27/2010 | 2,010 | 1.73e+04 |
| 9,227 | F | HHS | Department of Health and Human Services | School Health Services | Parttime-Regular | School Health Room Technician I | 03/05/2001 | 2,001 | 4.98e+04 |
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | gender | StringDtype | False | 17 (0.2%) | 2 (< 0.1%) | |||||
| 1 | department | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 2 | department_name | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 3 | division | StringDtype | False | 0 (0.0%) | 694 (7.5%) | |||||
| 4 | assignment_category | StringDtype | False | 0 (0.0%) | 2 (< 0.1%) | |||||
| 5 | employee_position_title | StringDtype | False | 0 (0.0%) | 443 (4.8%) | |||||
| 6 | date_first_hired | StringDtype | False | 0 (0.0%) | 2264 (24.5%) | |||||
| 7 | year_first_hired | Int64DType | False | 0 (0.0%) | 51 (0.6%) | 2.00e+03 | 9.33 | 1,965 | 2,005 | 2,016 |
| 8 | current_annual_salary | Float64DType | False | 0 (0.0%) | 3403 (36.9%) | 7.34e+04 | 2.91e+04 | 9.20e+03 | 6.94e+04 | 3.03e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
compare(
{"full_depth": report_full, "ordinal_encoder": report_ordinal}
).metrics.summarize(data_source="test").frame()
| estimator | full_depth | ordinal_encoder |
|---|---|---|
| metric | ||
| r2 | 0.900830 | 0.804144 |
| rmse | 9240.077297 | 12985.332649 |
| mae | 4332.343421 | 6681.730995 |
| mape | 0.058477 | 0.090946 |
| fit_time | 11.314932 | 2.627551 |
| predict_time | 0.299839 | 0.221712 |
The OrdinalEncoder cuts fit time and predict
time. Test R² shows what that speed costs relative to the
StringEncoder.
Reduce model complexity#
Keep the 100 trees and the full depth, and stop splits that would leave a
leaf with fewer than 20 training rows. min_samples_leaf shortens each
tree, which cuts fit time and predict time on the
StringEncoder pipeline.
report_leaf = evaluate(
tabular_pipeline(
RandomForestRegressor(
n_estimators=100,
min_samples_leaf=20,
random_state=42,
n_jobs=4,
)
),
X=X,
y=y,
splitter=splitter,
)
report_leaf
| Metric | RandomForestRegressor |
|---|---|
| R² | 0.835295 |
| RMSE | 11907.957168 |
| MAE | 5788.395857 |
| MAPE | 0.078879 |
| Fit time (s) | 7.048906 |
| Predict time (s) | 0.292267 |
- [SKD008] Highly correlated input features. 3 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
No tips were emitted for your report.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were muted.
Fast mode is on: expensive checks are skipped unless already cached.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
Pipeline(steps=[('tablevectorizer',
TableVectorizer(low_cardinality=OrdinalEncoder(handle_unknown='use_encoded_value',
unknown_value=-1))),
('randomforestregressor',
RandomForestRegressor(min_samples_leaf=20, n_jobs=4,
random_state=42))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
Fitted attributes
Parameters
Fitted attributes
['year_first_hired']
Parameters
['date_first_hired']
Parameters
['gender', 'department', 'department_name', 'assignment_category']
Parameters
['division', 'employee_position_title']
Parameters
69 features
| gender |
| department |
| department_name |
| division_00 |
| division_01 |
| division_02 |
| division_03 |
| division_04 |
| division_05 |
| division_06 |
| division_07 |
| division_08 |
| division_09 |
| division_10 |
| division_11 |
| division_12 |
| division_13 |
| division_14 |
| division_15 |
| division_16 |
| division_17 |
| division_18 |
| division_19 |
| division_20 |
| division_21 |
| division_22 |
| division_23 |
| division_24 |
| division_25 |
| division_26 |
| division_27 |
| division_28 |
| division_29 |
| assignment_category |
| employee_position_title_00 |
| employee_position_title_01 |
| employee_position_title_02 |
| employee_position_title_03 |
| employee_position_title_04 |
| employee_position_title_05 |
| employee_position_title_06 |
| employee_position_title_07 |
| employee_position_title_08 |
| employee_position_title_09 |
| employee_position_title_10 |
| employee_position_title_11 |
| employee_position_title_12 |
| employee_position_title_13 |
| employee_position_title_14 |
| employee_position_title_15 |
| employee_position_title_16 |
| employee_position_title_17 |
| employee_position_title_18 |
| employee_position_title_19 |
| employee_position_title_20 |
| employee_position_title_21 |
| employee_position_title_22 |
| employee_position_title_23 |
| employee_position_title_24 |
| employee_position_title_25 |
| employee_position_title_26 |
| employee_position_title_27 |
| employee_position_title_28 |
| employee_position_title_29 |
| date_first_hired_year |
| date_first_hired_month |
| date_first_hired_day |
| date_first_hired_total_seconds |
| year_first_hired |
Parameters
Fitted attributes
| gender | department | department_name | division | assignment_category | employee_position_title | date_first_hired | year_first_hired | current_annual_salary | |
|---|---|---|---|---|---|---|---|---|---|
| 0 | M | DOT | Department of Transportation | Transit Silver Spring Ride On | Fulltime-Regular | Bus Operator | 09/04/2006 | 2,006 | 5.07e+04 |
| 1 | M | COR | Correction and Rehabilitation | DS MCCF Unit 2 Security | Fulltime-Regular | Correctional Officer III (Corporal) | 11/14/2005 | 2,005 | 6.74e+04 |
| 2 | M | FRS | Fire and Rescue Services | Recruit Training | Fulltime-Regular | Firefighter/Rescuer I (Recruit) | 12/12/2016 | 2,016 | 4.53e+04 |
| 3 | F | HHS | Department of Health and Human Services | Child Welfare Services | Fulltime-Regular | Social Worker II | 06/06/2011 | 2,011 | 6.47e+04 |
| 4 | M | DGS | Department of General Services | Facilities Maintenance | Fulltime-Regular | HVAC Mechanic II | 12/10/2000 | 2,000 | 7.73e+04 |
| 9,223 | M | FRS | Fire and Rescue Services | Station 15 | Fulltime-Regular | Firefighter/Rescuer II | 09/22/2014 | 2,014 | 5.09e+04 |
| 9,224 | F | HHS | Department of Health and Human Services | Care Coordination | Fulltime-Regular | Community Health Nurse II | 10/16/2006 | 2,006 | 9.23e+04 |
| 9,225 | M | DOT | Department of Transportation | Traffic Engineering Studies Section | Fulltime-Regular | Engineer Technician II | 06/29/2003 | 2,003 | 6.46e+04 |
| 9,226 | M | HHS | Department of Health and Human Services | Adult Protective and Case Management Services | Parttime-Regular | Office Clerk | 04/27/2010 | 2,010 | 1.73e+04 |
| 9,227 | F | HHS | Department of Health and Human Services | School Health Services | Parttime-Regular | School Health Room Technician I | 03/05/2001 | 2,001 | 4.98e+04 |
gender
StringDtype- Null values
- 17 (0.2%)
- Unique values
- 2 (< 0.1%)
department
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
department_name
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 37 (0.4%)
division
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
694 (7.5%)
This column has a high cardinality (> 40).
assignment_category
StringDtype- Null values
- 0 (0.0%)
- Unique values
- 2 (< 0.1%)
employee_position_title
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
443 (4.8%)
This column has a high cardinality (> 40).
date_first_hired
StringDtype- Null values
- 0 (0.0%)
- Unique values
-
2,264 (24.5%)
This column has a high cardinality (> 40).
year_first_hired
Int64DType- Null values
- 0 (0.0%)
- Unique values
-
51 (0.6%)
This column has a high cardinality (> 40).
- Mean ± Std
- 2.00e+03 ± 9.33
- Median ± IQR
- 2,005 ± 14
- Min | Max
- 1,965 | 2,016
current_annual_salary
Float64DType- Null values
- 0 (0.0%)
- Unique values
-
3,403 (36.9%)
This column has a high cardinality (> 40).
- Mean ± Std
- 7.34e+04 ± 2.91e+04
- Median ± IQR
- 6.94e+04 ± 3.94e+04
- Min | Max
- 9.20e+03 | 3.03e+05
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
|
Column
|
Column name
|
dtype
|
Is sorted
|
Null values
|
Unique values
|
Mean
|
Std
|
Min
|
Median
|
Max
|
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | gender | StringDtype | False | 17 (0.2%) | 2 (< 0.1%) | |||||
| 1 | department | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 2 | department_name | StringDtype | False | 0 (0.0%) | 37 (0.4%) | |||||
| 3 | division | StringDtype | False | 0 (0.0%) | 694 (7.5%) | |||||
| 4 | assignment_category | StringDtype | False | 0 (0.0%) | 2 (< 0.1%) | |||||
| 5 | employee_position_title | StringDtype | False | 0 (0.0%) | 443 (4.8%) | |||||
| 6 | date_first_hired | StringDtype | False | 0 (0.0%) | 2264 (24.5%) | |||||
| 7 | year_first_hired | Int64DType | False | 0 (0.0%) | 51 (0.6%) | 2.00e+03 | 9.33 | 1,965 | 2,005 | 2,016 |
| 8 | current_annual_salary | Float64DType | False | 0 (0.0%) | 3403 (36.9%) | 7.34e+04 | 2.91e+04 | 9.20e+03 | 6.94e+04 | 3.03e+05 |
No columns match the selected filter: . You can change the column filter in the dropdown menu above.
Please enable javascript
The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").
SKD010 stays passed. The larger leaves shorten fit time on the same 100 trees.
report_leaf.checks.summarize()
- [SKD008] Highly correlated input features. 3 pair(s) of features have a Spearman correlation above 0.9. Highly correlated features can destabilize linear model coefficients and feature-importance estimates, and may cause collinearity-induced numerical issues.Dropping redundant features may also improve model performance.
- [SKD009] Model performance vs. HistGradientBoosting baseline. Test scores are significantly worse than a HistGradientBoosting baseline for 4/4 default predictive metrics. Baseline performance on the test set: MAE=4.51e+03, MAPE=0.0612, RMSE=8.8e+03, R²=0.91.
- [SKD012] Useless features. Feature(s) ['gender'] have permutation importance overlapping with zero and could likely be dropped without degrading performance. Dropping redundant features may also improve model performance.
- [SKD003] Inconsistent performance across splits. Not applicable to estimator reports.
- [SKD004] High class imbalance. ML task is not binary classification. Got regression.
- [SKD005] Underrepresented classes. ML task is not multiclass classification. Got regression.
- [SKD006] Coefficient interpretation. Estimator is not a linear model: it does not have a `coef_` attribute.
- [SKD011] Golden feature. Failed to create report from single feature.
- [SKD013] Train-test overlap in time series. No datetime column found.
- [SKD014] Hyperparameters at search edge. Estimator is not a BaseSearchCV instance. Got Pipeline.
- [SKD015] Hyperparameters worth tuning. Estimator is not a BaseSearchCV instance. Got Pipeline.
No checks were skipped in fast mode.
No checks were muted.
Mute a check by passing its code to ignore, e.g. .checks.summarize(ignore=['SKD001']).
report_leaf.metrics.summarize(
metric=["fit_time", "predict_time"],
data_source="test",
).frame()
metric
fit_time 7.048906
predict_time 0.292267
Name: RandomForestRegressor, dtype: float64
Compare the pipelines#
Fit time, test predict time, and test scores for every pipeline above. Large
leaves are the case that raises SKD010. Full depth is slower and justified.
The OrdinalEncoder and a moderate leaf limit
spend less time. RidgeCV is the speed
reference: SKD010 asks whether extra seconds buy a significant score gain
over that baseline.
comparison = compare(
{
"large_leaves": report_constrained,
"full_depth": report_full,
"ordinal_encoder": report_ordinal,
"moderate_leaves": report_leaf,
"ridge": report_linear,
}
)
comparison.metrics.summarize(data_source="test").frame()
| estimator | large_leaves | full_depth | ordinal_encoder | moderate_leaves | ridge |
|---|---|---|---|---|---|
| metric | |||||
| r2 | 0.737815 | 0.900830 | 0.804144 | 0.835295 | 0.758462 |
| rmse | 15024.106340 | 9240.077297 | 12985.332649 | 11907.957168 | 14420.395938 |
| mae | 8599.922385 | 4332.343421 | 6681.730995 | 5788.395857 | 9111.963211 |
| mape | 0.125128 | 0.058477 | 0.090946 | 0.078879 | 0.131599 |
| fit_time | 4.949131 | 11.314932 | 2.627551 | 7.048906 | 1.048433 |
| predict_time | 0.283453 | 0.299839 | 0.221712 | 0.292267 | 0.267647 |
Conclusion#
A usual forest can be slower than RidgeCV and
still pass SKD010 when test scores are significantly better. A cheaper
high-cardinality encoder, such as OrdinalEncoder,
or dropping a redundant column such as date_first_hired, cuts that cost.
min_samples_leaf shrinks the trees while keeping the same number of
trees. RidgeCV remains the speed reference when
a linear model is enough.
Total running time of the script: (1 minutes 33.161 seconds)