.. DO NOT EDIT. .. THIS FILE WAS AUTOMATICALLY GENERATED BY SPHINX-GALLERY. .. TO MAKE CHANGES, EDIT THE SOURCE PYTHON FILE: .. "auto_examples/technical_details/plot_cache_mechanism.py" .. LINE NUMBERS ARE GIVEN BELOW. .. only:: html .. note:: :class: sphx-glr-download-link-note :ref:`Go to the end ` to download the full example code. .. rst-class:: sphx-glr-example-title .. _sphx_glr_auto_examples_technical_details_plot_cache_mechanism.py: .. _example_cache_mechanism: =============== Cache mechanism =============== This example shows how :class:`~skore.EstimatorReport` and :class:`~skore.CrossValidationReport` use caching to speed up computations. .. GENERATED FROM PYTHON SOURCE LINES 13-18 Generating some data ==================== In this toy example, we create a large synthetic classification dataset that will let us see speed improvements easily. .. GENERATED FROM PYTHON SOURCE LINES 18-24 .. code-block:: Python import pandas as pd from sklearn.datasets import make_classification X, y = make_classification(n_samples=150_000, return_X_y=True) X = pd.DataFrame(X, columns=[str(i) for i in range(X.shape[1])]) .. GENERATED FROM PYTHON SOURCE LINES 25-26 Here is what the training data looks like: .. GENERATED FROM PYTHON SOURCE LINES 26-30 .. code-block:: Python from skrub import TableReport TableReport(X) .. raw:: html

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 31-32 And the target training data: .. GENERATED FROM PYTHON SOURCE LINES 32-34 .. code-block:: Python TableReport(y) .. raw:: html

Please enable javascript

The skrub table reports need javascript to display correctly. If you are displaying a report in a Jupyter notebook and you see this message, you may need to re-execute the cell or to trust the notebook (button on the top right or "File > Trust notebook").



.. GENERATED FROM PYTHON SOURCE LINES 35-37 We build a model using :func:`skrub.tabular_pipeline`: it is a simple predictive model that also performs basic feature engineering. .. GENERATED FROM PYTHON SOURCE LINES 37-42 .. code-block:: Python from skrub import tabular_pipeline model = tabular_pipeline("classifier") model .. raw:: html
Pipeline(steps=[('tablevectorizer',
                     TableVectorizer(low_cardinality=ToCategorical())),
                    ('histgradientboostingclassifier',
                     HistGradientBoostingClassifier())])
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.


.. GENERATED FROM PYTHON SOURCE LINES 43-48 Caching the predictions for fast metric computation =================================================== Let's explore how :class:`~skore.EstimatorReport` uses caching to speed up predictions. .. GENERATED FROM PYTHON SOURCE LINES 48-53 .. code-block:: Python from skore import evaluate report = evaluate(model, X, y, pos_label=1) report.help() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 54-55 We compute the accuracy on our test set and measure how long it takes: .. GENERATED FROM PYTHON SOURCE LINES 56-63 .. code-block:: Python import time start = time.time() report.metrics.accuracy() end = time.time() print(f"Time taken: {end - start:.3f} seconds") .. rst-class:: sphx-glr-script-out .. code-block:: none Time taken: 0.003 seconds .. GENERATED FROM PYTHON SOURCE LINES 64-65 For comparison, here's how scikit-learn computes the same accuracy score: .. GENERATED FROM PYTHON SOURCE LINES 66-73 .. code-block:: Python from sklearn.metrics import accuracy_score start = time.time() accuracy_score(report.y_test, report.estimator_.predict(report.X_test)) end = time.time() print(f"Time taken: {end - start:.2f} seconds") .. rst-class:: sphx-glr-script-out .. code-block:: none Time taken: 0.19 seconds .. GENERATED FROM PYTHON SOURCE LINES 74-77 skore outputs the result much faster than scikit-learn. How can this be? The answer lies in the EstimatorReport's state. When the EstimatorReport is created, it computes the model predictions, and caches them: .. GENERATED FROM PYTHON SOURCE LINES 77-79 .. code-block:: Python report._cache .. rst-class:: sphx-glr-script-out .. code-block:: none {('report', 'test', 'decision_function', None): array([[ 3.05225207, -3.05225207], [ 2.88216375, -2.88216375], [-3.65339324, 3.65339324], ..., [ 3.24317783, -3.24317783], [-5.10421486, 5.10421486], [-3.05419093, 3.05419093]], shape=(30000, 2)), ('report', 'test', 'predict', None): array([0, 0, 1, ..., 0, 1, 1], shape=(30000,)), ('report', 'test', 'predict_proba', None): array([[0.95487966, 0.04512034], [0.94695765, 0.05304235], [0.02524906, 0.97475094], ..., [0.96242719, 0.03757281], [0.00603447, 0.99396553], [0.04503688, 0.95496312]], shape=(30000, 2)), ('report', 'test', 'predict_log_proba', None): array([[-0.04616996, -3.09842203], [-0.05450091, -2.93666465], [-3.67896652, -0.02557328], ..., [-0.03829686, -3.28147469], [-5.11026761, -0.00605275], [-3.10027349, -0.04608256]], shape=(30000, 2)), ('metrics', 'test', 'accuracy', ('mapping', ())): 0.9345} .. GENERATED FROM PYTHON SOURCE LINES 80-82 The cache stores predictions by type and data source. This means that computing metrics that use the same type of predictions will be faster. .. GENERATED FROM PYTHON SOURCE LINES 84-89 Caching with :class:`~skore.CrossValidationReport` ================================================== Here we will demonstrate that :class:`~skore.CrossValidationReport` also benefits from caching. .. GENERATED FROM PYTHON SOURCE LINES 89-92 .. code-block:: Python report = evaluate(model, X=X, y=y, splitter=3, n_jobs=3) report.help() .. raw:: html


.. GENERATED FROM PYTHON SOURCE LINES 93-96 A :class:`~skore.CrossValidationReport` is essentially a list of :class:`~skore.EstimatorReport`, one for each split, so caching on the splits makes the calculation on the :class:`~skore.CrossValidationReport` faster as well. .. GENERATED FROM PYTHON SOURCE LINES 97-102 .. code-block:: Python start = time.time() report.metrics.summarize().frame() end = time.time() print(f"Time taken: {end - start:.2f} seconds") .. rst-class:: sphx-glr-script-out .. code-block:: none Time taken: 0.13 seconds .. GENERATED FROM PYTHON SOURCE LINES 103-104 The subsequent calls are even faster because the metrics themselves are cached: .. GENERATED FROM PYTHON SOURCE LINES 105-110 .. code-block:: Python start = time.time() report.metrics.summarize().frame() end = time.time() print(f"Time taken: {end - start:.2f} seconds") .. rst-class:: sphx-glr-script-out .. code-block:: none Time taken: 0.03 seconds .. GENERATED FROM PYTHON SOURCE LINES 111-113 By keeping the estimator together with the data, we are able to trade off some memory space for faster operations. .. rst-class:: sphx-glr-timing **Total running time of the script:** (0 minutes 14.374 seconds) .. _sphx_glr_download_auto_examples_technical_details_plot_cache_mechanism.py: .. only:: html .. container:: sphx-glr-footer sphx-glr-footer-example .. container:: sphx-glr-download sphx-glr-download-jupyter :download:`Download Jupyter notebook: plot_cache_mechanism.ipynb ` .. container:: sphx-glr-download sphx-glr-download-python :download:`Download Python source code: plot_cache_mechanism.py ` .. container:: sphx-glr-download sphx-glr-download-zip :download:`Download zipped: plot_cache_mechanism.zip ` .. only:: html .. rst-class:: sphx-glr-signature `Gallery generated by Sphinx-Gallery `_