Version 1.10#
Legend for changelogs
Major Feature something big that you couldn’t do before.
Feature something that you couldn’t do before.
Efficiency an existing feature now may not require as much computation or memory.
Enhancement a miscellaneous minor improvement.
Fix something that previously didn’t work as documented – or according to reasonable expectations – should now work.
API Change you will need to change your code to have the same effect in the future; or a feature will be removed in the future.
Version 1.10.dev0#
September 2026
Changed models#
preprocessing.QuantileTransformerandpreprocessing.quantile_transformnow estimate quantiles with theaveraged_inverted_cdfmethod and no longer capn_quantilesatn_samples. Dense subsampling is also done with replacement. These changes affectquantiles_even whensample_weightis not passed. By Karassay Raushanbek and Shruti Nath. #32761inspection.partial_dependenceandinspection.PartialDependenceDisplaynow usenumpy.nanquantile’s linear interpolation instead of the (unmaintained)scipy.stats.mstats.mquantiles’s Cunnane interpolation to compute the grid of values and deciles fromX. This may slightly shift the computed grid values and displayed deciles.np.nanvalues in a numerical column are now ignored (instead of being sorted last in the quantile calculation and hence possibly affecting the result depending on the requested percentiles). By Arthur Lacote. #34928
Support for Array API#
Additional estimators and functions have been updated to include support for all Array API compliant inputs.
See Array API support (experimental) for more details.
Feature
covariance.LedoitWolfnow supports array API compatible inputs. By Tim Head #33573Feature
covariance.OASnow supports array API compatible inputs. By Bruno Aristimunha #33600Feature
linear_model.LogisticRegressionCVnow supports array API compatible inputs withsolver="lbfgs"andsolver="newton-cg"and when the underlying scoring function is also compatible with the array API. By Omar Salman and Christian Lorentzen. #33906Feature
sklearn.metrics.matthews_corrcoefnow supports array API compatible inputs. By Mohamad Fazeli. #34423Feature
sklearn.metrics.cluster.contingency_matrixnow supports array API compatible inputs. By Yaroslav Korobko and Omar Salman. #34785Enhancement Add support for array inputs from mixed namespaces and devices for:
sklearn.metrics.det_curve,sklearn.metrics.roc_curve,sklearn.metrics.zero_one_loss,sklearn.metrics.jaccard_score,sklearn.metrics.balanced_accuracy_score,sklearn.metrics.cohen_kappa_score. By Lucy Liu. #34442Enhancement
linear_model.LogisticRegressionnow supports array API compatible inputs withsolver="newton-cg". By Christian Lorentzen and Omar Salman. #34412Fix
model_selection.GridSearchCV,model_selection.RandomizedSearchCV,model_selection.cross_validateandmodel_selection.cross_val_scoreno longer forceyintoX’s array namespace during cross-validation. Namespace and device handling ofyis delegated to the wrapped estimator, so an array APIXcan be combined with a NumPyy, including a NumPy stringywhich cannot be held by array API namespaces. By Tim Head. #33633Fix Fixed
metrics.median_absolute_errorto handle integer inputs when the input array namespace does not support integer inputs to itsquantilefunction (e.g., PyTorch). By Lucy Liu. #33778Fix Fixes support when
average="samples"andsample_weightis an array from a namespace or device different fromy_pred, for:sklearn.metrics.precision_recall_fscore_support,sklearn.metrics.recall,sklearn.metrics.precision,sklearn.metrics.f1_score,sklearn.metrics.fbeta_score. Also fixes several bugs in array API compliance ofsklearn.metrics.roc_curve. By Lucy Liu. #34442Fix
discriminant_analysis.LinearDiscriminantAnalysisnow supports array API inputs whereyis from a different namespace or device thanX, including string labels. By Kiyarash Fazeli. #34475Fix
linear_model.PoissonRegressornow acceptsyandsample_weightfrom a different array API namespace or device thanX. By Amineh Dadsetan. #34480Fix Fixed classification metrics to accept pandas labels when array API dispatch is enabled and predictions are from a different array API namespace or device. By Olivier Grisel. #34779
Metadata routing#
Refer to the Metadata Routing User Guide for more details.
Feature Feature Added support for auto-requesting metadata through
set_config(enable_metadata_auto_requests=True). This enables automatic routing of common metadata likesample_weight, orX_val. Consumers (estimators, scorers, and CV splitters) and consuming routers can now define which metadata requests to set automatically via theadd_auto_requestmethod, which complements the existing class-level defaults set through__metadata_request__*attributes. Refer to Default and Auto-Requested Metadata for more details. By Adrin Jalali and Stefanie Senger. #31413
Callbacks#
API Change In
callback.ScoringMonitor, the option to providescoring=Noneto have the default scorer of the estimator has been removed. By François Paugam. #34590
sklearn.calibration#
Feature
calibration.calibration_curve,calibration.CalibrationDisplay.from_estimator, andcalibration.CalibrationDisplay.from_predictionsnow supportn_bins="cube_root". This option sets the number of bins toceil(n_samples**(1/3)). By Zeyu Sun #33856Fix
calibration.CalibratedClassifierCVwithmethod="sigmoid"now converts predict_proba outputs to logits (log(p/(1-p))) before calibration instead of passing probabilities directly to the calibrator. Formethod="temperature", predict_proba outputs are converted to multinomial logits. Estimators that only expose decision_function are unaffected. Calibration now prefers predict_proba over decision_function when both are available. By Olivier Grisel and Antoine Baker. #34313
sklearn.cluster#
Efficiency
cluster.AgglomerativeClusteringandcluster.FeatureAgglomerationwithlinkage="single", as well ascluster.HDBSCAN, are now faster for inputs that produce deep union-find trees. Restored path compression avoids quadratic time while constructing the single-linkage hierarchy. By Chris Boseak. #34872
sklearn.datasets#
Fix
datasets.fetch_openmlwithparser="pandas"no longer raises aTypeErrorwhen a nominal column’s values are inferred by pandas as numeric or boolean categories instead of strings. By Dea María Léon. #34902
sklearn.discriminant_analysis#
Efficiency
discriminant_analysis.LinearDiscriminantAnalysisis faster to fit: the target is ordinal-encoded once infitand the number of classes is passed down to the internal helpers, removing a redundant pass over the target to recompute the unique class labels. Computing the class means is up to 1.4x faster for large numbers of samples and classes. By Kiyarash Fazeli. #34475
sklearn.ensemble#
Efficiency Part of the fit of
sklearn.ensemble.GradientBoostingClassifierwas optimized. The speed-up is mainly visible for deeper trees, which is not the primary use case for gradient boosting but can occur in grid searches. For very deep trees, this can be up to 20 times faster. By Arthur Lacote #32911Efficiency Fitting in
sklearn.ensemble.HistGradientBoostingClassifierandsklearn.ensemble.HistGradientBoostingRegressorwas optimized to be faster, by speeding up assignment of bins. By Itamar Turner-Trauring. #34194Efficiency Improved the performance of binning in
ensemble.HistGradientBoostingClassifierandensemble.HistGradientBoostingRegressor. Fitting with sample weights could be slow for datasets with fewer than 200,000 samples; this is now fixed. Binning without sample weights was also optimized. By Arthur Lacote. #34248Efficiency Prediction in
sklearn.ensemble.HistGradientBoostingClassifierandsklearn.ensemble.HistGradientBoostingRegressoris faster when using categoricals. By Itamar Turner-Trauring. #34486Efficiency
sklearn.ensemble.RandomForestClassifertraining is now faster on sparse datasets. By Itamar Turner-Trauring. #34586API Change The
sklearn.experimental.enable_hist_gradient_boostingmodule is deprecated and will be removed in 1.12. It is no longer needed sinceensemble.HistGradientBoostingClassifierandensemble.HistGradientBoostingRegressorare stable and can be imported normally fromsklearn.ensemble. By Guillaume Lemaitre. #34236
sklearn.feature_selection#
Fix Feature selectors such as
feature_selection.SelectKBestnow correctly reject non-finite values (infinity, and NaN unless allowed) intransformwhen usingset_output(transform="pandas"). By Abhimanyu Singh Shekhawat. #34513
sklearn.impute#
API Change
impute.IterativeImputeris no longer experimental. It can now be imported directly withfrom sklearn.impute import IterativeImputer, and importingsklearn.experimental.enable_iterative_imputeris no longer required (doing so now raises a warning and is a no-op). It also no longer raises aexceptions.ConvergenceWarningwhenmax_iteris reached without meeting thetolstopping criterion, as non-convergence of the round-robin imputation is expected. The User Guide now documents why the imputed values are not guaranteed to converge and why investing in better imputation often yields diminishing returns for prediction. By Guillaume Lemaitre. #34214
sklearn.inspection#
Enhancement The parameter
multiclass_colorswas deprecated in favour oftarget_colorsininspection.DecisionBoundaryDisplay. The attributemulticlass_colors_was also renamed totarget_colors_. Now they can be used for binary problems as well without causing confusion (which will be added in a follow-up PR). By Anne Beyer. #34092
sklearn.linear_model#
Major Feature The generalized linear models (GLM)
linear_model.GammaRegressor,linear_model.PoissonRegressorandlinear_model.TweedieRegressorare now able to fit L1 and Elastic-Net penalties with the new parameterl1_ratiotogether with the 2 new solvers"newton-cd"and"newton-cd-gram". These new solvers are also available forlinear_model.LogisticRegression. They are based on the Newton solver infrastructure introduced for"newton-cholesky"and take advantage of the already available and fast (e.g. via gap safe screening rules) coordinate descent solvers ofElasticNet."newton-cd-gram"constructs the full Hessian (gram) matrix and then usesElasticNet(precompute=True)as the inner solver. It is a good choice forn_samples >> n_features."newton-cd"avoids constructing the full Hessian matrix, but formulates the minimization problem for the inner solver as a penalized least squares problem (which amounts to taking the square root of the Hessian) and then usesElasticNet(precompute=False)on it. It is a good choice forn_features > n_samples.
Efficiency The
"newton-cholesky"solver oflinear_model.LogisticRegression,linear_model.GammaRegressor,linear_model.PoissonRegressorandlinear_model.TweedieRegressornow uses an improved backtracking line search. Instead of halving the step size in each line search iteration, it uses cubic or quadratic polynomial interpolation to find the next trial step size. By Christian Lorentzen. #34157Efficiency All solvers of
linear_model.LogisticRegressionCVare now faster on free-threaded Python when using parallelism. The only exceptions are"sag"and"saga"solvers, which will continue to run at the same speed. By Itamar Turner-Trauring. #34659Enhancement
linear_model.PoissonRegressor,linear_model.GammaRegressorandlinear_model.TweedieRegressorgot a new solver:solver="newton-cg". This is the same Newton conjugate gradient solver, a.k.a. truncated Newton, that is already available withlinear_model.LogisticRegression. By Christian Lorentzen. #33759Enhancement
linear_model.ElasticNetCV,linear_model.LassoCVandlinear_model.MultiTaskElasticNetCVas well asenet_pathandlasso_pathare now faster. The (relative) speed-up is larger for lower tolerance (larger values oftol) and achieved by avoiding to re-calculate the residuals and the column norms ofXfor each penaltyalphaalong the path. By Christian Lorentzen. #34572Enhancement
linear_model.ElasticNetCV,linear_model.LassoCVandlinear_model.MultiTaskElasticNetCVnow avoid prematurely convertingXto F-contiguous array or sparse CSC because the inner cross validation take subsamples ofXwhich then need the same conversion again, anyway. So this PR can avoid memory copies ofX. By Christian Lorentzen. #34613Fix The
solver="newton-cg"forlinear_model.LogisticRegression,linear_model.PoissonRegressorand other GLMs is now more robust. Many convergence issues have been fixed, e.g. for certain ill-conditioned problems with low regularization and nearly collinear features. The fix is about treating small or even negative curvature. By Christian Lorentzen. #34412
sklearn.metrics#
Efficiency
metrics.nan_euclidean_distancesis now several times faster on dense data, with a more moderate speed-up forimpute.KNNImputerand estimators usingmetric="nan_euclidean"such asneighbors.NearestNeighbors. By Roman Yurchak. #1140Enhancement
confusion_matrix_at_thresholds,roc_curve,precision_recall_curve,det_curve, androc_auc_scorenow use exact integer cumulative sums for unweighted inputs (sample_weight=None), preventing floating-point precision saturation onfloat32-only array API devices when processing datasets exceeding \(2^{24}\) samples. By Kaynup. #34817Fix Fixed
metrics.average_precision_scoreto handle list-typey_scorefor multiclass data. By Lucy Liu. #34083API Change The default value of the
averageparameter ofprecision_recall_fscore_supportwill change fromNoneto"binary"in version 1.12. By François Paugam. #34190API Change In
metrics.DetCurveDisplaytheestimator_nameparameter is deprecated in favour ofnameand will be removed in 1.12. The**kwargsparameter ofmetrics.DetCurveDisplay.plot,metrics.DetCurveDisplay.from_estimatorandmetrics.DetCurveDisplay.from_predictionsis also deprecated in favour ofcurve_kwargsand will be removed in 1.12. By @AnneBeyer. #34443
sklearn.mixture#
Efficiency
mixture.GaussianMixturewithcovariance_type="tied"is now faster, with the speed-up most visible for many components or high-dimensional data. By Roman Yurchak. #1140
sklearn.model_selection#
API Change
model_selection.HalvingGridSearchCVandmodel_selection.HalvingRandomSearchCVnow expose the cross-validation results of the last halving iteration incv_results_. The full per-iteration history is available in the newall_cv_results_attribute. By Guillaume Lemaitre. #34250
sklearn.neighbors#
Feature
neighbors.NeighborhoodComponentsAnalysisnow supports sparse input matrices forinitin['pca', 'random', 'identity']. By Arturo Amor. #34122Efficiency
neighbors.KNeighborsRegressor.predictis now faster for regression with many output targets. By Roman Yurchak. #1140
Efficiency
sklearn.neighbors.KNeighborsClassifieris now faster when running with multiple threads. The statistics returned fromsklearn.neighbors.BallTreeandsklearn.neighbors.KDTreeno longer reset on each query, and are no longer incorrect when the tree is used from multiple threads. By Itamar Turner-Trauring. #34187Efficiency
sklearn.neighbors.KNeighborsClassifieris now faster when fitting when used withsklearn.neighbors.BallTreeorsklearn.neighbors.KDTree. By Itamar Turner-Trauring. #34564
API Change Getting statistics and number of calls from
sklearn.neighbors.BallTreeandsklearn.neighbors.KDTreeis deprecated. The relevant APIs will be removed in version 1.12. #34242
sklearn.pipeline#
Fix The default value for the
transform_inputparameter ofPipelinewas changed fromNoneto("X_val",)so that the validation set, when passed tofit, is always transformed alongsideX, to prevent easy to miss mistakes. By Jérémie du Boisberranger. #34263
sklearn.preprocessing#
Efficiency Improved fit time of
preprocessing.OneHotEncoderandpreprocessing.OrdinalEncoderon object/string categorical features when category counts are needed, for instance withmin_frequencyormax_categories. By Itamar Turner-Trauring. #34386Efficiency
preprocessing.OneHotEncoder,preprocessing.OrdinalEncoder, andpreprocessing.TargetEncoderare now faster by using a Fortran-contiguous memory layout in their encoding paths. By Arthur Lacote. #34392Enhancement
preprocessing.QuantileTransformercan now handle sample weight values by subsampling using sample weights if subsampling is enabled or otherwise using_weighted_percentile. The default method of calculating percentiles for both weighted and unweighted cases is now fixed to “averaged_inverted_cdf”. Support is only implemented for dense X. By Karassay Raushanbek and Shruti Nath. #32761Fix A few edge cases with missing values and unique knot values in
preprocessing.SplineTransformerare now handled correctly. For instance, if a feature is constant, i.e., has only one unique value at fit time,transformnow always returns all zeros for that feature. By Christian Lorentzen. #33953Fix The warning raised by
preprocessing.OneHotEncodernow distinguishes columns where unknown categories are encoded as the infrequent category from columns where they are encoded as all zeros. By oathbreaker. #34861
sklearn.tree#
Major Feature
tree.DecisionTreeRegressor,tree.DecisionTreeClassifiernow have native support for categorical features for binary classification and single-output regression. All criteria are supported except for'absolute_error'. Categorical features can be specified with thecategorical_featuresparameter. Up to 256 categories per features are supported. By Adam Li, Arthur Lacote, Adrin Jalali and Christian Lorentzen #33354Major Feature
tree.DecisionTreeRegressor,tree.DecisionTreeClassifier,tree.ExtraTreeRegressor, andtree.ExtraTreeClassifiernow have native support for categorical features via thecategorical_featuresparameter. Decision trees support binary classification and single-output regression (all criteria except'absolute_error') with up to 256 categories per feature. Extra trees also support multi-class / multi-output targets with up to2**24 - 1categories via random hash splits. By Adam Li, Arthur Lacote, Adrin Jalali and Christian Lorentzen #33972Feature Added a
fill_colorsparameter totree.plot_treeandtree.export_graphvizto allow user-defined class colors when rendering filled decision trees. By Simon-Martin Schröder. #33810
Code and documentation contributors
Thanks to everyone who has contributed to the maintenance and improvement of the project since version 1.9, including:
TODO: update at the time of the release.