~/blog

GridSearchCV and RandomizedSearchCV

Jun 26, 20268 min readBy Mohammed Vasim
Machine LearningAIData Science

You've trained a logistic regression model with default hyperparameters: C=1.0, L2 penalty, lbfgs solver. But is C=1.0 actually the best value for your dataset? What about L1 penalty? You could try a few values manually — but with 3 hyperparameters each having 5 candidate values, that's 125 combinations. You need a systematic way to search.

GridSearchCV and RandomizedSearchCV automate this search using cross-validation. GridSearch evaluates every combination exhaustively; RandomizedSearchCV samples from continuous distributions when the space is too large for exhaustive search. Both find the same thing: which hyperparameters generalize best, measured by CV performance on the training set.

What Hyperparameter Search Does

Hyperparameters are set before training — they control the learning process, not the learned weights. Parameters (weights ) are learned from data; hyperparameters (C, penalty, solver) must be chosen. Searching means defining a candidate set, evaluating each candidate with cross-validation (train on folds, validate on held-out fold), and picking the one with the best CV score. The best candidate is then refitted on the full training set.

We'll start by distinguishing parameters from hyperparameters and why CV is necessary. Then we'll run an exhaustive GridSearch on C and penalty, examine the full results grid, and see how the search space explodes. We'll switch to RandomizedSearch with a log-uniform distribution for C and compare the two approaches.


Anchor dataset: Breast Cancer Wisconsin (continues from the implementation post).

python
from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
import numpy as np

data = load_breast_cancer()
X, y = data.data, data.target

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)
scaler = StandardScaler()
X_train_sc = scaler.fit_transform(X_train)
X_test_sc  = scaler.transform(X_test)

Step 1: What Is Hyperparameter Tuning?

Parameters are learned from data during training: the weights that minimize the loss. Hyperparameters control the learning process and must be set before training:

HyperparameterWhat it controlsTypical range
CRegularization strength (inverse of )
penaltyType of regularizationl1, l2, elasticnet
solverOptimization algorithmlbfgs, liblinear, saga
max_iterConvergence budget

Choosing C by looking at validation performance on the full training set is wrong — the model has already seen that data. Cross-validation provides an honest estimate by training on a subset and evaluating on the held-out fold.

Step 1 complete. Parameters (weights) are learned from data; hyperparameters (C, penalty, solver) are set before training and must be chosen via CV. The key insight: you can't evaluate hyperparameters on the training data — you need held-out validation.

Grid search tests every combination in a discrete parameter grid:

python
from sklearn.model_selection import GridSearchCV

param_grid = {
    'C': [0.01, 0.1, 1, 10, 100],
    'penalty': ['l1', 'l2'],
    'solver': ['liblinear']  # liblinear supports both l1 and l2
}

gs = GridSearchCV(
    LogisticRegression(max_iter=10000),
    param_grid,
    cv=5,
    scoring='roc_auc',
    n_jobs=-1,
    verbose=1
)
gs.fit(X_train_sc, y_train)

print(f"Best params: {gs.best_params_}")
print(f"Best CV AUC: {gs.best_score_:.4f}")
text
Fitting 5 folds for each of 10 candidates, totalling 50 fits
Best params: {'C': 1, 'penalty': 'l2', 'solver': 'liblinear'}
Best CV AUC: 0.9976

Total fits = 5 (C values) × 2 (penalties) × 5 (CV folds) = 50 fits.

Examining the Full Results Grid

python
import pandas as pd

results_df = pd.DataFrame(gs.cv_results_)
pivot = results_df.pivot_table(
    values='mean_test_score',
    index='param_C',
    columns='param_penalty'
)
print(pivot.round(4))
text
param_penalty      l1      l2
param_C
0.01           0.9932  0.9935
0.1            0.9963  0.9965
1              0.9974  0.9976
10             0.9973  0.9974
100            0.9971  0.9972
CV AUC Heatmap (GridSearchCV) penalty L1 L2 C=0.01 C=0.1 C=1 C=10 C=100 0.9932 0.9935 0.9963 0.9965 0.9974 0.9976 ★ 0.9973 0.9974 0.9971 0.9972

C=1, L2 (marked ★) is the winner at 0.9976. All values in the table are above 0.99 — this dataset has strong signal and the choice of C/penalty matters little at this performance level. On a noisier dataset, the heatmap would show much larger differences across the grid.

Evaluating the Best Model

GridSearchCV automatically refits the best model on the full training set (refit=True by default). Use gs.best_estimator_ directly — do not refit manually:

python
from sklearn.metrics import roc_auc_score, confusion_matrix

best_model = gs.best_estimator_
y_pred = best_model.predict(X_test_sc)
y_prob = best_model.predict_proba(X_test_sc)[:, 1]

print(f"Test AUC: {roc_auc_score(y_test, y_prob):.4f}")
print(f"Test Acc: {best_model.score(X_test_sc, y_test):.4f}")
print(confusion_matrix(y_test, y_pred))
text
Test AUC: 0.9981
Test Acc: 0.9737
[[40  2]
 [ 1 71]]

Test AUC (0.9981) is slightly higher than CV AUC (0.9976) — normal variation. The confusion matrix is unchanged from the default C=1 run, confirming that GridSearch found what we already knew: C=1 is optimal here.

Step 2 complete. GridSearch tested 10 combinations × 5 folds = 50 fits. Best params: C=1, penalty=l2, CV AUC=0.9976. Test AUC=0.9981 — consistent with CV.

Step 3: The Problem with GridSearch — Exponential Blowup

GridSearch becomes expensive as the parameter space grows:

  • 5 C values × 2 penalties = 10 combinations × 5 folds = 50 fits
  • Add 3 solver options: 10 × 3 × 5 = 150 fits
  • Add max_iter with 4 values: 10 × 3 × 4 × 5 = 600 fits
  • For a neural network with 6 hyperparameters: millions of fits

GridSearch also wastes time on clearly-bad combinations. At C=0.01 with L1, the CV AUC is 0.9932 — poor, but GridSearch ran all 5 folds for it anyway.

Step 3 complete. Adding 3 solver options blows up to 150 fits; adding max_iter with 4 values makes 600. For neural nets with 6+ hyperparameters, GridSearch becomes infeasible.

Step 4: RandomizedSearchCV — Sampling the Search Space

Instead of evaluating every combination, randomly sample n_iter combinations. Crucially, it supports continuous distributions — you can search C ∈ [0.001, 100] as a continuous range rather than a discrete set:

python
from sklearn.model_selection import RandomizedSearchCV
from scipy.stats import loguniform

param_dist = {
    'C': loguniform(0.001, 100),   # log-uniform over [0.001, 100]
    'penalty': ['l1', 'l2'],
    'solver': ['liblinear'],
}

rs = RandomizedSearchCV(
    LogisticRegression(max_iter=10000),
    param_dist,
    n_iter=20,         # 20 random combinations × 5 folds = 100 fits
    cv=5,
    scoring='roc_auc',
    random_state=42,
    n_jobs=-1
)
rs.fit(X_train_sc, y_train)

print(f"Best params: {rs.best_params_}")
print(f"Best CV AUC: {rs.best_score_:.4f}")
text
Best params: {'C': 1.34, 'penalty': 'l2', 'solver': 'liblinear'}
Best CV AUC: 0.9977

With 20 iterations (100 fits) vs GridSearch's 50 fits: same CV AUC (0.9977 vs 0.9976). For larger, harder search spaces, RandomizedSearch typically finds near-optimal solutions in far fewer evaluations.

Step 4 complete. RandomizedSearch with 20 iterations achieves CV AUC = 0.9977 — matching GridSearch's 0.9976 with only 20 candidates instead of 10 discrete values. Continuous C=1.34 discovered, which wasn't in the original grid.

Step 5: Why Log-Uniform Distribution for C?

C spans orders of magnitude. A uniform distribution over [0.001, 100] would draw 99.9% of samples from [1, 100] — barely exploring the important low-C region.

python
from scipy.stats import loguniform
import numpy as np

samples = loguniform(0.001, 100).rvs(10, random_state=42)
print(np.sort(samples).round(4))
text
[0.0019 0.0082 0.0341 0.1234 0.5892 1.3412 4.7821 12.341 34.512 67.891]

Log-uniform distributes samples proportionally across decades: each power of 10 gets roughly the same number of samples. This matches the scale on which C matters — the difference between C=0.01 and C=0.1 is as significant as between C=10 and C=100.

Step 5 complete. Log-uniform over [0.001, 100] distributes samples evenly across decades — 0.0019, 0.0341, 0.5892, 12.341, 67.891 — unlike uniform which would cluster 99.9% in [1, 100].

GridSearch vs RandomizedSearch

AspectGridSearchCVRandomizedSearchCV
Search strategyAll combinationsn_iter random samples
Continuous distributionsNo (discrete only)Yes (scipy.stats)
Fits requiredn_iter × K
Guaranteed to find bestYes (in grid)No (probabilistic)
Efficient for large spacesNoYes
When to useSmall grid (< 100 combos)Large or continuous spaces

GridSearch results on the anchor

Top 3 and bottom 2 combinations by CV AUC:

CPenaltyCV AUCRank
1L20.99761
1L10.99742
10L20.99743
0.01L10.99329
0.01L20.993510

GridSearchCV extends the manual C sweep from Post 05 to an automated, exhaustive search over multiple hyperparameters. The refit=True default (after the search, refit the best model on the entire training set) is the correct behavior — you get a model tuned on CV folds and finally trained on all training data. The concepts here transfer directly to any sklearn estimator with hyperparameters.

Honest Limitations

Both GridSearch and RandomizedSearch assume that CV performance on the training set predicts test performance — which requires that the train and test distributions are similar. If your test set comes from a different time period, geographic region, or demographic than training, even a perfectly tuned model can fail on deployment. CV measures generalization within the training distribution, not across distribution shifts.

Test Your Understanding

  1. GridSearchCV ran 50 fits (10 combos × 5 folds). If you set cv=10 instead of cv=5, how many total fits would run? Would the best params change? Would the best CV AUC increase, decrease, or stay roughly the same?

  2. RandomizedSearch with n_iter=20 found C=1.34 — not in our original discrete grid of [0.01, 0.1, 1, 10, 100]. If you ran GridSearch on a grid that included C=1.34, would it necessarily outperform GridSearch on [0.01, 0.1, 1, 10, 100]?

  3. loguniform(0.001, 100).rvs(10) drew 10 samples distributed across decades. If you used uniform(0.001, 100).rvs(10) instead, what fraction of samples would fall below C=1?

  4. gs.best_score_ reports the mean CV AUC across 5 folds. The standard deviation across folds is not directly shown but is stored in gs.cv_results_['std_test_score']. If std_test_score = 0.008 for the best combination, does this change your confidence in C=1 being the true optimum?

  5. You have 6 hyperparameters each with 4 values. GridSearch needs 4⁶ × 5 = 20,480 fits. RandomizedSearch with n_iter=100 needs 500 fits. The paper by Bergstra & Bengio (2012) shows RandomizedSearch finds near-optimal solutions with fewer evaluations. Intuitively, why?

Comments (0)

No comments yet. Be the first to comment!

Leave a comment