Evaluation practice · E08 · Evaluation audit

When your test set quietly becomes a validation set

A compact model-selection loop makes an important boundary visible: the data that chooses the winner cannot also provide its untouched final score.

scikit-learn metricsCustom predictor adapterSerialization layer
The moment a score selects a model, that data has entered a feedback loop and no longer serves as an untouched test.
Figure 1. The test set has a job. The moment a score selects a model, that data has entered a feedback loop and no longer serves as an untouched test. Selection schematic. Original vector illustration.

Follow the information

From input to outcome

Candidate scores feed the selection decision. Reusing that partition to choose a winner makes it validation data; serialization does not restore its independence.

Candidate scores feed the selection decision. Reusing that partition to choose a winner makes it validation data; serialization does not restore its independence.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The moment a score selects a model, that data has entered a feedback loop and no longer serves as an untouched test. The module map and layer-level figures below expand the operations in this route.

When your test set quietly becomes a validation set: system and evaluation mapConfigurations: Estimator + parameters → Fit candidate: Training partition → Score candidate: Selection partition → Select maximum: Consumed feedback → Serialize winner: Model + preprocessing. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E08 / MODULE MAP01 INPUTConfigurationsEstimator + parameters02 MODULEFit candidateTraining partition03 MODULEScore candidateSelection partition04 MODULESelect maximumConsumed feedback05 OUTPUTSerialize winnerModel + preprocessing
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Configurations — Estimator + parameters

The architecture in context

What this comparison asks

The archived utility evaluates configurations through a common predictor interface. Each candidate is fitted, produces predictions and returns an accuracy score; a small outer loop keeps the largest score. The design is convenient because different estimator families can share orchestration and serialization instead of duplicating experiment code.

Who does what in the stack

scikit-learn metrics
Computes candidate accuracy.
Custom predictor adapter
Fits alternative configurations through one interface.
Serialization layer
Must retain model, preprocessing and feature schema.

The project-specific part is the adapter between parameter dictionaries and the predictor’s internal methods, plus the saved-model path. This is useful engineering glue, but its naming should reflect how evidence is used. Once a split ranks configurations, that split is serving model selection.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Model selection is itself a fitted procedure

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The source uses y_test while selecting the best candidate. Whatever the variable is named, those labels then serve as validation data. Serialization preserves the chosen object but does not restore independence. The saved artifact also needs the preprocessing state and feature schema used to produce its inputs.

The mathematical contract

m^=arg⁡max⁡m∈MS(m,Dselect)\widehat m=\arg\max_{m\in\mathcal M} S(m,D_{\rm select})

Increasing the search space can overfit the selection partition without changing the final estimator’s apparent size. A small model chosen from thousands of tries is a larger statistical selection procedure than a single frozen model. An untouched final set evaluates that entire procedure.

Implementation and resource card

Capacity / budget
Compute scales with candidate fits, not merely final-model parameters. The selected file does not establish an untouched test score.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Trace every call that reads the selection labels. Save the selected hyperparameters, candidate count and split identifiers with the model. Reload it in a clean process and compare predictions on a fixed synthetic feature matrix before declaring the artifact portable.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt computes accuracy against y_test inside evaluate_model_configuration, then calls that function repeatedly while updating best_score. The issue is not the formula for accuracy. It is the feedback path from the nominal test labels into the choice of configuration.

Python · file · lines 71–96
def evaluate_model_configuration(predictor, classifier_params):
    """
    Fit and evaluate the model based on the given classifier parameters.
    Returns the score and the parameters for comparison.
    """
    print(f"{datetime.now()}: {classifier_params.get('type')}: Evaluating configuration: {classifier_params['params']}...")

    predictor.ml_model._create_and_fit_model(classifier_params=classifier_params)
    predictor.ml_model._predict()
    score = accuracy_score(predictor.ml_model.y_test, predictor.ml_model.y_pred)

    print(f"DONE {classifier_params.get('type')}: {classifier_params['params']}, Score: {score}")

    return score, classifier_params


def optimize_classifier_group(predictor, classifier_type, configurations):
    print(f"Starting optimization for {classifier_type} with {len(configurations)} configurations...")
    
    # For other classifiers, proceed with the existing sequential process
    best_score, best_config = -1, None
    for config in configurations:
        current_score, _ = evaluate_model_configuration(predictor, {'type': classifier_type, 'params': config})
        if current_score > best_score:
            best_score = current_score
            best_config = config

Verbatim archive excerpt from fit_and_evaluate.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

The selected maximum is an optimistic description of the search outcome, not an independent estimate of the chosen model. Re-labeling the same split after the fact does not restore independence. A fresh final partition must remain outside this loop.

Keep building

Other posts of interest