Evaluation practice · E07 · Implementation

A reusable ML pipeline starts before the estimator

Scaling and PCA are learned components too. This archive example gets the fit/transform boundary right while showing why a date filter is not automatically a forward test.

scikit-learn StandardScaler / PCAPandasCustom estimator wrapper
A fitted coordinate system changes every downstream example. The projection illustrates why preprocessing must be learned on training data only.
Figure 1. Learning the coordinate system. A fitted coordinate system changes every downstream example. The projection illustrates why preprocessing must be learned on training data only. Illustrative 3D projection. Original vector illustration.

Follow the information

From input to outcome

Training rows determine preprocessing and the fitted estimator. Evaluation rows pass through the stored transform. The lower route supplies the evaluation data, not additional fitting observations.

Training rows determine preprocessing and the fitted estimator. Evaluation rows pass through the stored transform. The lower route supplies the evaluation data, not additional fitting observations.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A fitted coordinate system changes every downstream example. The projection illustrates why preprocessing must be learned on training data only. The module map and layer-level figures below expand the operations in this route.

A reusable ML pipeline starts before the estimator: architectureFeature table: Numeric columns + dates → Partition: Training / evaluation rows → Train-fitted scaler: Transform both partitions → Optional PCA: Fit on training only → Estimator: Predict evaluation labels. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E07 / MODULE MAP01 INPUTFeature tableNumeric columns + dates02 MODULEPartitionTraining / evaluation rows03 MODULETrain-fitted scalerTransform both partitions04 MODULEOptional PCAFit on training only05 OUTPUTEstimatorPredict evaluation labels
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Feature table — Numeric columns + dates

The architecture in context

The system we are building

A useful tabular baseline is a pipeline rather than a naked classifier. Feature columns must retain the same order, the scaler must retain training statistics, and an optional PCA transform must retain the same coordinate system at inference. Otherwise the estimator receives a different problem from the one it learned.

Who does what in the stack

scikit-learn StandardScaler / PCA
Learn a training-defined coordinate system.
Pandas
Preserves date metadata and feature columns.
Custom estimator wrapper
Stores preprocessing and model state together.

The wrapper stores preprocessing objects alongside train and test matrices and supports alternative estimators. The inspected scaling code fits only the training matrix and transforms the held-out matrix with the fitted object. The PCA path similarly separates fitting from application. These are practical building blocks for a reusable baseline.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Fit the transform on the training rows

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

A scaler and PCA basis are learned components even though they are not neural layers. Their fitted state must be saved with the estimator, in the same feature order. The inspected standardization and PCA fits use training rows. However, a date complement can put later observations into training when a historical block is held out; that is not equivalent to a forward-only prediction task.

The mathematical contract

z=(x−μtrain)/strain,u=VtrainTzz=(x-\mu_{\rm train})/s_{\rm train},\qquad u=V_{\rm train}^Tz

PCA can reduce downstream optimization cost but maximizes explained input variance, not label usefulness. For an SVM, prediction cost depends on retained support vectors as well as feature dimension. Scaling changes the geometry seen by regularization and kernels, so it belongs inside validation rather than being tuned on the full dataset.

Implementation and resource card

Capacity / budget
Dimensions, support-vector count and estimator capacity depend on the selected branch. A source file containing multiple estimators is not one fixed model card.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Alter only held-out feature values and verify scaler and PCA state remain identical. Then inspect min/max timestamps of both partitions. A random fallback or future-containing complement should be labeled explicitly rather than called chronological generalization.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The important distinction in the snippet is fit_transform versus transform. The first estimates moments; the second uses existing moments. Dropping Date after preserving it for reporting also prevents a timestamp column from entering numeric preprocessing accidentally.

Python · file · lines 752–769
        self.test_dates = self.X_test_unscaled['Date']
        self.X_test_unscaled = self.X_test_unscaled.drop(columns=['Date'])
        self.X_train_unscaled = self.X_train_unscaled.drop(columns=['Date'])
        
        #print("DEBUG: ", self.X_train_unscaled.to_string())

        # Scale the data
        self.scaler = StandardScaler()
        self.X_train = self.scaler.fit_transform(self.X_train_unscaled)
        self.X_test = self.scaler.transform(self.X_test_unscaled)

        if bool_pca_features:
            # Perform PCA transformation
            print("transform_params = ", self.transform_params)

            if self.transform_params == None:
                pca = PCA(n_components=min(len(self.X.columns), len(self.X_train)) -1) # -1 because of the Date-column in X.columns    

Verbatim archive excerpt from ml_model_svm.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

The surrounding splitter can use the complement of a test date interval as training data, including later dates. It also contains a random-split fallback. Neither establishes forward-only forecasting. A correct preprocessing boundary cannot repair an unsuitable chronology.

Keep building

Other posts of interest