Evaluation practice · E06 · Evaluation

Give a classifier an abstain option—and account for it

A probability-to-decision layer separates model confidence from action. Its denominator is as important as its threshold.

NumPyPandas / metadataCustom evaluation utility
An abstain region trades prediction coverage against the behavior of accepted decisions; it must remain visible in evaluation.
Figure 1. A decision can abstain. An abstain region trades prediction coverage against the behavior of accepted decisions; it must remain visible in evaluation. Illustrative probability distribution. Original vector illustration.

Follow the information

From input to outcome

Predictions first pass numerical validity checks and then decision thresholds. Labels and grouping identities are used for evaluation; abstained examples must remain visible in coverage accounting.

Predictions first pass numerical validity checks and then decision thresholds. Labels and grouping identities are used for evaluation; abstained examples must remain visible in coverage accounting.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: An abstain region trades prediction coverage against the behavior of accepted decisions; it must remain visible in evaluation. The module map and layer-level figures below expand the operations in this route.

Give a classifier an abstain option—and account for it: system and evaluation mapImages + metadata: Labels · entity · date → Classifier: Two probabilities → Validity screen: Finite / range checks → Decision thresholds: Negative · abstain · positive → Evaluation table: Accuracy + coverage. A high-level module map; comparison branches and training details are explained in the article.EVALUATION PRACTICE / E06 / MODULE MAP01 INPUTImages + metadataLabels · entity · date02 MODULEClassifierTwo probabilities03 MODULEValidity screenFinite / range checks04 MODULEDecision thresholdsNegative · abstain · positive05 OUTPUTEvaluation tableAccuracy + coverage
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Images + metadata — Labels · entity · date

The architecture in context

What this comparison asks

A classifier need not make a committed decision on every sample. The evaluation utility joins image paths, labels, identities and predicted probabilities, then applies lower and upper thresholds to the probability of class one. The interval between them is a deliberate abstention region. This adapter belongs after the model: the network estimates a score, while a separate rule decides how to use it.

Who does what in the stack

NumPy
Probability checks and threshold decisions.
Pandas / metadata
Joins outcomes to dates and entities.
Custom evaluation utility
Owns abstention semantics and reporting.

The custom evaluator carries source metadata into grouped analysis and converts one-hot labels to class indices. That makes it possible to ask whether a threshold works consistently across dates or merely selects a favorable subset. It is an evaluation interface rather than an architecture change.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

A classifier is not yet a decision rule

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

A prediction can be correct, wrong or not acted on. The evaluator turns a probability vector into a class and optionally abstains. Accuracy on accepted examples conditions on that selection. It must be paired with the eligible denominator and abstention rate. The class-range filter in the source can remove records before scoring, so the denominator needs to be reconciled with the original prediction batch.

The mathematical contract

y^=arg⁡max⁡kpk,a=1[max⁡kpk≥t],coverage=∑iaiN\widehat y=\arg\max_k p_k,\quad a=\mathbf 1[\max_k p_k\geq t],\quad \mathrm{coverage}=\frac{\sum_i a_i}{N}

A higher threshold can raise conditional precision while rejecting most useful opportunities. Probability calibration and decision utility are separate from classification accuracy. Choose threshold and class mapping on a development partition, then carry them unchanged into evaluation; otherwise the reported operating point is another fitted model.

Implementation and resource card

Capacity / budget
No additional neural parameters or training epochs. The threshold is a decision parameter; threshold selection consumes validation evidence.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Use a three-row probability matrix containing a confident correct prediction, a confident error and a low-confidence case. Compute all counts by hand. Add an invalid class and require an explicit rejected-record count rather than silently improving the confusion matrix.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt converts predictions to float32, filters rows whose probabilities lie outside [0,1], and extracts class-one probabilities. This makes malformed predictions visible in code, but removing them changes the evaluated population. Range checks alone also do not establish that class probabilities sum to one.

Python · file · lines 178–198
        # Check predictions for validity (in range [0, 1])
        valid_mask = np.all((predictions >= 0) & (predictions <= 1), axis=1)
        valid_predictions = predictions[valid_mask]
        valid_labels = labels[valid_mask]
        valid_tickers = np.array(tickers)[valid_mask]
        valid_dates = np.array(dates)[valid_mask]

        # Extract probabilities for class 1
        valid_probabilities = valid_predictions[:, 1]

        # Determine predicted classes based on thresholds
        if lower_threshold is None and upper_threshold is None:
            # No thresholds, use argmax to determine classes
            predicted_classes = np.argmax(valid_predictions, axis=1)
        else:
            # Apply thresholds to get predicted classes
            predicted_classes = pd.cut(valid_probabilities, 
                                    bins=[-float('inf'), lower_threshold, upper_threshold, float('inf')], 
                                    labels=[0, 'unclassified', 1]).astype(str)

        # Create a DataFrame

Verbatim archive excerpt from cnn_eval.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

Accuracy among accepted predictions can rise while useful coverage collapses. Invalid, abstained and scored rows need separate counts against the same original sample total. This historical classifier analysis is not an execution-qualified strategy or a profitability claim.

Keep building

Other posts of interest