The complete map
61 ways into the implementation.
One entry per selected case study. Shared systems are connected—not counted as independent discoveries.
← ML Engineering home| ID / type | Article | Example illustrated |
|---|---|---|
| E01 Implementation | Build a patch transformer—and make every pixel count | Chart image → Patch extraction → Token embedding → Transformer × 3 → Pooling + MLP → Class probabilities |
| E02 Implementation | Porting a transformer is more than translating layer names | Chart image → Unfold patches → Token projection → Transformer × 4 → Flatten + MLP → Class logits |
| E03 Implementation | One encoder, many entities: building pairwise attention | Entity images + IDs → Shared CNN → Identity + category → All ordered pairs → Pair attention → Pair predictions |
| E04 Implementation audit | Building graph attention from node and edge signals | Node / edge histories → Temporal encoders → Node-edge features → Attention × 2 → Node classifier → Predictions |
| E05 Implementation audit | A padded sequence needs more than zero-valued inputs | Sequence + validity → Multiply by mask → LSTM stack → Dense representation → Classifier |
| E06 Evaluation | Give a classifier an abstain option—and account for it | Images + metadata → Classifier → Validity screen → Decision thresholds → Evaluation table |
| E07 Implementation | A reusable ML pipeline starts before the estimator | Feature table → Partition → Train-fitted scaler → Optional PCA → Estimator |
| E08 Evaluation audit | When your test set quietly becomes a validation set | Configurations → Fit candidate → Score candidate → Select maximum → Serialize winner |
| E09 Implementation | Turning a time series into an image is a modeling decision | Historical table → Context selection → Chart channels → Raster renderer → Image dataset → CNN or transformer |
| E10 Implementation | Evolve a population of networks with one vectorized evaluator | Parameter population → Unflatten parameters → Policy batch → Fitness evaluation → PGPE update |
| E11 Implementation audit | Build a foveal time-series state without borrowing the future | Five-second bars → Resolution bank → OHLCV aggregation → Relative features → Policy input |
| E12 Implementation | Keep an evolving network and its feature schema in sync | Historical features → NEAT configuration → Genome population → Fitness + selection → Winner checkpoint |
| E13 Evaluation | An OOD notebook needs an explicit distribution boundary | Frozen genome → Evaluation cohort → Feature adapter → Inference runner → Grouped outcomes |
| E14 Implementation | Add prioritized replay without changing what a transition means | Sequence observation → Stacked LSTMs → Action + transition → Priority replay → Weighted TD update → Target network |
| E15 Implementation | Separate the three clocks in a recurrent DQN | Observed sequence → LSTM × 2 → Q head → Replay batch → TD target → Optimizer step |
| E16 Upstream reference | Tensorized neuroevolution starts with an upstream dependency map | Genome representation → TensorNEAT upstream → JAX / Flax stack → Task environments → Fitness interface |
| E17 Implementation prototype | Represent variable neural graphs with fixed-size tensors | Bounded genome arrays → Graph scheduling → Input injection → Masked aggregation → Activation + output |
| E18 Implementation | Bridge an evolutionary optimizer and a Flax policy | Setup snapshot → Flax MLP → Five sigmoid outputs → Episode simulator → Population fitness |
| E19 Implementation | Keep the policy snapshot separate from the future tape | Observed bar prefix → Geometry snapshot → Policy observation → Future bar tape → Simulation outcome |
| E20 Implementation | Put the state machine inside Numba, not the experiment design | Aligned bar arrays → Parameter tuple → Numba state machine → Per-case results → Search orchestrator |
| E21 Evaluation | Aggregate a strategy evaluation without losing the quiet days | Daily event records → Sequential state → Session summaries → Monthly aggregation → Diagnostics |
| E22 Evaluation audit | Ten thousand bootstrap draws are still twenty observed days | Observed daily values → IID index sampler → Synthetic annual sums → Tail summaries → Interpretation |
| E23 Implementation | Parallel hyperparameter search needs one experiment ledger | Persistent study → Worker processes → Suggested parameters → Fixed evaluator → Trial result |
| E24 Implementation | A data callback is not a durable dataset | Feed callbacks → Per-symbol buffers → Periodic save check → CSV writers → Offline dataset |
| E25 Implementation | When does an online bar actually become observable? | Incoming source bar → Bucket assignment → Current aggregate → Later bucket arrives → Downstream consumer |
| E26 Implementation | Idempotent event handling starts with the right identity | Transport messages → Typed event parser → Semantic identity → Bounded seen cache → State handlers |
| E27 Implementation | Learn an invariant encoder from neighboring views | MNIST image → Two rotations → Shared CNN encoder → Normalized embeddings → Contrastive objective → Frozen ridge probe |
| E28 Evaluation | Evaluate invariance where the labeled examples are scarce | Pretrained encoder → Canonical support → Frozen features → Ridge readout → Rotated test images |
| E29 Implementation | Train one convolutional block at a time | Two augmented images → Frozen prefix → Current block → Local projection head → NT-Xent loss → Next block |
| E30 Negative result | A local learning rule still has to beat random features | CIFAR-10 images → Representation arms → Frozen feature probes → Accuracy comparison → Decision |
| E31 Implementation | Distill a representation, not just a class prediction | Training image → Student CNN → Feature projection → Teacher feature target → MSE / cross-entropy → Student evaluation |
| E32 Negative result | A smaller student is not automatically a compressed teacher | Teacher features → Compact student → Training arms → Readout choices → Paired evaluation |
| E33 Implementation | Add classes by accumulating statistics over frozen features | Cached DINOv2 features → Fixed Fourier map → Class batches → Sufficient statistics → Ridge solve → Prediction |
| E34 Evaluation | Match the joint readout without claiming universal zero forgetting | Same cached features → Task-wise Gram arm → Joint Gram control → Other head controls → Final accuracy |
| E35 Implementation | Take control of PyTorch’s reverse pass, one block at a time | Four stacked frames → Convolutional blocks → Dense Q head → TD error seed → Reverse block sweep → Optimizer update |
| E36 Negative result | A nine-second RL run can validate plumbing, not competence | Pong observations → CNN Q network → 3,000 training steps → Greedy evaluation → Diagnostics |
| E37 Implementation | Put a learned actuator model between policy and physics | World + force sensors → SAC actor → Actuator model → Physical transition → Reward + next state |
| E38 Implementation | Make each checkpoint an experiment receipt | Training callback → Policy state → Safetensors snapshot → Frozen validation → Receipt JSON |
| E39 Negative result | Prediction gains do not automatically become control gains | Frozen actuator arms → SAC training seeds → Selected policies → Held-out resets → Joint capability gate |
| E40 Implementation | Infer a state-space model’s poles before fitting its readout | Training output sequences → Hankel snapshot pairs → Truncated SVD → DMD eigenvalues → Complex recurrence → Ridge readout |
| E41 Evaluation | Stress a dynamical model along two independent axes | Synthetic driven dynamics → Noise panel → Representation arms → Fit on 96 steps → Evaluate at 96 / 192 |
| E42 Implementation | Build a sinusoidal coordinate network with the right initialization | Spatial coordinates → First sine layer → Hidden sine layers → Linear field head → Training loss |
| E43 Implementation | Differentiate a field four times without losing the graph | Coordinates → Tanh network → Nested derivatives → Physics residual → Boundary + data loss |
| E44 Implementation audit | Package an operator-aware basis as a PyTorch layer | Physical coordinates → Cardinal grid map → Local trig atoms → Global sin / cos → Concatenated basis |
| E45 Engineering practice | A research project needs a current state, not just a long history | Experiment protocol → Run artifacts → Verification → Current state ledger → Next authorized action |
| E46 Engineering practice | Turn an overbroad memory claim into a regression test | Old observation → Ridge state → Conflicting new label → Pooled statistics → Updated optimum |
| E47 Evaluation | A compressed memory needs an answer budget as well as a byte budget | Observation blocks → Native accumulator → Retained capsule → Later model query → Answer gate |
| E48 Implementation audit | Collect a transformer’s internal state at the correct boundary | Prompt tokens → Pretrained GPT-2 → KV cache extraction → Query reconstruction → Offline record |
| E49 Engineering practice | Design a KV-compression benchmark before designing a codec | Per-head K / V → Candidate codec → Reconstructed K / V → Attention proxy → Accounting |
| E50 Implementation audit | Four-bit arithmetic is not yet a four-bit representation | Float tensor → Symmetric scale → Round + clip → int32 workspace → Float reconstruction → Estimated bytes |
| E51 Implementation audit | Evaluate compression at the operation that consumes it | Q, K, V tensors → Scaled dot products → Row softmax → Output → Reconstructed K / V → Error + cosine |
| E52 Evaluation | Nineteen thousand metric rows are not nineteen thousand experiments | Saved metric table → Group identities → Matched comparisons → Prompt-level summaries → Interpretation |
| E53 Implementation prototype | Keep the audio callback small; estimate the motion elsewhere | Speaker carrier → Acoustic path → Audio callback → Windowed estimator → Tracking / filtering |
| E54 Engineering practice | A hundred updates per second is not a hundred independent measurements | 48 kHz samples → Overlapping windows → 10 ms hop → Queue + filters → Displayed trace |
| E55 Engineering practice | Design a sensor data contract before training a sensor model | Wi-Fi observation → Embedded packet → Host parser → Aligned dataset → Future model |
| E56 Implementation audit | A ring buffer is a concurrency protocol, not just an array | CSI callback → Packet framing → PSRAM ring → USB writer task → Host stream |
| E57 Implementation prototype | Carry the filter state across speech frames | Speech frame → Analysis → Excitation generator → Stateful all-pole filter → Energy matching → Reconstructed waveform |
| E58 Engineering practice | Verify a resource limit where the process actually runs | Service definition → Service manager → Container process → Resource controller → Observed counters |
| E59 Implementation | Make a read-only service small by construction | Pinned transport packages → Limited source tree → Python module entry → Read-only transport role → External runtime policy |
| E60 Implementation | What happens when a text classifier reads pixels instead of tokens? | Raw text → Pillow renderer → CNN blocks → Training schedule → Class logits → Text-native control |
| E61 Evaluation | Before improving the network, test the representation | Same document task → Visual branch → Training comparison → Random-feature control → Raw-text branch → Saved accuracy |