Build a patch transformer—and make every pixel count
A TensorFlow/Keras classifier turns a narrow chart image into tokens. Padding, residual paths and head dimensions determine the model you actually built.
ImplementationA new branch of Research & Algorithms
The architecture, the tooling, the custom code—and the detail that makes it work.
Hands-on engineering from my project archive. From patch transformers and evolutionary policies to streaming sensors and compact model memory.
What reaches the model?
Where does information flow?
Who owns the update?
What did the implementation establish?
61 source-grounded case studies
Original code excerpts · Inspectable module maps
Start here
A TensorFlow/Keras classifier turns a narrow chart image into tokens. Padding, residual paths and head dimensions determine the model you actually built.
ImplementationA CNN DQN explicitly propagates vector–Jacobian products across detached blocks. This is custom gradient orchestration—not an autograd replacement.
ImplementationA pretrained representation, random Fourier features and a ridge solve form a simple continual readout. Its guarantees stop at the fixed feature map.
ImplementationAn acoustic prototype combines a high-frequency carrier, queued microphone blocks and a three-sample sinusoidal identity. It is not a validated breathing monitor.
Implementation prototypeJAX handles the population axis, Flax defines the policy and EvoJAX updates a search distribution. The interfaces—not just the optimizer—make the experiment work.
ImplementationA tiny quantizer makes the distinction between simulated distortion, resident array storage and a real packed payload impossible to miss.
Implementation auditConnected ideas, practical routes
From image patches to relational attention—and the tensor contracts between them.
Population axes, replay buffers, gradient routing and reproducible checkpoints.
Contrastive learning, distillation and fixed-feature memory, with the controls that changed the interpretation.
Causal data, bounded memory, live sensing and deliberately small services.
The complete engineering library
61 articles
A TensorFlow/Keras classifier turns a narrow chart image into tokens. Padding, residual paths and head dimensions determine the model you actually built.
ImplementationThe PyTorch version makes the training loop explicit—and exposes three architectural differences that a line-by-line port can hide.
ImplementationA shared CNN, identity embeddings and an explicit pair constructor turn a set of images into relation tokens. The interesting engineering decision is where the quadratic expansion happens.
ImplementationSeparate temporal encoders feed a custom graph-attention layer. A fourteen-line excerpt shows why sparse topology must survive the softmax.
Implementation auditA three-layer LSTM classifier illustrates the difference between masking a measurement and masking a recurrent state transition.
Implementation auditA probability-to-decision layer separates model confidence from action. Its denominator is as important as its threshold.
EvaluationScaling and PCA are learned components too. This archive example gets the fit/transform boundary right while showing why a date filter is not automatically a forward test.
ImplementationA compact model-selection loop makes an important boundary visible: the data that chooses the winner cannot also provide its untouched final score.
Evaluation auditBefore a CNN sees a chart, plotting code has chosen the context, coordinate system and information budget. Treat that renderer as part of the model.
ImplementationJAX handles the population axis, Flax defines the policy and EvoJAX updates a search distribution. The interfaces—not just the optimizer—make the experiment work.
ImplementationA multi-resolution feature builder combines fast detail and slower context. Its aggregation boundary reveals a subtle source of look-ahead.
Implementation auditA NEAT training notebook shows why a genome checkpoint is only one part of a reproducible policy.
ImplementationLoading a trained genome is straightforward. Establishing that its evaluation is genuinely out of distribution takes a separate evidence trail.
EvaluationA recurrent DQN prototype adds priority sampling and importance weights. The replay buffer becomes part of the learning algorithm, not just storage.
ImplementationEnvironment steps, replay updates and target synchronization are different events. A small PyTorch agent makes their roles visible.
ImplementationTensorNEAT is third-party software, not an original model from this archive. Its package metadata provides a useful map of the stack a local adaptation must respect.
Upstream referenceA PyTorch adaptation stores nodes and connections in bounded arrays. The central challenge is preserving graph semantics while making population evaluation regular.
Implementation prototypeA nine-line flatten/unflatten adapter is the hinge between a structured neural network and a population optimizer.
ImplementationAn episode dataset has two time directions: observations describe what was known; the future tape measures what happened next.
ImplementationA compiled simulation kernel can accelerate repeated evaluations. It cannot turn an in-sample grid search into walk-forward validation.
ImplementationSession-level accounting, zero-action days and regime labels belong in the result—not just the winning events.
EvaluationA small resampling notebook illustrates the difference between Monte Carlo precision and the amount of independent evidence.
Evaluation auditOptuna workers can share a study, but they must also share a frozen objective, data identity and resource budget.
ImplementationA scheduled market-data recorder separates callbacks, buffers and file writes. Each boundary has a different failure mode.
ImplementationA timestamp-bucket aggregator closes a bar when the next bucket arrives—not merely when a clock crosses its nominal boundary.
ImplementationTyped deduplication keys prevent repeated messages from becoming repeated state changes. A bounded cache is useful—but it is not exactly-once delivery.
ImplementationA PyTorch contrastive experiment asks which changes an embedding should ignore before fitting a small supervised readout.
ImplementationThe support distribution and test transformations matter as much as the encoder. A saved few-shot study makes that distinction concrete.
EvaluationGreedy local contrastive learning freezes earlier blocks and gives the current block its own projection head. That is different from splitting an ordinary backward pass.
ImplementationA saved CIFAR-10 comparison turns a disappointing result into a useful engineering decision: isolate the representation before expanding the training recipe.
Negative resultA compact CNN learns against teacher features with an optional classification loss. The feature interface is the real contract between teacher and student.
ImplementationParameter accounting and capability retention are separate tests. A two-run distillation record makes both visible.
Negative resultA pretrained representation, random Fourier features and a ridge solve form a simple continual readout. Its guarantees stop at the fixed feature map.
ImplementationThe saved DINOv2-feature experiment supports a precise result: additive statistics recover the same fixed-feature fit. It does not freeze every earlier prediction.
EvaluationA CNN DQN explicitly propagates vector–Jacobian products across detached blocks. This is custom gradient orchestration—not an autograd replacement.
ImplementationA Pong smoke test completes on MPS but does not learn a useful policy. Its action distribution is more informative than its runtime headline.
Negative resultA SAC experiment separates the controller, actuator law and simulated world. That modularity makes model transfer testable.
ImplementationWeights, validation traces, resource measurements and an immutable-world check turn a checkpoint into something another engineer can inspect.
ImplementationA multi-arm locomotion gate tests whether an improved model enables a useful policy—not just whether a fitted curve looks better.
Negative resultDelay embeddings and a reduced eigensystem provide one route from observed dynamics to a compact recurrent feature bank.
ImplementationNoise robustness and length generalization are different questions. A saved state-space panel keeps both visible.
EvaluationA SIREN baseline maps coordinates to a field. Its frequency scale and weight initialization are part of the architecture—not incidental optimizer settings.
ImplementationA classical PINN baseline exposes the cost of high-order residuals and the distinction between an analytical derivative and an automatically differentiated network.
ImplementationLocal trigonometric atoms and global null-space functions can share one feature interface. Their derivatives do not share one universal shortcut.
Implementation auditA project-state ledger connects code, evidence and closed decisions so an old experimental queue does not silently become today’s plan.
Engineering practiceRetaining sufficient statistics is not the same as preserving old predictions. An eleven-line counterexample keeps that distinction executable.
Engineering practiceA bounded streaming capsule keeps a small state, but its numerical assurance can fail before memory runs out.
EvaluationA KV-cache experiment needs more than hidden-state tensors. Queries, keys and values must come from the same actual attention computation.
Implementation auditReconstruction error, attention error and memory accounting test different promises. A benchmark should retain all three.
Engineering practiceA tiny quantizer makes the distinction between simulated distortion, resident array storage and a real packed payload impossible to miss.
Implementation auditAttention-output error is more task-connected than tensor error—but only if the proxy matches the attention semantics you intend to preserve.
Implementation auditA head-by-head compression table needs hierarchy-aware aggregation before it can support a model-level claim.
EvaluationAn acoustic prototype combines a high-frequency carrier, queued microphone blocks and a three-sample sinusoidal identity. It is not a validated breathing monitor.
Implementation prototypeA sensing report becomes more useful when it separates sampling rate, window length, hop size, queue delay and physiological validation.
Engineering practiceThe hardware roadmap is a plan, not a completed multimodal experiment. Its concrete starting point is a timestamped observation packet.
Engineering practiceAn ESP32 CSI prototype reveals three low-level issues that can quietly invalidate a downstream ML dataset: wraparound, ownership and clock width.
Implementation auditA source–filter speech experiment shows why continuity is an engineering property of the whole decoder, not just a good fit inside each frame.
Implementation prototypeA container operations lesson generalizes to ML services: configuring a CPU limit and observing that limit are separate tasks.
Engineering practiceA minimal Python image copies only the bridge implementation. Dependency boundaries can reduce operational capability as well as image size.
ImplementationA language-as-image experiment turns rendering into a lossy front end and compares two ways of training the resulting CNN.
ImplementationA saved text-as-image comparison shows close visual training results—but the straightforward raw-text baseline is stronger.
EvaluationNo matching articles. Try a framework name or clear the filters.