The architecture in context
The system we are building
A recurrent classifier has a state as well as an input. The notebook accepts a sequence and a same-shaped numeric mask, multiplies them, and feeds the result through LSTMs of increasing width. Dropout and a large dense layer sit between the learned temporal representation and the three-class output. This is a useful example of composing multiple inputs in the Keras functional API.
Who does what in the stack
- Keras functional API
- Connects sequence and validity inputs.
- Keras LSTM
- Maintains recurrent hidden and cell states.
- Dropout / Dense
- Regularizes and classifies the final representation.
The custom input branch makes missingness explicit at the model boundary. It does not, however, pass a boolean temporal mask into the recurrent layers. That distinction matters whenever examples are padded to a common length: an LSTM can update its state on a zero input because gates, recurrent weights and biases remain active.
From module map to executable structure
Inside Three-layer recurrent classifier
Single-channel sequence; three classes; hidden widths 128, 256 and 512.
| Layer or branch | Output shape | Implementation detail |
|---|---|---|
| Input × numeric mask | B × T × 1 | Multiplication zeros padded values; it does not stop recurrent state transitions. |
| LSTM128 | B × T × 128 | Return all states → LeakyReLU(.01) → dropout .2. |
| LSTM256 | B × T × 256 | Return all states → LeakyReLU(.01) → dropout .3. |
| LSTM512 | B × 512 | Return final state → LeakyReLU(.01) → dropout .5. |
| Dense classifier | B × 3 | 512→1024 → LeakyReLU → dropout .5 → Dense 3 softmax. |
Each LSTM has four affine transformations of the input and previous hidden state: input, forget and output gates plus a candidate state. Its parameter count is 4h(d+h+1). Zero input still leaves recurrent weights and biases active. Multiplying a mask into the input is therefore different from telling Keras to skip a padded time step.
The equation and the update
Training cell specifies batch 128, maximum 100 epochs and learning rate 1e-4. Preserve the categorical target convention when rebuilding the classifier; recurrent padding semantics must be fixed before interpreting a training comparison.
Implementation card / no invented benchmarks
Capacity, budget and execution evidence
- Parameters / retained state
- 2,564,099 scalars for Keras four-gate LSTMs with one bias vector per gate.
- Duration and hardware evidence
- No verified convergence duration or historical hardware identity attached.
- Source coordinates
- E05 cell 15, lines 98–137; cell 18
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.
What these design choices change
Increasing hidden width costs quadratically in recurrent parameters and also enlarges every stored activation across time. The final 1024-unit head contributes 528,387 parameters including its three-class output. The archive does not establish that these widths are optimal; narrower recurrent layers and explicit masking are distinct experiments.
Reproduction and measurement protocol
Compare a sequence alone with the same sequence followed by zero padding. If predictions change, that is expected for this implementation and is evidence against treating the numeric mask as sequence-length masking. Keep sequence orientation, padding side and label alignment in the batch specification. Test gradients on short sequences before increasing length.
For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.
Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.
A closer look at the implementation
The code that carries the idea
The excerpt shows precisely where validity enters: Multiply changes sequence values, and the resulting tensor is passed to LSTM. There is no mask argument in that call. Orthogonal recurrent initialization and He input initialization are separate choices; neither changes padding semantics.
def create_lstm_model_masked(sequence_length):
# Inputs for sequences and masks
sequence_input = Input(shape=(sequence_length, 1), name="sequence_input")
mask_input = Input(shape=(sequence_length, 1), name="mask_input") # Mask input
# Apply the mask by multiplying the sequence with it
masked_sequence = Multiply()([sequence_input, mask_input])
# First LSTM layer
x = LSTM(128, return_sequences=True,
kernel_initializer='he_normal',
recurrent_initializer='orthogonal')(masked_sequence)
x = LeakyReLU(alpha=0.01)(x)
x = Dropout(0.2)(x)Verbatim archive excerpt from lstm_model_lp2_3class.ipynb. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
This implementation should not be described as padding-invariant solely because its function name contains “masked”. A real zero observation and an absent timestep may have different meanings. Changing to true masking also needs a decision about left padding, right padding and interior gaps.