Model architecture · E05 · Implementation audit

A padded sequence needs more than zero-valued inputs

A three-layer LSTM classifier illustrates the difference between masking a measurement and masking a recurrent state transition.

Keras functional APIKeras LSTMDropout / Dense
Zeroing a padded input does not zero the recurrent computation: hidden state, cell state, weights and biases remain active.
Figure 1. Zero input ≠ frozen state. Zeroing a padded input does not zero the recurrent computation: hidden state, cell state, weights and biases remain active. Recurrent-state schematic. Original vector illustration.

Follow the information

From input to outcome

The validity mask multiplies input values. It is not a recurrent-state freeze: biases, hidden state and cell state still participate in each LSTM transition.

The validity mask multiplies input values. It is not a recurrent-state freeze: biases, hidden state and cell state still participate in each LSTM transition.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Zeroing a padded input does not zero the recurrent computation: hidden state, cell state, weights and biases remain active. The module map and layer-level figures below expand the operations in this route.

A padded sequence needs more than zero-valued inputs: architectureSequence + validity: B × T × 1 → Multiply by mask: Zero padded values → LSTM stack: 128 → 256 → 512 → Dense representation: 1,024 units → Classifier: Three-way softmax. A high-level module map; comparison branches and training details are explained in the article.MODEL ARCHITECTURE / E05 / MODULE MAP01 INPUTSequence + validityB × T × 102 MODULEMultiply by maskZero padded values03 MODULELSTM stack128 → 256 → 51204 MODULEDense representation1,024 units05 OUTPUTClassifierThree-way softmax
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Sequence + validity — B × T × 1

The architecture in context

The system we are building

A recurrent classifier has a state as well as an input. The notebook accepts a sequence and a same-shaped numeric mask, multiplies them, and feeds the result through LSTMs of increasing width. Dropout and a large dense layer sit between the learned temporal representation and the three-class output. This is a useful example of composing multiple inputs in the Keras functional API.

Who does what in the stack

Keras functional API
Connects sequence and validity inputs.
Keras LSTM
Maintains recurrent hidden and cell states.
Dropout / Dense
Regularizes and classifies the final representation.

The custom input branch makes missingness explicit at the model boundary. It does not, however, pass a boolean temporal mask into the recurrent layers. That distinction matters whenever examples are padded to a common length: an LSTM can update its state on a zero input because gates, recurrent weights and biases remain active.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

From module map to executable structure

Inside Three-layer recurrent classifier

Single-channel sequence; three classes; hidden widths 128, 256 and 512.

Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table.
Layer-level implementation. B denotes batch size; parameter and shape conventions are expanded in the table. Open full-size SVG ↗
Layer / tensor / operation ledger
Layer or branchOutput shapeImplementation detail
Input × numeric maskB × T × 1Multiplication zeros padded values; it does not stop recurrent state transitions.
LSTM128B × T × 128Return all states → LeakyReLU(.01) → dropout .2.
LSTM256B × T × 256Return all states → LeakyReLU(.01) → dropout .3.
LSTM512B × 512Return final state → LeakyReLU(.01) → dropout .5.
Dense classifierB × 3512→1024 → LeakyReLU → dropout .5 → Dense 3 softmax.

Each LSTM has four affine transformations of the input and previous hidden state: input, forget and output gates plus a candidate state. Its parameter count is 4h(d+h+1). Zero input still leaves recurrent weights and biases active. Multiplying a mask into the input is therefore different from telling Keras to skip a padded time step.

One LSTM cell opened into its gates. This cell is repeated over time and within each recurrent layer.
One LSTM cell opened into its gates. This cell is repeated over time and within each recurrent layer. Open full-size SVG ↗

The equation and the update

ct=ft⊙ct−1+it⊙c~t,ht=ot⊙tanh⁡(ct)c_t=f_t\odot c_{t-1}+i_t\odot\widetilde c_t,\qquad h_t=o_t\odot\tanh(c_t)

Training cell specifies batch 128, maximum 100 epochs and learning rate 1e-4. Preserve the categorical target convention when rebuilding the classifier; recurrent padding semantics must be fixed before interpreting a training comparison.

Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries.
Learning or solution path. A parameter-update path is different from the forward inference path; see text for target-network, frozen-feature and local-loss boundaries. Open full-size SVG ↗

Implementation card / no invented benchmarks

Capacity, budget and execution evidence

Parameters / retained state
2,564,099 scalars for Keras four-gate LSTMs with one bias vector per gate.
Duration and hardware evidence
No verified convergence duration or historical hardware identity attached.
Source coordinates
E05 cell 15, lines 98–137; cell 18
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

Counts above are calculated from the stated layer shapes unless identified as saved measurements. They exclude optimizer state and nontrainable buffers. No archived training was rerun for this revision.

What these design choices change

Increasing hidden width costs quadratically in recurrent parameters and also enlarges every stored activation across time. The final 1024-unit head contributes 528,387 parameters including its three-class output. The archive does not establish that these widths are optimal; narrower recurrent layers and explicit masking are distinct experiments.

Reproduction and measurement protocol

Compare a sequence alone with the same sequence followed by zero padding. If predictions change, that is expected for this implementation and is evidence against treating the numeric mask as sequence-length masking. Keep sequence orientation, padding side and label alignment in the batch specification. Test gradients on short sequences before increasing length.

For a new run, save the resolved Python/framework versions, backend, dtype, seed, input shapes, batch size and exact source revision. Start with one batch and one update. Log training steps separately from epochs or environment steps. Do not equate the configured maximum with a completed budget or convergence.

Measure initialization/compilation, data preparation, warmed forward pass, training updates and evaluation separately. Synchronize accelerator work around timed regions using the chosen framework’s supported mechanism. Report peak process memory and framework allocation separately; parameter bytes exclude activations, gradients, optimizer state and input buffers. On a shared machine, begin with a single CPU worker and a small batch rather than claiming all available resources.

A closer look at the implementation

The code that carries the idea

The excerpt shows precisely where validity enters: Multiply changes sequence values, and the resulting tensor is passed to LSTM. There is no mask argument in that call. Orthogonal recurrent initialization and He input initialization are separate choices; neither changes padding semantics.

Python · cell 15 · lines 98–111
def create_lstm_model_masked(sequence_length):
    # Inputs for sequences and masks
    sequence_input = Input(shape=(sequence_length, 1), name="sequence_input")
    mask_input = Input(shape=(sequence_length, 1), name="mask_input")  # Mask input
    
    # Apply the mask by multiplying the sequence with it
    masked_sequence = Multiply()([sequence_input, mask_input])
    
    # First LSTM layer
    x = LSTM(128, return_sequences=True,
             kernel_initializer='he_normal',
             recurrent_initializer='orthogonal')(masked_sequence)
    x = LeakyReLU(alpha=0.01)(x)
    x = Dropout(0.2)(x)

Verbatim archive excerpt from lstm_model_lp2_3class.ipynb. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

This implementation should not be described as padding-invariant solely because its function name contains “masked”. A real zero observation and an absent timestep may have different meanings. Changing to true masking also needs a decision about left padding, right padding and interior gaps.

Keep building

Other posts of interest