Sensing & deployment · E57 · Implementation prototype

Carry the filter state across speech frames

A source–filter speech experiment shows why continuity is an engineering property of the whole decoder, not just a good fit inside each frame.

SciPy signalNumPyCustom frame protocol
Filter state carries the preceding frame into the next one; resetting it changes the synthesis even when frame-level parameters match.
Figure 1. Carry state across frames. Filter state carries the preceding frame into the next one; resetting it changes the synthesis even when frame-level parameters match. Illustrative impulse response. Original vector illustration.

Follow the information

From input to outcome

Analysis supplies excitation parameters and filter poles. The filter state carries across frame boundaries; resetting it for each frame changes the reconstructed waveform.

Analysis supplies excitation parameters and filter poles. The filter state carries across frame boundaries; resetting it for each frame changes the reconstructed waveform.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: Filter state carries the preceding frame into the next one; resetting it changes the synthesis even when frame-level parameters match. The module map and layer-level figures below expand the operations in this route.

Carry the filter state across speech frames: architectureSpeech frame: 16 kHz · 25 ms → Analysis: Voicing / pitch / poles → Excitation generator: Pulses + noise → Stateful all-pole filter: Carry zi → zf → Energy matching: Frame-level scaling → Reconstructed waveform: Needs codec validation. A high-level module map; comparison branches and training details are explained in the article.SENSING & DEPLOYMENT / E57 / MODULE MAP01 INPUTSpeech frame16 kHz · 25 ms02 MODULEAnalysisVoicing / pitch / poles03 MODULEExcitation generatorPulses + noise04 MODULEStateful all-pole filterCarry zi → zf05 MODULEEnergy matchingFrame-level scaling06 OUTPUTReconstructed waveformNeeds codec validation
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Speech frame — 16 kHz · 25 ms

The architecture in context

The system we are building

The speech prototype analyzes short frames, estimates a stable filter and synthesizes speech-like output from voiced or unvoiced excitation. The system is model-based signal processing rather than a neural vocoder. It is still relevant ML engineering: a learned parameter predictor would have to respect the same state and waveform interfaces.

Who does what in the stack

SciPy signal
Implements analysis filters and stateful synthesis.
NumPy
Computes recurrence fits, polynomial roots and stabilization.
Custom frame protocol
Carries excitation phase, filter state and energy across frames.

The custom analysis regularizes a recurrence fit, adjusts pole radii and selects excitation behavior. SciPy implements filtering; NumPy implements the frame-level algebra. The synthesizer carries both filter state and pulse phase between frames to reduce discontinuities.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Carry synthesis state across frame boundaries

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

The vocal-tract approximation is an all-pole filter driven by voiced/noise excitation. Carrying lfilter’s final state into the next call avoids resetting the filter at every frame boundary. Pole stabilization and excitation phase continuity solve different problems; both can influence clicks and synthetic quality.

The mathematical contract

A(z)=1+∑k=118akz−k,y=filter⁡(1,A,e)A(z)=1+\sum_{k=1}^{18}a_kz^{-k},\qquad y=\operatorname{filter}(1,A,e)

Clipping pole radii below .99 prevents unstable synthesis but alters the spectral envelope. Quantization, codebook metadata and excitation parameters would all belong in a real bitstream budget. A reconstruction plot does not establish a bitrate until that format is specified.

Implementation and resource card

Capacity / budget
16 kHz audio;400 samples per 25 ms frame; order 18; codebook setting 64. No neural training budget; a complete serialized codec is not defined.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Split one synthetic excitation into frames and compare stateful filtering with a continuous call under fixed coefficients. Then deliberately reset state to expose discontinuities. Changing coefficients between frames needs its own state-transition and stability check.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt supplies zi to lfilter and returns zf for the next frame. Resetting that state every 25 ms would repeatedly erase the filter’s memory. Output energy is then matched to the source frame, with a separate unvoiced adjustment; that scaling also changes how continuity should be evaluated.

Python · file · lines 115–126
    if len(zi) != len(poles) - 1:
        zi = np.zeros(len(poles) - 1)
    
    synth, zf = signal.lfilter([1.0], poles, source, zi=zi)
    
    synth_energy = np.mean(synth**2) + 1e-9
    scaling_factor = np.sqrt(energy / synth_energy)
    if not is_voiced:
        scaling_factor *= 0.5 
        
    synth = synth * scaling_factor
    return synth, zf, phase_offset

Verbatim archive excerpt from auto_benchmark.py. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

The file is an analysis/synthesis benchmark, not a complete compressed bitstream specification. Pole coefficients, pitch, voicing, energy, codebooks and framing all consume bits. A reconstructed WAV and a favorable spectral metric do not establish a deployed bitrate or perceptual quality.

Keep building

Other posts of interest