Sensing & deployment · E56 · Implementation audit

A ring buffer is a concurrency protocol, not just an array

An ESP32 CSI prototype reveals three low-level issues that can quietly invalidate a downstream ML dataset: wraparound, ownership and clock width.

ESP-IDF CSI callbackFreeRTOS taskCustom C ring / USB output
The critical ring-buffer case is a write that crosses the physical end of storage while remaining one logical packet.
Figure 1. The dangerous byte is at the boundary. The critical ring-buffer case is a write that crosses the physical end of storage while remaining one logical packet. Illustrative circular buffer. Original vector illustration.

Follow the information

From input to outcome

The producer and consumer share ring state. A wrap boundary can split a packet physically without splitting it semantically, so length, capacity and concurrency rules must remain consistent.

The producer and consumer share ring state. A wrap boundary can split a packet physically without splitting it semantically, so length, capacity and concurrency rules must remain consistent.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: The critical ring-buffer case is a write that crosses the physical end of storage while remaining one logical packet. The module map and layer-level figures below expand the operations in this route.

A ring buffer is a concurrency protocol, not just an array: architectureCSI callback: Variable-length radio payload → Packet framing: Header + timestamp + RSSI → PSRAM ring: 4 MB intended capacity → USB writer task: Consumes available bytes → Host stream: Needs framing and loss checks. A high-level module map; comparison branches and training details are explained in the article.SENSING & DEPLOYMENT / E56 / MODULE MAP01 INPUTCSI callbackVariable-length radio payload02 MODULEPacket framingHeader + timestamp + RSSI03 MODULEPSRAM ring4 MB intended capacity04 MODULEUSB writer taskConsumes available bytes05 OUTPUTHost streamNeeds framing and loss checks
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
CSI callback — Variable-length radio payload

The architecture in context

The system we are building

The firmware separates a radio callback from a USB-output task using a shared PSRAM buffer. That is a useful producer/consumer architecture: capture should not wait for a slow host. But variable-length packets make the buffer a protocol with invariants, not merely two modulo indices.

Who does what in the stack

ESP-IDF CSI callback
Provides radio measurements.
FreeRTOS task
Drains data independently of the producer.
Custom C ring / USB output
Owns framing, wraparound and synchronization correctness.

The project implements binary framing and a custom ring rather than using a ready-made message queue. The callback writes a packed header and payload; the consumer emits contiguous regions. This makes throughput-oriented design choices visible, while also exposing the correctness work still required.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Nine header bytes do not make a safe ring buffer

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

Producer and consumer coordinate through indices into a circular byte buffer. If a record crosses the physical end, a contiguous memcpy needs splitting or a wrap strategy. Testing only total free capacity does not prove that one contiguous destination span is large enough. The source’s volatile indices also do not by themselves define safe synchronization.

The mathematical contract

9=2magic+4timestamp+1RSSI+2length9=2_{\rm magic}+4_{\rm timestamp}+1_{\rm RSSI}+2_{\rm length}

A larger buffer delays overflow but does not fix straddling writes, timestamp rollover or race conditions. The host parser needs a framing recovery policy after corrupted length fields. Header packing makes transport compact while increasing the importance of exact width and byte-order documentation.

Implementation and resource card

Capacity / budget
ESP32-S3 prototype,4 MB PSRAM buffer; uint 32 microsecond timestamp wraps after about 71.58 minutes.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Use a tiny synthetic buffer to force a header and a payload across the boundary. Test exact-fit, overflow and clock wrap. Keep proposed repairs distinct from the inspected firmware; no live device capture is initiated here.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The excerpt tests only whether the new write position equals the read position, then copies a full packet contiguously. That does not establish sufficient free space, and a packet can straddle the end of the allocation. Modulo-updating write_ptr after the copy does not repair an out-of-bounds write.

C · file · lines 42–58
    size_t packet_size = sizeof(csi_packet_t) + info->len;
    
    // Check for overflow before writing to PSRAM
    if ((write_ptr + packet_size) % CSI_RING_BUFFER_SIZE == read_ptr) {
        // Overflow! Mac isn't reading fast enough.
        return;
    }

    // Wrap the CSI data in our binary frame
    csi_packet_t *pkt = (csi_packet_t *)(ring_buffer + write_ptr);
    pkt->magic = PACKET_HEADER;
    pkt->timestamp = (uint32_t)esp_timer_get_time(); // Microsecond precision for Splines
    pkt->rssi = info->rx_ctrl.rssi;
    pkt->csi_len = info->len;
    memcpy(pkt->payload, info->buf, info->len);

    write_ptr = (write_ptr + packet_size) % CSI_RING_BUFFER_SIZE;

Verbatim archive excerpt from main.c. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

volatile indices do not provide a complete synchronization protocol. The uint32 microsecond timestamp wraps after roughly 71.6 minutes. These are concrete prototype limitations; the article does not provide the firmware as a validated acquisition appliance.

Keep building

Other posts of interest