The architecture in context
The system we are building
The firmware separates a radio callback from a USB-output task using a shared PSRAM buffer. That is a useful producer/consumer architecture: capture should not wait for a slow host. But variable-length packets make the buffer a protocol with invariants, not merely two modulo indices.
Who does what in the stack
- ESP-IDF CSI callback
- Provides radio measurements.
- FreeRTOS task
- Drains data independently of the producer.
- Custom C ring / USB output
- Owns framing, wraparound and synchronization correctness.
The project implements binary framing and a custom ring rather than using a ready-made message queue. The callback writes a packed header and payload; the consumer emits contiguous regions. This makes throughput-oriented design choices visible, while also exposing the correctness work still required.
Open up the implementation
Nine header bytes do not make a safe ring buffer
Producer and consumer coordinate through indices into a circular byte buffer. If a record crosses the physical end, a contiguous memcpy needs splitting or a wrap strategy. Testing only total free capacity does not prove that one contiguous destination span is large enough. The source’s volatile indices also do not by themselves define safe synchronization.
The mathematical contract
A larger buffer delays overflow but does not fix straddling writes, timestamp rollover or race conditions. The host parser needs a framing recovery policy after corrupted length fields. Header packing makes transport compact while increasing the importance of exact width and byte-order documentation.
Implementation and resource card
- Capacity / budget
- ESP32-S3 prototype,4 MB PSRAM buffer; uint 32 microsecond timestamp wraps after about 71.58 minutes.
- Execution evidence
- This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
- Current reproduction context
- Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.
From explanation to a reproducible check
Use a tiny synthetic buffer to force a header and a payload across the boundary. Test exact-fit, overflow and clock wrap. Keep proposed repairs distinct from the inspected firmware; no live device capture is initiated here.
Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.
A closer look at the implementation
The code that carries the idea
The excerpt tests only whether the new write position equals the read position, then copies a full packet contiguously. That does not establish sufficient free space, and a packet can straddle the end of the allocation. Modulo-updating write_ptr after the copy does not repair an out-of-bounds write.
size_t packet_size = sizeof(csi_packet_t) + info->len;
// Check for overflow before writing to PSRAM
if ((write_ptr + packet_size) % CSI_RING_BUFFER_SIZE == read_ptr) {
// Overflow! Mac isn't reading fast enough.
return;
}
// Wrap the CSI data in our binary frame
csi_packet_t *pkt = (csi_packet_t *)(ring_buffer + write_ptr);
pkt->magic = PACKET_HEADER;
pkt->timestamp = (uint32_t)esp_timer_get_time(); // Microsecond precision for Splines
pkt->rssi = info->rx_ctrl.rssi;
pkt->csi_len = info->len;
memcpy(pkt->payload, info->buf, info->len);
write_ptr = (write_ptr + packet_size) % CSI_RING_BUFFER_SIZE;Verbatim archive excerpt from main.c. Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.
The boundary that matters
volatile indices do not provide a complete synchronization protocol. The uint32 microsecond timestamp wraps after roughly 71.6 minutes. These are concrete prototype limitations; the article does not provide the firmware as a validated acquisition appliance.