Sensing & deployment · E58 · Engineering practice

Verify a resource limit where the process actually runs

A container operations lesson generalizes to ML services: configuring a CPU limit and observing that limit are separate tasks.

Service manager / cgroupsContainer imageObserved counters
A resource limit is meaningful only if it applies to the actual workload’s execution scope.
Figure 1. A limit must reach the process. A resource limit is meaningful only if it applies to the actual workload’s execution scope. Resource-control schematic. Original vector illustration.

Follow the information

From input to outcome

A requested limit must reach the cgroup that owns the running workload. Counters are observations of enforcement, not a substitute for checking the process hierarchy.

A requested limit must reach the cgroup that owns the running workload. Counters are observations of enforcement, not a substitute for checking the process hierarchy.
Figure 2. Information flow. Solid arrows carry observations, tensors or artifacts; other routes are explicitly labelled. Signal shapes, matrices and network icons are schematic, not measured samples or literal neuron counts. Open full-size SVG ↗ On narrow screens, scroll the diagram horizontally.

Read this alongside Figure 1: A resource limit is meaningful only if it applies to the actual workload’s execution scope. The module map and layer-level figures below expand the operations in this route.

Verify a resource limit where the process actually runs: system and evaluation mapService definition: Image + resource policy → Service manager: Actual execution context → Container process: Real cgroup membership → Resource controller: CPU / memory enforcement → Observed counters: Throttle and usage evidence. A high-level module map; comparison branches and training details are explained in the article.SENSING & DEPLOYMENT / E58 / MODULE MAP01 INPUTService definitionImage + resource policy02 MODULEService managerActual execution context03 MODULEContainer processReal cgroup membership04 MODULEResource controllerCPU / memory enforcement05 OUTPUTObserved countersThrottle and usage evidence
Source-grounded module map. Boxes summarize operations, not individual neurons; comparison arms and training paths are detailed below. On a small screen, scroll the diagram horizontally.
Service definition — Image + resource policy

The architecture in context

The system we are building

The operational record describes a resource-limit investigation in which a test process was initially launched in the wrong control-group context. That is a useful lesson for model serving and background experimentation: switching a user identity does not necessarily reproduce the service manager’s execution hierarchy.

Who does what in the stack

Service manager / cgroups
Own runtime resource enforcement.
Container image
Defines packaged code and dependencies, not CPU policy.
Observed counters
Verify that the intended process is actually constrained.

The project combined service manifests, container images and measured controller state. Private operational commands, host details and deployment settings are intentionally not reproduced. The short companion excerpt instead shows the limited image boundary; runtime resource enforcement belongs outside that image.

Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph.
Framework responsibility map. Each row maps a library or custom component to its job; rows are not a sequential inference graph. Open full-size SVG ↗

Open up the implementation

Resource limits belong outside the inference loop

A concrete operation-level view of this implementation; no unobserved neural architecture is implied.
A concrete operation-level view of this implementation; no unobserved neural architecture is implied. Open full-size SVG ↗

Application-level thread settings and operating-system resource controls operate at different layers. A declared limit can be attached to the wrong process group or user context, leaving the actual workload unaffected. Verify effective limits from the running process hierarchy rather than trusting a configuration file.

The mathematical contract

effectivecapacity≠requestedconfiguration\mathrm{effective capacity}\neq\mathrm{requested configuration}

A CPU quota is not the same as a thread count, and unified memory is shared with other processes. A fair model benchmark records the effective allocation and competing load. Private operational addresses and service commands are intentionally excluded from the publication.

Implementation and resource card

Capacity / budget
Deployment evidence, not neural performance. The inspected incident involved a mismatched user/service resource scope.
Execution evidence
This revision inspects and explains the archived implementation. It does not rerun the original workload. No unrecorded convergence time, throughput or accelerator result is supplied.
Current reproduction context
Current workstation, supplied by the author: Apple M4, 128 GB unified RAM, 40 GPU cores and 16 CPU cores. This is context for prospective reproduction, not attribution of every archived run. Python and framework versions are not fully locked for these historical sources; declarations, when available, are identified separately.

From explanation to a reproducible check

Use an isolated dummy workload to confirm the intended process scope and effective quota. Record observed utilization without touching live services. Do not infer deployment safety from a successful model forward pass.

Preserve input identities, configuration and failure records with the result. A successful numerical check only establishes the operation it exercises: it does not certify an entire dataset, model or deployed system. Reproduce the interface on a small deterministic input before optimizing throughput or increasing workload size.

A closer look at the implementation

The code that carries the idea

The Containerfile pins two Python package versions and copies a narrow source subtree. Nothing in those lines imposes CPU or memory limits. Those controls must be configured and verified in the actual runtime context, with the observed process identity and controller counters recorded.

Containerfile · file · lines 3–12
FROM docker.io/library/python:3.12-slim

RUN pip install --no-cache-dir pyzmq==26.2.0 msgpack==1.1.0

WORKDIR /app
COPY src/bridge/ /app/src/bridge/
RUN touch /app/src/__init__.py /app/src/bridge/__init__.py 2>/dev/null || true

# Neither entrypoint can place an order: there is no code here that could.
ENTRYPOINT ["python", "-m"]

Verbatim archive excerpt from Containerfile.bridge (companion source E59). Context-dependent historical code, not a standalone runnable program. Comments retain their original wording; the article distinguishes implemented behavior from stale or overbroad comments.

The boundary that matters

A low average CPU reading does not prove enforcement, and a limit configured in one hierarchy may not govern the intended process. Persistent configuration also needs a restart test. This article is a design case study, not an instruction to change this machine’s service manager.

Keep building

Other posts of interest