Repository navigation
Conversation
… decision to settle rossoctl#328 started as three candidate fixes. Two cheap tests reduced it to one, so this records the decisions rather than leaving the next session to rediscover them. What the measurements found: the pre-socket wait is 82-85% the jailer building a fresh chroot per VM, not Firecracker, which binds its socket in ~1.5 ms. The jailer reuses an existing jail directory, and that halves the pre-socket window (8275 us cold, then 4025 and 4008 us). The note fixes the order of work, because it is not arbitrary: 1. Split the phase accounting first. The `jailer` phase times fork/exec RETURNING, so 6-8 ms of chroot construction is attributed to `sockwait`. That mis-attribution already cost two rounds of analysis chasing "Firecracker is slow to bind"; left in place it would equally obscure whether the fix worked. 2. Implement the pool mirroring cgroup_pool.go, whose rossoctl#319 review comments are the specification (verify on pop, never reclaim a name after a failed mint, refuse a dirty one and count the leak). 3. Re-measure the floor AND the slope — the reuse test says nothing about the load-dependent half, 9.78 ms at c=8 rising to 22.31 ms at c=64. Also records the design decision to settle before writing code: a pooled jail is a directory the previous VM had write access to, which is the only place in rossoctl#328 where a performance change touches isolation. The copied firecracker binary is most of the saving, so the residue policy has to permit that file to persist while guaranteeing nothing else does. And what not to re-investigate: `--no-api` (the config schema is fresh-boot only and silently ignores a snapshot key, so it fails open), telling the jailer to skip a mknod (no such flag), and two figures in rossoctl#328's own body that were inflated by host state or taken at a different slot count. Refs rossoctl#328 rossoctl#319 rossoctl#307 rossoctl#274 Assisted-By: Claude (Anthropic AI) <noreply@anthropic.com> Signed-off-by: Paolo Dettori <dettori@us.ibm.com>
Closing — #332 refuted this note's central recommendationClosing rather than merging. This note told the next session that pooling the jail was the Where the recommendation came from, and what was wrong with itThe 2× jail-reuse figure in #328 was measuring its own leftover socket. That test deliberately Two further claims of mine do not survive #332's measurements:
What did hold upThe strace split. #332's independent instrument puts And the rejection mechanism is worth reading, because it is not one this note anticipated: base SalvagedFour rig traps from this note are recorded in neither #332's design note nor anywhere on
Authoritative record for this work: |
…empty result Salvaged from the #331 handoff note, whose thesis #332 refuted. These four are recorded in neither #332's design note nor anywhere else on main, and they share a shape worth naming: the instrument reports something plausible, so a clean number is indistinguishable from a broken probe. - `RESULTS=` is the only name for a per-rung results dir. There is no SH_E11_RESULTS_DIR; an unrecognised name is silently ignored, every rung falls back to the default .results/, and the worker log is TRUNCATED per invocation. An eight-rung ladder can exit 0 everywhere and yield phase data for one rung. This cost a full ladder. - `sched_schedstats` is off on this host, so /proc/<pid>/sched carries no se.statistics.* fields. A probe reading wait_sum to detect CPU starvation measures nothing and reports nothing, which reads as "never starved". - Never time `sudo <thing>`: it is setuid doing PAM work and accounted for ~13 ms of a 16 ms reading. Also calibrate the harness, because each $(date +%s%N) is a fork+exec and puts the floor of a shell-timed measurement at ~2.4 ms — the same order as phases worth measuring. - A leftover socket reads as this run's socket coming up, because waitForUnixSocket dials. fcJailOccupied exists to refuse exactly this and its comment says so. The signature is repeated attempts agreeing far too closely: two runs within 17 us is a constant, not reproducibility. One measurement reported a 2x jail-reuse win on that basis; the real figure was -31%, and the change it motivated cost 9% throughput. Cross-references section 3a for the existing kill-by-cgroup-or-port guidance and adds that pgrep -f self-matches through its own ssh and sudo argv. Refs #328 #332 Assisted-By: Claude (Anthropic AI) <noreply@anthropic.com> Signed-off-by: Paolo Dettori <dettori@us.ibm.com>
Summary
Docs only. #328 started as three candidate fixes; two cheap tests reduced it to one, so this
records the decisions instead of leaving the next session to rediscover them.
The finding: the wait for Firecracker's API socket — 72-85% of every restore — is 82-85% the
jailer building a fresh chroot per VM, not Firecracker, which binds its socket in ~1.5 ms. The
jailer reuses an existing jail directory, and that halves the pre-socket window: 8275 us cold,
then 4025 and 4008 us reused.
What the note pins down
The order of work, which is not arbitrary. Splitting the phase accounting comes first because it
is the measuring instrument for everything after it: the
jailerphase timestime.Sinceaftercmd.Start()returns — fork/exec returning — so 6-8 ms of chroot construction is attributed tosockwait. That mis-attribution already cost two rounds of analysis chasing "Firecracker is slow tobind"; left in place it would equally obscure whether the fix worked.
The one design decision to settle before writing code. A pooled jail is a directory the previous
VM had write access to — the only place in #328 where a performance change touches isolation. The
awkward part is that the copied
firecrackerbinary inside the jail is most of the saving, so theresidue policy has to permit that file to persist while guaranteeing nothing else does. The note poses
the three questions and recommends an allowlist that fails closed, mirroring how
cgroupPoolrefusesa dirty cgroup and counts the leak rather than reusing optimistically.
What to measure. The reuse test says nothing about the load-dependent half (
sockwait9.78 ms atc=8 rising to 22.31 ms at c=64), so report the floor and the slope. Halving the floor while leaving
the slope alone is still a win, but a different one from what a single 64-slot number would suggest.
What not to re-investigate, each with the evidence:
--no-api/--config-file— the schema requiresdrives+boot-sourcethen asks for the kernelimage; a
snapshotkey is silently ignored, so it fails open into a fresh boot. Restore isAPI-only in v1.17.0, and we need the API twice anyway (
PUT /snapshot/load, thenPATCH /vm {state: Resumed}, because standbys are deliberately paused)./dev/net/tun— no such flag; the fourmknods are hardcoded.reimplementing isolation.
Two figures in #328's own body are wrong and are corrected here because they are the ones likeliest
to be quoted:
sockwaitat 64 slots is 22-25 ms, not 38.07 (the original was inflated by droppedcaches once at the start plus ~1000 accumulated dying cgroups), and the 238.72/869.43 ms figures were
at 128/256 slots — over-subscription, a different configuration rather than more load.
Rig notes that would otherwise cost an hour each, including the one that cost a full ladder in
this session: per-rung results need
RESULTS=<dir>— there is noSH_E11_RESULTS_DIR, an unknown nameis silently ignored, and the worker log is truncated per invocation so only the last rung survives.
Testing
make fmtandmake lintpass. No code changes.Related
Depends on nothing; independent of #300, #301, #327 and #330. Linked from #328.