Skip to content

docs(328): handoff note for jail pooling — one task, fixed order, one decision to settle - #331

Closed
pdettori wants to merge 1 commit into
rossoctl:mainfrom
pdettori:docs/328-jail-pooling-handoff
Closed

pdettori wants to merge 1 commit into
rossoctl:mainfrom
pdettori:docs/328-jail-pooling-handoff

Conversation

@pdettori

Copy link
Copy Markdown
Member

Summary

Docs only. #328 started as three candidate fixes; two cheap tests reduced it to one, so this
records the decisions instead of leaving the next session to rediscover them.

The finding: the wait for Firecracker's API socket — 72-85% of every restore — is 82-85% the
jailer building a fresh chroot per VM
, not Firecracker, which binds its socket in ~1.5 ms. The
jailer reuses an existing jail directory, and that halves the pre-socket window: 8275 us cold,
then 4025 and 4008 us reused.

What the note pins down

The order of work, which is not arbitrary. Splitting the phase accounting comes first because it
is the measuring instrument for everything after it: the jailer phase times time.Since after
cmd.Start() returns — fork/exec returning — so 6-8 ms of chroot construction is attributed to
sockwait. That mis-attribution already cost two rounds of analysis chasing "Firecracker is slow to
bind"; left in place it would equally obscure whether the fix worked.

The one design decision to settle before writing code. A pooled jail is a directory the previous
VM had write access to — the only place in #328 where a performance change touches isolation. The
awkward part is that the copied firecracker binary inside the jail is most of the saving, so the
residue policy has to permit that file to persist while guaranteeing nothing else does. The note poses
the three questions and recommends an allowlist that fails closed, mirroring how cgroupPool refuses
a dirty cgroup and counts the leak rather than reusing optimistically.

What to measure. The reuse test says nothing about the load-dependent half (sockwait 9.78 ms at
c=8 rising to 22.31 ms at c=64), so report the floor and the slope. Halving the floor while leaving
the slope alone is still a win, but a different one from what a single 64-slot number would suggest.

What not to re-investigate, each with the evidence:

  • --no-api / --config-file — the schema requires drives + boot-source then asks for the kernel
    image; a snapshot key is silently ignored, so it fails open into a fresh boot. Restore is
    API-only in v1.17.0, and we need the API twice anyway (PUT /snapshot/load, then
    PATCH /vm {state: Resumed}, because standbys are deliberately paused).
  • Telling the jailer to skip /dev/net/tun — no such flag; the four mknods are hardcoded.
  • Skipping the jailer entirely — deprioritised, not rejected: pooling gets most of the win without
    reimplementing isolation.

Two figures in #328's own body are wrong and are corrected here because they are the ones likeliest
to be quoted: sockwait at 64 slots is 22-25 ms, not 38.07 (the original was inflated by dropped
caches once at the start plus ~1000 accumulated dying cgroups), and the 238.72/869.43 ms figures were
at 128/256 slots — over-subscription, a different configuration rather than more load.

Rig notes that would otherwise cost an hour each, including the one that cost a full ladder in
this session: per-rung results need RESULTS=<dir> — there is no SH_E11_RESULTS_DIR, an unknown name
is silently ignored, and the worker log is truncated per invocation so only the last rung survives.

Testing

make fmt and make lint pass. No code changes.

Related

Depends on nothing; independent of #300, #301, #327 and #330. Linked from #328.

… decision to settle

rossoctl#328 started as three candidate fixes. Two cheap tests reduced it to one, so this
records the decisions rather than leaving the next session to rediscover them.

What the measurements found: the pre-socket wait is 82-85% the jailer building a
fresh chroot per VM, not Firecracker, which binds its socket in ~1.5 ms. The
jailer reuses an existing jail directory, and that halves the pre-socket window
(8275 us cold, then 4025 and 4008 us).

The note fixes the order of work, because it is not arbitrary:

1. Split the phase accounting first. The `jailer` phase times fork/exec
   RETURNING, so 6-8 ms of chroot construction is attributed to `sockwait`. That
   mis-attribution already cost two rounds of analysis chasing "Firecracker is
   slow to bind"; left in place it would equally obscure whether the fix worked.
2. Implement the pool mirroring cgroup_pool.go, whose rossoctl#319 review comments are
   the specification (verify on pop, never reclaim a name after a failed mint,
   refuse a dirty one and count the leak).
3. Re-measure the floor AND the slope — the reuse test says nothing about the
   load-dependent half, 9.78 ms at c=8 rising to 22.31 ms at c=64.

Also records the design decision to settle before writing code: a pooled jail is
a directory the previous VM had write access to, which is the only place in rossoctl#328
where a performance change touches isolation. The copied firecracker binary is
most of the saving, so the residue policy has to permit that file to persist
while guaranteeing nothing else does.

And what not to re-investigate: `--no-api` (the config schema is fresh-boot only
and silently ignores a snapshot key, so it fails open), telling the jailer to
skip a mknod (no such flag), and two figures in rossoctl#328's own body that were
inflated by host state or taken at a different slot count.

Refs rossoctl#328 rossoctl#319 rossoctl#307 rossoctl#274

Assisted-By: Claude (Anthropic AI) <noreply@anthropic.com>
Signed-off-by: Paolo Dettori <dettori@us.ibm.com>
@pdettori

Copy link
Copy Markdown
Member Author

Closing — #332 refuted this note's central recommendation

Closing rather than merging. This note told the next session that pooling the jail was the
promising direction
and to expect the pre-socket window to fall from ~8 ms toward ~4 ms. #332
implemented it, measured it, and rejected it: −9.24% throughput at 64 slots
(559.5/556.5/552.3 → 522.0/497.1/505.1 Exec/s, non-overlapping reps). Merging this would put
actively misleading guidance into docs/notes/ aimed at a regression.

Where the recommendation came from, and what was wrong with it

The 2× jail-reuse figure in #328 was measuring its own leftover socket. That test deliberately
kept the jail directory between attempts — but the directory contains api.sock, and
waitForUnixSocket dials, so attempt 1's still-live listener satisfied the check instantly. The
signature was in the output and went unread: attempts 2 and 3 agreed to 17 µs, which is a
constant rather than reproducibility. fcJailOccupied already exists to refuse exactly this and its
comment spells it out; it wasn't consulted.

Two further claims of mine do not survive #332's measurements:

  • "The jailer happily reuses an existing jail" — it does not. mknod /dev/net/tun: File exists
    unless root/dev/ is wiped first, so wiping is a correctness requirement of any pool and the four
    mknods are re-paid on every reuse.
  • The baseline was wrong. Restore already MkdirAlls the jail root and hardlinks six files in
    before the jailer starts, so the jail directory always exists in production. Pre-creating it is
    worth −3%, i.e. nothing, and the honest prize was −31% of the pre-socket window, not −50%.

What did hold up

The strace split. #332's independent instrument puts jailer_setup at 93.7% of sockwait with
fc_bind's median at 46 µs, corroborating #328's 82–85% in the stronger direction: the wait is
the jailer, and Firecracker is not the cost. The phase-split prerequisite this note insisted on going
first was also right, and #332 is exactly that instrument.

And the rejection mechanism is worth reading, because it is not one this note anticipated: base
creates a fresh exec-file copy that Destroy unlinks ~30 ms later, so most dirty pages are dropped
before writeback reaches them, while a pool keeps the file and the overwrite cannot be elided.
Keeping that file was the entire saving, so the saving and the cost are the same decision.

Salvaged

Four rig traps from this note are recorded in neither #332's design note nor anywhere on main, so
they move to METAL-RUNBOOK.md in #333: RESULTS= being the only name for a per-rung results dir
(an unknown name silently ignored, worker log truncated per invocation — this cost a full ladder),
sched_schedstats being off so wait_sum probes measure nothing, never timing sudo, and the
leftover-socket signature above.

--no-api / --config-file is covered in #332's note, so it is not duplicated.

Authoritative record for this work: docs/notes/2026-09-22-328-jail-pool-design.md (#332).

@pdettori pdettori closed this Sep 23, 2026
@pdettori
pdettori deleted the docs/328-jail-pooling-handoff branch September 23, 2026 02:58
pdettori added a commit that referenced this pull request Sep 24, 2026
…empty result

Salvaged from the #331 handoff note, whose thesis #332 refuted. These four are
recorded in neither #332's design note nor anywhere else on main, and they share
a shape worth naming: the instrument reports something plausible, so a clean
number is indistinguishable from a broken probe.

- `RESULTS=` is the only name for a per-rung results dir. There is no
  SH_E11_RESULTS_DIR; an unrecognised name is silently ignored, every rung falls
  back to the default .results/, and the worker log is TRUNCATED per invocation.
  An eight-rung ladder can exit 0 everywhere and yield phase data for one rung.
  This cost a full ladder.
- `sched_schedstats` is off on this host, so /proc/<pid>/sched carries no
  se.statistics.* fields. A probe reading wait_sum to detect CPU starvation
  measures nothing and reports nothing, which reads as "never starved".
- Never time `sudo <thing>`: it is setuid doing PAM work and accounted for ~13 ms
  of a 16 ms reading. Also calibrate the harness, because each $(date +%s%N) is a
  fork+exec and puts the floor of a shell-timed measurement at ~2.4 ms — the same
  order as phases worth measuring.
- A leftover socket reads as this run's socket coming up, because
  waitForUnixSocket dials. fcJailOccupied exists to refuse exactly this and its
  comment says so. The signature is repeated attempts agreeing far too closely:
  two runs within 17 us is a constant, not reproducibility. One measurement
  reported a 2x jail-reuse win on that basis; the real figure was -31%, and the
  change it motivated cost 9% throughput.

Cross-references section 3a for the existing kill-by-cgroup-or-port guidance and
adds that pgrep -f self-matches through its own ssh and sudo argv.

Refs #328 #332

Assisted-By: Claude (Anthropic AI) <noreply@anthropic.com>
Signed-off-by: Paolo Dettori <dettori@us.ibm.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant