experiment: herdr integration for cloud-agents - #1176
Conversation
Adds `railway ca herdr {install,new,agents,sync,bootstrap}` so cloud
agents show in herdr 0.9 as saved SSH machines. `install` writes and
links a herdr plugin manifest whose every command calls back into these
verbs; `bootstrap` provisions the VM in app mode and prepares its herdr
server.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
…n recovery `install` writes and updates the keybindings itself. Bootstrap links a plugin on the VM too, so prefix+shift+a and prefix+shift+s (sleep this agent) work while a machine is selected, and starts the chosen harness in the /app pane. Local prefix+shift+s wakes. `sync` runs on workspace focus (debounced), remembers agent status, waits for the ssh relay before enabling a machine, and reconnects machines herdr parked in Attention by reading the client log. Picker closes after connect/new/wake, notifies before sleep and wake, caches project names. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
`railway ca herdr watch` subscribes to backboard's cloudAgentInvalidation per environment (the unpublished /graphql/internal graph the apps use) and reconciles machines on each event, so a sleep from anywhere disables the machine before herdr's reconnect can park it in Attention and a wake re-enables it once the relay answers. One watcher per herdr session, started by install and the startup hook, nudged by SIGUSR1, no polling. Also: prefix+shift+y runs sync with a toast; the sweep hook moves to pane.agent_status_changed (focus events are client-local in herdr 0.9 and never reach hooks); the client-log reading is removed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
Watcher pidfile only counts a live `railway ca herdr watch` process, so a reused pid is never signalled or trusted; nudge happens after the debounce; the detached watcher logs quietly with rotation and exits when the session socket stops answering. `sync` claims its window before the relay wait, removes only machines it recorded itself and never on an empty agent list, keeps the old status when an Enable/Kick fails so it is retried, and reports failures in its summary and toast. `new` waits for the relay before `herdr machine add` like every other path. Bootstrap ends in BOOTSTRAP-FAILED when a step fails, accepts `[table] # comment` headers, and feeds the token to curl on stdin. `install --remove` works with herdr stopped and stops the watcher; known_hosts is appended, never rewritten. herdr verbs are no longer tracked per invocation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
…ilway Bindings shrink to prefix+shift+a (the picker, which now leads with "new agent" and "sync now") and prefix+shift+s (wake). Bootstrap upgrades the VM's railway in place when it is older than the one running bootstrap, so the remote picker works as soon as a release carries `ca herdr`. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
herdr reconnects with strict host-key checking and no config of its own. A Host block that sends ssh.railway.com to UserKnownHostsFile /dev/null parks every Railway machine in Attention after the first network blip, which herdr never retries. `install` now runs `ssh -G` on the relay host and says so, naming the fix. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
Socket liveness, the SIGUSR1 nudge and the ps/kill pid checks are unix only; elsewhere they compile to no-ops. The plugin declares linux and macos, so on Windows the binary builds and the herdr verbs simply do nothing extra. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
The fake is a shebang script, which Windows cannot execute. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
|
Opened #1187 targeting this branch with the pressure-tested reliability fixes: account/backend-scoped and atomic state updates, stale-plan checks around relay waits, server-authoritative wake/sleep requests, reconnect refetches plus periodic reconciliation, SSH identity handoff, and empty-picker/bootstrap retry fixes. All 1,418 tests pass with injected Railway token overrides cleared from the test process; formatting and Clippy checks pass (existing lint warnings). The contribution includes the reproductions and the remaining integration limits: Herdr 0.9 already retries ordinary network failures, but its machine-list API cannot identify live Attention state, and live per-harness create/sleep/wake restoration still needs end-to-end validation. |
* fix(run): cover execution-search variables in the host env filter (#1175) The local-spawn filter listed shell hooks, runtime hooks and the loader prefixes, but not the names that relocate where code is searched for. PATH picks which binary a bare command name resolves to and HOME picks which startup file an interactive shell sources, so both choose what the child runs without naming a command themselves. Add that category, plus the GIT_CONFIG prefix (GIT_CONFIG_KEY_<n> sets core.pager and core.sshCommand with no file involved, and the <n> makes the family open-ended) and the JVM, gem, lua and R startup hooks that were missing next to the Python and Perl ones. Dropping is still the right primitive: the spawn merges what survives onto the inherited environment, so a dropped name leaves the machine's own value in place for the child. * chore: Release railwayapp version 5.49.6 * feat(ca): add OpenCode desktop and local client connections (#1173) * feat(ca): add OpenCode desktop support * feat(ca): connect OpenCode Desktop through agent HTTPS * feat(ca): configure OpenCode Desktop automatically * fix(ca): protect saved OpenCode Desktop credentials * feat(ca): configure OpenCode Beta desktop stores * feat(ca): isolate OpenCode editions and launch the latest OpenCode2 Beta * feat(code): add authenticated OpenCode remote servers * feat(code): launch local OpenCode clients and discover remote servers * fix(opencode): highlight shared connection details * fix(ca): handle OpenCode Desktop restart failures after saving * fix(ca): shorten OpenCode Desktop setup output * fix(ca): carry OpenCode2 provider sign-ins to cloud agents * fix(ca): stream setup scripts over SSH stdin * feat(ca): give OpenCode agents recognizable lowercase names * fix(ca): hide empty session hints in the agent menu * fix(ca): show session loading on the agent status icon * chore: Release railwayapp version 5.50.0 * fix: honor update opt-out and unify CLI/skill upgrades (#1178) * fix: honor update opt-out and sync skills after CLI upgrades * feat: unify CLI and managed skill update reporting * chore: Release railwayapp version 5.50.1 * fix(release): unblock i686 Windows GNU builds (#1182) * fix(release): use MinGW import libraries for i686 Windows builds * fix(release): pin modern MinGW image and align CI build steps * test(upgrade): serialize executable fixture launches on Unix * chore: Release railwayapp version 5.50.2 * Add OpenCode Desktop auto-configuration and railway code get-config (#1183) * feat: automatically configure detected OpenCode desktop clients * refactor: update OpenCode connection output without desktop restarts * refactor: present OpenCode results after clearing setup output * fix: show one OpenCode result panel after canceling connection * feat: replay the last cloud agent connection with railway code get-config * chore: Release railwayapp version 5.51.0 * fix: show saved code configuration in one results panel (#1185) * chore: Release railwayapp version 5.51.1 * fix: harden Herdr reconciliation and bootstrap retries Contribute scoped state, serialized reconciliation, reconnect recovery, and SSH identity handoff to #1176. Cover stale intent and interrupted onboarding with regression tests. * fix(ca): preserve agent identity across uncertain VM observations (#1186) * fix(ca): preserve agent identity across uncertain VM observations * fix(ca): harden operation watches against stale responses * chore: Release railwayapp version 5.51.2 * feat: add Codex remote clients, Desktop setup, and named config replay (#1188) * feat: add Codex remote clients and named connection config replay * fix: isolate remote Codex client settings and keep Desktop setup in background * fix: show Railway reconnect commands instead of raw client launch commands * fix: align Codex naming and saved-config results with OpenCode * chore: Release railwayapp version 5.52.0 * feat(code): opt in to code endpoints and restore from checkpoints (#1192) * feat(code): use the shared cloud agent code endpoint * feat(code): request optional endpoints and support checkpoint restores * chore: Release railwayapp version 5.52.1 * feat(ca): embed native clients and default harness launches to fresh VMs (#1193) * feat(ca): embed native clients and default harness launches to fresh VMs * feat(ca): browse Claude and Grok threads and open OpenCode home * fix(ca): keep native thread titles and sidebar refreshes stable * fix(ca): refresh cached conversations on demand * feat(ca): delete native conversations optimistically * fix(ca): preserve native client terminal colors * feat(ca): add bootstrap workflows and improve terminal navigation (#1195) * feat(ca): save and launch from named environment bootstraps * feat(ca): store bootstrap defaults locally and reuse internal APIs * feat(ca): create project bootstraps from the launcher * feat(ca): polish bootstrap flows and terminal navigation * test(ca): wait for complete PTY output on Windows * test(ca): canonicalize linked project fixture on macOS --------- Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com> * chore: Release railwayapp version 5.53.0 * Add public domains to sandbox create and fork (#1197) * Add public domains to sandbox create and fork * Keep sandbox usage documentation in the docs site * chore: Release railwayapp version 5.54.0 * fix(test): avoid reverse DNS in Codex bootstrap fixture (#1202) * Focus cloud-agent and sandbox command help (#1201) * Focus cloud-agent and sandbox command help * Complete bootstrap command help and examples --------- Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com> * chore: Release railwayapp version 5.54.1 * Run the local Railway TUI inside the CA flow (#1200) * Local Railway TUI over the public cloud-agent gate * Fix Codex hyperlinks and full-access permissions in CA * chore: Release railwayapp version 5.55.0 * feat(agent): add Pi support to setup tooling (#912) * chore: Release railwayapp version 5.56.0 * Commonify TUI theming: one shared Theme for every ratatui screen (#1204) * Extract shared Theme into src/tui_theme.rs Moves the `railway ca` TUI's Theme struct (previously commands/cloud_agent/tui/theme.rs) to a top-level module so every ratatui screen can share it instead of each redefining its own colors. Adds a `danger` field (used by delete confirmations, 5xx/p99, Cancel buttons — the one role none of the existing fields covered) and a shared, persisted theme preference at ~/.railway/tui-prefs.json, migrated once from cloud-agent's older agent-prefs.json so existing users don't see their theme reset. cloud_agent::tui::theme re-exports the new module so its existing consumers are unaffected; this step has no visible behavior change for `railway ca`. * Thread the shared Theme through the scale TUI Replaces railway scale's module-level Color constants with theme.<field> lookups from the shared preference, and adds a 't' key to cycle themes (previewing immediately and persisting the choice for every screen). Smallest of the remaining screens, so this proves the App-threading + cycle-key + shared-persistence pattern end-to-end before the larger ones. * Thread the shared Theme through the volume file browser Same pattern as the scale TUI: module-level Color constants become theme.<field> lookups, and 't' cycles the shared theme. Exercises the new `danger` field for the first time — delete confirmations now read as destructive (danger) while upload/download overwrite confirmations stay cautionary (pending/warning), a distinction the old single Yellow border couldn't make. * Thread the shared Theme through the metrics TUI Largest of the remaining screens, and the reason Theme grows a `series: &[Color]` field: metrics charts draw several simultaneous lines (CPU/memory limits aside, which reuse `accent`) — egress/ingress, p50/p90/p95/p95/p99, 2xx/3xx/4xx/5xx — that must stay visually distinct from each other, which the existing semantic slots (accent/running/pending/danger) don't have enough of on their own. Each theme now defines an ordered 4-color palette picked by index instead of the metrics TUI naming a `CPU_COLOR`/`P90_COLOR`/`STATUS_5XX` constant per metric. Both MetricsApp (single service) and ProjectApp (--all) get their own theme field, cycled with 'v' (metrics already uses 't' for the time range). * Thread the shared Theme through the templates picker The picker keeps its independent light/dark terminal-background detection (TerminalTheme, detect_terminal_theme, OSC background query) — that answers a different question ("is the user's terminal light or dark") than "which brand theme is selected," and stays exactly as it was. The PickerApp field that held it is renamed `bg` so it doesn't collide with the new `theme` field carrying the shared brand Theme. Only the picker's few brand-colored spots (the search input's border, the loading spinner, the verified badge, health-score coloring) move to the theme. Since every plain key types into the search box here, cycling the theme uses Ctrl+T instead of the bare letter the other screens use. * Thread the shared Theme through the dev/run TUI Last of the six screens. Leaves the colored::Color -> ratatui::Color passthrough mapper (convert_color) untouched — it renders subprocess log output colors, not chrome, and was never in scope for this pass. Everything else (tab bar, log selection highlight, info pane, help bar) now reads from the shared theme, cycled with 't' like the scale TUI and volume browser. * chore: Release railwayapp version 5.57.0 * fix(code): fall back to installed client when the latest check is rate-limited (#1205) * fix(code): fall back to installed client when the latest check is rate-limited `railway code --railway` resolves the newest railway-agent-tui on every launch by calling the unauthenticated GitHub API (api.github.com/.../agent-releases/ releases/latest), which is capped at 60 requests/hr per IP. When that quota is exhausted — a burst of launches, a shared NAT, other tooling on the same IP — the call returns 403 and `.error_for_status()` aborted the whole command, even though a perfectly good client was already installed under ~/.railway/runtimes/railway-tui/<version>/. Scope a fallback to the release-resolution step: when fetching the latest release metadata fails (rate limit, network, offline), use the newest installed client that still runs instead of blocking the launch. The fallback is deliberately limited to this network step — a later failure such as a download checksum mismatch is a real integrity problem and still surfaces, never papered over with an older binary. Adds tests for the rate-limited fallback (picks the newest runnable installed version, ignores non-version and empty dirs) and for the no-installed-client case (still errors, with a clear message). The existing corrupt-download test continues to assert a hard failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(code): retry the latest check with a GitHub token only when rate-limited Extends the installed-client fallback: when the anonymous latest-release check hits GitHub's rate limit (403/429 with x-ratelimit-remaining: 0), retry once authenticated before giving up, raising the cap from 60/hr per IP to 5000/hr for the token's account. The token is used ONLY to get past the rate limit — never attached to the normal request — so ordinary launches stay anonymous. It is resolved from GH_TOKEN, then GITHUB_TOKEN (matching the gh CLI's own precedence), then the signed-in gh CLI (`gh auth token`), best-effort and skipped silently when gh isn't installed or logged in. With no token, behaviour is unchanged: the error propagates and the newest installed client is used. Adds a GH_TOKEN > GITHUB_TOKEN > blank-falls-through precedence test; the rate-limit integration tests now serve GitHub's 403 + x-ratelimit-remaining: 0 shape so the outcome is identical whether or not the test host has a token. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * chore: Release railwayapp version 5.57.1 * Create fresh Railway agents by default and add new-VM project selection (#1203) * Default code to fresh Railway agents and add creation project selector * Keep new-agent picker spacing stable across focus changes * Make Option+N configure and create a fresh cloud agent * docs(code): explain client update fallback (#1207) * chore: Release railwayapp version 5.57.2 * fix(database): authenticate the generic HA switchover against the node's health server (#1177) * fix(database): send the node's HEALTH_API_PASSWORD on the generic HA switchover Clusters on the generic node contract (redis-ha, mysql-ha, mongo-ha) gate POST /switchover with HTTP Basic once the node carries HEALTH_API_PASSWORD. The switchover command now resolves that credential inside the target container (HEALTH_API_USERNAME defaults to railway) and hands curl -u user:pass, or nothing when the node has no password, so one command spans a cluster mid-rollout and no secret enters the exec payload or a log. * test(cluster_probe): run the sh-shim switchover tests on unix only The emitted command only ever runs inside a Linux container; Windows has no sh to check it against, so the shim-backed tests are gated like the Patroni ones. The prelude-shape test still runs everywhere. * fix(database): keep the HA switchover credential out of curl's argv (#1180) The generic HA switchover authenticates against the node's health server inside its container, and #1177 had the credential reach curl as `-u "$HEALTH_API_USER:$HEALTH_API_PW"`. Argv is public inside the container: /proc/<pid>/cmdline (ps) shows every process's arguments to every other process in the PID namespace for as long as the request runs, and under the ONE PASSWORD design that argument is the engine's root password. Hand curl the credential as a one-line config document on stdin instead (`user = "user:pass"` piped into `curl -K -`), escaped for curl's config parser by a pure-POSIX shell function (`\` and `"` backslashed, newline to `\n`). printf is a builtin of the shells the data images ship (dash, bash), so no process argv carries the secret at any point. Twin of the Patroni prelude's change. HEALTH_API_PASSWORD/HEALTH_API_USERNAME resolution, the bare request for an open node, and the error surfacing of the node's status and body are unchanged. The sh-shim tests now capture curl's stdin as well as its argv and assert the password is in the document and in no argument; a new test round-trips a password with every character curl's parser treats specially. * fix(postgres): keep the Patroni switchover credential out of curl's argv (#1179) The switchover authenticates against Patroni's REST API inside the member's container, and since #1171 the credential reached curl as `-u "$PATRONI_REST_USER:$PATRONI_REST_PW"`. Argv is public inside the container: /proc/<pid>/cmdline (ps) shows every process's arguments to every other process in the PID namespace for as long as the request runs, and under the ONE PASSWORD design that argument is the superuser password. Hand curl the credential as a one-line config document on stdin instead (`user = "user:pass"` piped into `curl -K -`), escaped for curl's config parser by a pure-POSIX shell function (`\` and `"` backslashed, newline to `\n`). printf is a builtin of the shells the data images ship (dash, bash), so no process argv carries the secret at any point. Resolution precedence, the bare request for a member with no password, and the error surfacing of Patroni's status and body are unchanged. The sh-shim tests now capture curl's stdin as well as its argv and assert the password is in the document and in no argument; a new test round- trips a password with every character curl's parser treats specially. * chore: Release railwayapp version 5.57.3 * chore: Release railwayapp version 5.57.4 * fix(postgres): wait out Patroni's switchover poll before declaring failure (#1208) Patroni holds POST /switchover open for up to ~20s while it polls the failover result. The CLI's 8s curl budget cut that short, so a completed handoff still surfaced as exit 28 / HTTP_STATUS:000. Raise the budget above Patroni's window and point operators at `ha status` when a wait still times out. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: Release railwayapp version 5.57.5 * rendergraphasrailway now seeds the resource ident pool (names) with… (#1209) Task: Seed pull codegen ident pool with source aliases Attempt 2 of 3 Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * railway config migrate treated platform configFile values like… (#1210) Task: Normalize leading slash in migrate configFile paths Attempt 3 of 3 Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * Set ok=false and exit non-zero on failed apply (#1211) * src/iac/engine.rs: after applychangeset, ok is now recomputed via… Task: Set ok=false and exit non-zero on failed apply Attempt 6 of 6 * The local branch had reset to master, so I restored the PR #1211… Task: Set ok=false and exit non-zero on failed apply Attempt 7 of 9 Follow-up: The operator re-ran this task. Review the current state of the branch, address anything outstanding, make the gate pass, then call report again. --------- Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * Classify imported databases by provenance, match by name (#1212) * Imported services are now classified as database only when the… Task: Classify imported databases by provenance, match by name Attempt 2 of 3 * Fixed the remaining defect in diffgraphs (src/iac/changeset.rs): a… Task: Classify imported databases by provenance, match by name Attempt 5 of 9 Follow-up: The operator re-ran this task. Review the current state of the branch, address anything outstanding, make the gate pass, then call report again. * The Installer (POSIX sh) failure is install.sh hitting GitHub's… Task: Classify imported databases by provenance, match by name Attempt 7 of 9 --------- Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * Replaced the "Last resort" comment above the named-partial export in… (#1213) Task: Reword migrate's 'last resort' partial comment Attempt 3 of 3 Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * chore: Release railwayapp version 5.57.6 * fix(code): upgrade cloud OpenCode V2 servers to match local clients (#1215) * chore: Release railwayapp version 5.57.7 * fix(iac): validate regions before creating buckets (#1206) Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com> Co-authored-by: Victor Ramirez <victor@railway.com> * fix(IaC): only match an image as a database if it matches a known name (#1214) * chore: Release railwayapp version 5.57.8 * fix(code): unify agent startup and project selection (#1216) * fix(code): unify agent startup and project selection * fix(code): trust Grok cloud-agent sessions by default * fix(ca): consolidate discovery into the global refresh shortcut * chore: Release railwayapp version 5.57.9 * feat(ca): offer to wake and resume an agent slept under a local client (#1219) An agent slept from outside the TUI (an idle timer, `railway ca sleep` elsewhere) takes its Codex or OpenCode server with it. The local client keeps retrying an address that no longer answers: the code domain does not wake a slept VM, and a woken one has no server until `reconnect` respawns it. On each agent refresh, an agent that went from awake to sleeping while this TUI holds a live local client pane on it raises the confirm modal: wake it and resume? A yes wakes the agent and, once the wake lands, respawns the harness server on the saved port with the saved credentials, so the client's own retries succeed with no restart. Sleeps this TUI asked for are not offered back; offers are capped at two per agent per run. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * chore: Release railwayapp version 5.57.10 * fix(code): repair OpenCode V2 readiness and saved history fallback (#1220) * chore: Release railwayapp version 5.57.11 * fix(up): show build errors and failure reasons (#1222) * fix(up): show build errors and failure reasons * fix(up): clean up deployment links * chore: Release railwayapp version 5.57.12 * feat(code): make stable OpenCode V2 the default (#1221) * feat(code): default new OpenCode launches to stable V2 * fix(code): use recorded OpenCode protocol when respawning after wake * fix(code): use OpenCode labels and oc agent names * chore: Release railwayapp version 5.58.0 * feat: add IaC partial ownership commands (#1226) * chore: Release railwayapp version 5.59.0 * feat(code): cut OpenCode over to V2 and retire the opencode2 shim (#1229) * feat(code): cut OpenCode over to V2 and retire the opencode2 shim OpenCode V2 is stable and the cloud-agent image's `opencode` is being hard cut over to it, so the CLI now launches the image's native `opencode` directly and drops the Railway-managed `~/.local/bin/opencode2` shim that downloaded `@opencode/cli` from npm at launch. - Fold Agent::OpenCode2 into Agent::OpenCode (slug/binary `opencode`). The V1 "OpenCode 1 - legacy" agent and its auth.json seeding are gone; V1 remains only as Protocol::V1 for reconnecting/upgrading managed servers started before the cutover. - Keep `--opencode2` as a hidden, deprecated alias of `--opencode` (with a stderr note); normalize the `opencode2` slug to `opencode` wherever it is read from prefs, panes, saved connections, VM history or backboard. Protocol::harness() now writes `opencode` for V2 connections. - Move opencode2/auth.rs to opencode/auth.rs and add opencode/import_auth.py, which has V2 create/migrate its own store (opencode api --standalone GET /api/info) and merges staged provider credentials into `credential` without overwriting remote accounts, backing up V1-only storage first. - Runtime seed: verify `opencode --version` is 2.x; on an older image run the official installer (https://opencode.ai/v2/install, fetched first, stdin closed) which upgrades ~/.opencode/bin/opencode in place, then remove the old shim. The Desktop/server bootstrap does the same and upgrades when the image is behind a saved server's release. Resuming a saved conversation on a 1.x VM explains how to upgrade instead. - Put ~/.opencode/bin first on HARNESS_PATH; update README and docs. Tests: drop tests/opencode2.py, retarget the credential import test, rework tests/opencode_desktop.py around in-place installs with a stubbed installer, add seed/alias/history coverage. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(code): drop OpenCode 1 support and the in-VM V2 auto-upgrade The cutover is a hard one: the image's `opencode` is V2 and the CLI runs it directly. Remove everything that existed only to run, reconnect, or migrate OpenCode 1, and never install or upgrade OpenCode on a VM. - Remove the in-place V2 install from the runtime seed and the server bootstrap. A VM whose `opencode --version` is not 2.x fails fast with: "This cloud agent is on an older image whose OpenCode is not V2. Create a new agent with railway code --opencode --new." The resume guard prints the same message. - Delete Protocol::V1 (and the Protocol enum), legacy_runtime, the V1 reconnect and `upgrade` client action, V1 configure_permissions, the V1 local-client npm install, V1 API paths/event shapes in the bridge and client sessions, the V1 auth.json credential conversion, and the V1 storage backup in import_auth.py. V2 credential import stays. - Remove the --opencode2 flag from `railway code` and `railway ca desktop`. - Keep read-side normalization of stored `opencode2` -> `opencode` (prefs, saved server.json, snapshots, history, panes). A saved OpenCode 1 connection record is dropped on read so the rest of the snapshot archive stays usable; the VM bootstrap refuses to reconnect a server started as OpenCode 1 but still stops it. - Docs and tests follow: V2-only bootstrap tests, old-image and V1 server refusals, V1 snapshot records rendered as plain SSH. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore: Release railwayapp version 5.59.1 * fix: correct OpenCode label in agent picker (#1228) * chore: Release railwayapp version 5.59.2 * feat: add `railway trace` for tracing settings and traces (#1230) `railway trace enable|disable|inherit` flip a service's tracing override, or the project default with `--project-default`, through `serviceUpdate` and `projectUpdate`. `status`, `list` and `get` read the settings, the per-service span activity, the trace list and one trace's spans; `list --json` and `get --json` print one object per line like `railway logs --json`. * chore: Release railwayapp version 5.60.0 * feat(database): add MySQL point-in-time recovery (#1144) * feat(database): add MySQL point-in-time recovery MySQL PITR is binlog archiving into a Railway bucket, enabled by the `mysql-pitr` composable overlay. The generic `pitr` tree already drives everything an engine's archive needs from its declared contract, so this is mostly the declaration plus the two ways MySQL genuinely differs from Postgres. Both differences become declarations rather than branches on the engine name: - `supports_ha: false`. The image's archiver refuses to run whenever the Group Replication seed list is set, so MySQL PITR is standalone-only. Every path that would take (or follow) the rolling HA workflow now checks this first -- `progress`/`cancel`/`clear` refuse before any network call, and `enable`/`disable` refuse before the HA mutation -- so the user gets one clear sentence instead of a server error from a workflow that was never going to start. - `probe_kind: None`. The live coverage probe in `pitr status` is pgBackRest shelling into the container, which MySQL has no equivalent tool for. Selecting the probe by declared kind means MySQL renders no coverage section at all, rather than a probe run with another engine's tooling or a fake "unavailable" for something it simply does not implement. The archive variable contract needs no code: `BINLOG_ARCHIVE_` as the declared prefix is enough for enabled-state detection and the enable overlay, and the image-eligibility rules ship with the `mysql-pitr` template record. Notably that template does NOT require a floating major tag, unlike postgres-pitr -- its image matrix only publishes exact minors, so a hardcoded "minor pins are bad" rule would have refused every MySQL image. Restore, backups and schedules ride the volume-instance mutations, which are engine-agnostic server-side and needed no change. * fix(database): MySQL PITR rolls HA clusters too mysql-ha archives from whichever member is the writable primary and backboard's enable-pitr-ha dispatches the Group Replication choreography on the registry's haRolloutKind (railwayapp/mono#38063, mysql-ha#45), so the standalone-only refusal on MySQL cluster roots is stale: enable/disable on a cluster root and progress/cancel/clear drive the rolling workflow, the same as Postgres. The gate stays for engines that declare no rolling workflow; the test now exercises it through a synthetic declaration. * fix(database): update the stale standalone-only assumption CI caught capability_declarations_match_what_ships still asserted MySQL PITR's supports_ha as false, which the previous commit had already flipped to true -- the test was written for the old invariant and never updated alongside it, so CI failed on the assertion. Also fixes the same stale assumption in `railway mysql`'s module doc comment and --help text, which still told users PITR is standalone-only on MySQL and cannot be enabled on an HA cluster. * chore: Release railwayapp version 5.61.0 * fix: avoid GitHub API rate limits during CLI installation (#1232) * fix(herdr): retry failed bootstrap launches and profile removals * ci: validate pull requests targeting stacked branches --------- Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com> Co-authored-by: Noah <hi@nd.mt> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com> Co-authored-by: Jake Runzer <jakerunzer@gmail.com> Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com> Co-authored-by: Arnav Gupta <championswimmer@gmail.com> Co-authored-by: Arnav Gupta <arnav@railway.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paulo Cabral Sanz <paulo@railway.app> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Víctor Ramírez <hey@futurepastori.com> Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> Co-authored-by: Fandu <113630375+mrfandu1@users.noreply.github.com> Co-authored-by: Victor Ramirez <victor@railway.com> Co-authored-by: Nebula <nebula@railway.app> Co-authored-by: Brody Over <10548119+brody192@users.noreply.github.com> Co-authored-by: Luca Casonato <hello@lcas.dev>
What changed
railway ca herdr {install,new,agents,sync,bootstrap}: cloud agents appear in herdr 0.9 as saved SSH machines.installwrites and links a herdr plugin manifest whose every command calls back into these verbs, so the plugin ships with the CLI and has no code of its own.bootstrapprovisions the VM in app mode (the samecode::prepareasca desktop), installs herdr's Claude and Codex integrations, sets[terminal] shell_mode/default_shelland[experimental] pane_history, adds a~/.profileblock sourcing the carried token, and creates an/appworkspace.~/.ssh/known_hostsfrom the CLI's~/.railway/known_hosts_relay.Why
herdr 0.9 runs one server per machine, and a cloud agent VM is exactly that: herdr owns the panes and detects the agent natively, and after a sleep its own restore brings the layout and
claude --resumeback. herdr cannot create, wake, sleep or delete the VM, which is what these verbs do. Relay durable sessions are not used.The ssh target is
ssh://agent%3A<env>%3A<id>@ssh.railway.com: herdr rejects a literal colon in the user part as a password, OpenSSH decodes the percent form. herdr's background reconnects run withStrictHostKeyChecking=yes, so the relay key has to be in~/.ssh/known_hosts. Pane shells on the VM fall back to/bin/shbecause the server inherits noSHELL, hencedefault_shell.Notes
newagainst a real create not yet run.lifecycle::place_namesandcode::HARNESS_PATHmadepub(crate).~/claude-brain/projects/2026-09-08-herdr-railway-ca-plugin.🤖 Generated with Claude Code
https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu