fix(database): authenticate the generic HA switchover against the node's health server - #1177
Merged
Merged
Conversation
…switchover Clusters on the generic node contract (redis-ha, mysql-ha, mongo-ha) gate POST /switchover with HTTP Basic once the node carries HEALTH_API_PASSWORD. The switchover command now resolves that credential inside the target container (HEALTH_API_USERNAME defaults to railway) and hands curl -u user:pass, or nothing when the node has no password, so one command spans a cluster mid-rollout and no secret enters the exec payload or a log.
The emitted command only ever runs inside a Linux container; Windows has no sh to check it against, so the shim-backed tests are gated like the Patroni ones. The prelude-shape test still runs everywhere.
This was referenced Sep 9, 2026
paulocsanz
added a commit
that referenced
this pull request
Sep 16, 2026
The generic HA switchover authenticates against the node's health server inside its container, and #1177 had the credential reach curl as `-u "$HEALTH_API_USER:$HEALTH_API_PW"`. Argv is public inside the container: /proc/<pid>/cmdline (ps) shows every process's arguments to every other process in the PID namespace for as long as the request runs, and under the ONE PASSWORD design that argument is the engine's root password. Hand curl the credential as a one-line config document on stdin instead (`user = "user:pass"` piped into `curl -K -`), escaped for curl's config parser by a pure-POSIX shell function (`\` and `"` backslashed, newline to `\n`). printf is a builtin of the shells the data images ship (dash, bash), so no process argv carries the secret at any point. Twin of the Patroni prelude's change. HEALTH_API_PASSWORD/HEALTH_API_USERNAME resolution, the bare request for an open node, and the error surfacing of the node's status and body are unchanged. The sh-shim tests now capture curl's stdin as well as its argv and assert the password is in the document and in no argument; a new test round-trips a password with every character curl's parser treats specially.
paulocsanz
added a commit
that referenced
this pull request
Sep 16, 2026
…1180) The generic HA switchover authenticates against the node's health server inside its container, and #1177 had the credential reach curl as `-u "$HEALTH_API_USER:$HEALTH_API_PW"`. Argv is public inside the container: /proc/<pid>/cmdline (ps) shows every process's arguments to every other process in the PID namespace for as long as the request runs, and under the ONE PASSWORD design that argument is the engine's root password. Hand curl the credential as a one-line config document on stdin instead (`user = "user:pass"` piped into `curl -K -`), escaped for curl's config parser by a pure-POSIX shell function (`\` and `"` backslashed, newline to `\n`). printf is a builtin of the shells the data images ship (dash, bash), so no process argv carries the secret at any point. Twin of the Patroni prelude's change. HEALTH_API_PASSWORD/HEALTH_API_USERNAME resolution, the bare request for an open node, and the error surfacing of the node's status and body are unchanged. The sh-shim tests now capture curl's stdin as well as its argv and assert the password is in the document and in no argument; a new test round-trips a password with every character curl's parser treats specially.
codyde
added a commit
that referenced
this pull request
Sep 23, 2026
* fix(run): cover execution-search variables in the host env filter (#1175) The local-spawn filter listed shell hooks, runtime hooks and the loader prefixes, but not the names that relocate where code is searched for. PATH picks which binary a bare command name resolves to and HOME picks which startup file an interactive shell sources, so both choose what the child runs without naming a command themselves. Add that category, plus the GIT_CONFIG prefix (GIT_CONFIG_KEY_<n> sets core.pager and core.sshCommand with no file involved, and the <n> makes the family open-ended) and the JVM, gem, lua and R startup hooks that were missing next to the Python and Perl ones. Dropping is still the right primitive: the spawn merges what survives onto the inherited environment, so a dropped name leaves the machine's own value in place for the child. * chore: Release railwayapp version 5.49.6 * feat(ca): add OpenCode desktop and local client connections (#1173) * feat(ca): add OpenCode desktop support * feat(ca): connect OpenCode Desktop through agent HTTPS * feat(ca): configure OpenCode Desktop automatically * fix(ca): protect saved OpenCode Desktop credentials * feat(ca): configure OpenCode Beta desktop stores * feat(ca): isolate OpenCode editions and launch the latest OpenCode2 Beta * feat(code): add authenticated OpenCode remote servers * feat(code): launch local OpenCode clients and discover remote servers * fix(opencode): highlight shared connection details * fix(ca): handle OpenCode Desktop restart failures after saving * fix(ca): shorten OpenCode Desktop setup output * fix(ca): carry OpenCode2 provider sign-ins to cloud agents * fix(ca): stream setup scripts over SSH stdin * feat(ca): give OpenCode agents recognizable lowercase names * fix(ca): hide empty session hints in the agent menu * fix(ca): show session loading on the agent status icon * chore: Release railwayapp version 5.50.0 * fix: honor update opt-out and unify CLI/skill upgrades (#1178) * fix: honor update opt-out and sync skills after CLI upgrades * feat: unify CLI and managed skill update reporting * chore: Release railwayapp version 5.50.1 * fix(release): unblock i686 Windows GNU builds (#1182) * fix(release): use MinGW import libraries for i686 Windows builds * fix(release): pin modern MinGW image and align CI build steps * test(upgrade): serialize executable fixture launches on Unix * chore: Release railwayapp version 5.50.2 * Add OpenCode Desktop auto-configuration and railway code get-config (#1183) * feat: automatically configure detected OpenCode desktop clients * refactor: update OpenCode connection output without desktop restarts * refactor: present OpenCode results after clearing setup output * fix: show one OpenCode result panel after canceling connection * feat: replay the last cloud agent connection with railway code get-config * chore: Release railwayapp version 5.51.0 * fix: show saved code configuration in one results panel (#1185) * chore: Release railwayapp version 5.51.1 * fix: harden Herdr reconciliation and bootstrap retries Contribute scoped state, serialized reconciliation, reconnect recovery, and SSH identity handoff to #1176. Cover stale intent and interrupted onboarding with regression tests. * fix(ca): preserve agent identity across uncertain VM observations (#1186) * fix(ca): preserve agent identity across uncertain VM observations * fix(ca): harden operation watches against stale responses * chore: Release railwayapp version 5.51.2 * feat: add Codex remote clients, Desktop setup, and named config replay (#1188) * feat: add Codex remote clients and named connection config replay * fix: isolate remote Codex client settings and keep Desktop setup in background * fix: show Railway reconnect commands instead of raw client launch commands * fix: align Codex naming and saved-config results with OpenCode * chore: Release railwayapp version 5.52.0 * feat(code): opt in to code endpoints and restore from checkpoints (#1192) * feat(code): use the shared cloud agent code endpoint * feat(code): request optional endpoints and support checkpoint restores * chore: Release railwayapp version 5.52.1 * feat(ca): embed native clients and default harness launches to fresh VMs (#1193) * feat(ca): embed native clients and default harness launches to fresh VMs * feat(ca): browse Claude and Grok threads and open OpenCode home * fix(ca): keep native thread titles and sidebar refreshes stable * fix(ca): refresh cached conversations on demand * feat(ca): delete native conversations optimistically * fix(ca): preserve native client terminal colors * feat(ca): add bootstrap workflows and improve terminal navigation (#1195) * feat(ca): save and launch from named environment bootstraps * feat(ca): store bootstrap defaults locally and reuse internal APIs * feat(ca): create project bootstraps from the launcher * feat(ca): polish bootstrap flows and terminal navigation * test(ca): wait for complete PTY output on Windows * test(ca): canonicalize linked project fixture on macOS --------- Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com> * chore: Release railwayapp version 5.53.0 * Add public domains to sandbox create and fork (#1197) * Add public domains to sandbox create and fork * Keep sandbox usage documentation in the docs site * chore: Release railwayapp version 5.54.0 * fix(test): avoid reverse DNS in Codex bootstrap fixture (#1202) * Focus cloud-agent and sandbox command help (#1201) * Focus cloud-agent and sandbox command help * Complete bootstrap command help and examples --------- Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com> * chore: Release railwayapp version 5.54.1 * Run the local Railway TUI inside the CA flow (#1200) * Local Railway TUI over the public cloud-agent gate * Fix Codex hyperlinks and full-access permissions in CA * chore: Release railwayapp version 5.55.0 * feat(agent): add Pi support to setup tooling (#912) * chore: Release railwayapp version 5.56.0 * Commonify TUI theming: one shared Theme for every ratatui screen (#1204) * Extract shared Theme into src/tui_theme.rs Moves the `railway ca` TUI's Theme struct (previously commands/cloud_agent/tui/theme.rs) to a top-level module so every ratatui screen can share it instead of each redefining its own colors. Adds a `danger` field (used by delete confirmations, 5xx/p99, Cancel buttons — the one role none of the existing fields covered) and a shared, persisted theme preference at ~/.railway/tui-prefs.json, migrated once from cloud-agent's older agent-prefs.json so existing users don't see their theme reset. cloud_agent::tui::theme re-exports the new module so its existing consumers are unaffected; this step has no visible behavior change for `railway ca`. * Thread the shared Theme through the scale TUI Replaces railway scale's module-level Color constants with theme.<field> lookups from the shared preference, and adds a 't' key to cycle themes (previewing immediately and persisting the choice for every screen). Smallest of the remaining screens, so this proves the App-threading + cycle-key + shared-persistence pattern end-to-end before the larger ones. * Thread the shared Theme through the volume file browser Same pattern as the scale TUI: module-level Color constants become theme.<field> lookups, and 't' cycles the shared theme. Exercises the new `danger` field for the first time — delete confirmations now read as destructive (danger) while upload/download overwrite confirmations stay cautionary (pending/warning), a distinction the old single Yellow border couldn't make. * Thread the shared Theme through the metrics TUI Largest of the remaining screens, and the reason Theme grows a `series: &[Color]` field: metrics charts draw several simultaneous lines (CPU/memory limits aside, which reuse `accent`) — egress/ingress, p50/p90/p95/p95/p99, 2xx/3xx/4xx/5xx — that must stay visually distinct from each other, which the existing semantic slots (accent/running/pending/danger) don't have enough of on their own. Each theme now defines an ordered 4-color palette picked by index instead of the metrics TUI naming a `CPU_COLOR`/`P90_COLOR`/`STATUS_5XX` constant per metric. Both MetricsApp (single service) and ProjectApp (--all) get their own theme field, cycled with 'v' (metrics already uses 't' for the time range). * Thread the shared Theme through the templates picker The picker keeps its independent light/dark terminal-background detection (TerminalTheme, detect_terminal_theme, OSC background query) — that answers a different question ("is the user's terminal light or dark") than "which brand theme is selected," and stays exactly as it was. The PickerApp field that held it is renamed `bg` so it doesn't collide with the new `theme` field carrying the shared brand Theme. Only the picker's few brand-colored spots (the search input's border, the loading spinner, the verified badge, health-score coloring) move to the theme. Since every plain key types into the search box here, cycling the theme uses Ctrl+T instead of the bare letter the other screens use. * Thread the shared Theme through the dev/run TUI Last of the six screens. Leaves the colored::Color -> ratatui::Color passthrough mapper (convert_color) untouched — it renders subprocess log output colors, not chrome, and was never in scope for this pass. Everything else (tab bar, log selection highlight, info pane, help bar) now reads from the shared theme, cycled with 't' like the scale TUI and volume browser. * chore: Release railwayapp version 5.57.0 * fix(code): fall back to installed client when the latest check is rate-limited (#1205) * fix(code): fall back to installed client when the latest check is rate-limited `railway code --railway` resolves the newest railway-agent-tui on every launch by calling the unauthenticated GitHub API (api.github.com/.../agent-releases/ releases/latest), which is capped at 60 requests/hr per IP. When that quota is exhausted — a burst of launches, a shared NAT, other tooling on the same IP — the call returns 403 and `.error_for_status()` aborted the whole command, even though a perfectly good client was already installed under ~/.railway/runtimes/railway-tui/<version>/. Scope a fallback to the release-resolution step: when fetching the latest release metadata fails (rate limit, network, offline), use the newest installed client that still runs instead of blocking the launch. The fallback is deliberately limited to this network step — a later failure such as a download checksum mismatch is a real integrity problem and still surfaces, never papered over with an older binary. Adds tests for the rate-limited fallback (picks the newest runnable installed version, ignores non-version and empty dirs) and for the no-installed-client case (still errors, with a clear message). The existing corrupt-download test continues to assert a hard failure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(code): retry the latest check with a GitHub token only when rate-limited Extends the installed-client fallback: when the anonymous latest-release check hits GitHub's rate limit (403/429 with x-ratelimit-remaining: 0), retry once authenticated before giving up, raising the cap from 60/hr per IP to 5000/hr for the token's account. The token is used ONLY to get past the rate limit — never attached to the normal request — so ordinary launches stay anonymous. It is resolved from GH_TOKEN, then GITHUB_TOKEN (matching the gh CLI's own precedence), then the signed-in gh CLI (`gh auth token`), best-effort and skipped silently when gh isn't installed or logged in. With no token, behaviour is unchanged: the error propagates and the newest installed client is used. Adds a GH_TOKEN > GITHUB_TOKEN > blank-falls-through precedence test; the rate-limit integration tests now serve GitHub's 403 + x-ratelimit-remaining: 0 shape so the outcome is identical whether or not the test host has a token. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> * chore: Release railwayapp version 5.57.1 * Create fresh Railway agents by default and add new-VM project selection (#1203) * Default code to fresh Railway agents and add creation project selector * Keep new-agent picker spacing stable across focus changes * Make Option+N configure and create a fresh cloud agent * docs(code): explain client update fallback (#1207) * chore: Release railwayapp version 5.57.2 * fix(database): authenticate the generic HA switchover against the node's health server (#1177) * fix(database): send the node's HEALTH_API_PASSWORD on the generic HA switchover Clusters on the generic node contract (redis-ha, mysql-ha, mongo-ha) gate POST /switchover with HTTP Basic once the node carries HEALTH_API_PASSWORD. The switchover command now resolves that credential inside the target container (HEALTH_API_USERNAME defaults to railway) and hands curl -u user:pass, or nothing when the node has no password, so one command spans a cluster mid-rollout and no secret enters the exec payload or a log. * test(cluster_probe): run the sh-shim switchover tests on unix only The emitted command only ever runs inside a Linux container; Windows has no sh to check it against, so the shim-backed tests are gated like the Patroni ones. The prelude-shape test still runs everywhere. * fix(database): keep the HA switchover credential out of curl's argv (#1180) The generic HA switchover authenticates against the node's health server inside its container, and #1177 had the credential reach curl as `-u "$HEALTH_API_USER:$HEALTH_API_PW"`. Argv is public inside the container: /proc/<pid>/cmdline (ps) shows every process's arguments to every other process in the PID namespace for as long as the request runs, and under the ONE PASSWORD design that argument is the engine's root password. Hand curl the credential as a one-line config document on stdin instead (`user = "user:pass"` piped into `curl -K -`), escaped for curl's config parser by a pure-POSIX shell function (`\` and `"` backslashed, newline to `\n`). printf is a builtin of the shells the data images ship (dash, bash), so no process argv carries the secret at any point. Twin of the Patroni prelude's change. HEALTH_API_PASSWORD/HEALTH_API_USERNAME resolution, the bare request for an open node, and the error surfacing of the node's status and body are unchanged. The sh-shim tests now capture curl's stdin as well as its argv and assert the password is in the document and in no argument; a new test round-trips a password with every character curl's parser treats specially. * fix(postgres): keep the Patroni switchover credential out of curl's argv (#1179) The switchover authenticates against Patroni's REST API inside the member's container, and since #1171 the credential reached curl as `-u "$PATRONI_REST_USER:$PATRONI_REST_PW"`. Argv is public inside the container: /proc/<pid>/cmdline (ps) shows every process's arguments to every other process in the PID namespace for as long as the request runs, and under the ONE PASSWORD design that argument is the superuser password. Hand curl the credential as a one-line config document on stdin instead (`user = "user:pass"` piped into `curl -K -`), escaped for curl's config parser by a pure-POSIX shell function (`\` and `"` backslashed, newline to `\n`). printf is a builtin of the shells the data images ship (dash, bash), so no process argv carries the secret at any point. Resolution precedence, the bare request for a member with no password, and the error surfacing of Patroni's status and body are unchanged. The sh-shim tests now capture curl's stdin as well as its argv and assert the password is in the document and in no argument; a new test round- trips a password with every character curl's parser treats specially. * chore: Release railwayapp version 5.57.3 * chore: Release railwayapp version 5.57.4 * fix(postgres): wait out Patroni's switchover poll before declaring failure (#1208) Patroni holds POST /switchover open for up to ~20s while it polls the failover result. The CLI's 8s curl budget cut that short, so a completed handoff still surfaced as exit 28 / HTTP_STATUS:000. Raise the budget above Patroni's window and point operators at `ha status` when a wait still times out. Co-authored-by: Cursor <cursoragent@cursor.com> * chore: Release railwayapp version 5.57.5 * rendergraphasrailway now seeds the resource ident pool (names) with… (#1209) Task: Seed pull codegen ident pool with source aliases Attempt 2 of 3 Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * railway config migrate treated platform configFile values like… (#1210) Task: Normalize leading slash in migrate configFile paths Attempt 3 of 3 Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * Set ok=false and exit non-zero on failed apply (#1211) * src/iac/engine.rs: after applychangeset, ok is now recomputed via… Task: Set ok=false and exit non-zero on failed apply Attempt 6 of 6 * The local branch had reset to master, so I restored the PR #1211… Task: Set ok=false and exit non-zero on failed apply Attempt 7 of 9 Follow-up: The operator re-ran this task. Review the current state of the branch, address anything outstanding, make the gate pass, then call report again. --------- Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * Classify imported databases by provenance, match by name (#1212) * Imported services are now classified as database only when the… Task: Classify imported databases by provenance, match by name Attempt 2 of 3 * Fixed the remaining defect in diffgraphs (src/iac/changeset.rs): a… Task: Classify imported databases by provenance, match by name Attempt 5 of 9 Follow-up: The operator re-ran this task. Review the current state of the branch, address anything outstanding, make the gate pass, then call report again. * The Installer (POSIX sh) failure is install.sh hitting GitHub's… Task: Classify imported databases by provenance, match by name Attempt 7 of 9 --------- Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * Replaced the "Last resort" comment above the named-partial export in… (#1213) Task: Reword migrate's 'last resort' partial comment Attempt 3 of 3 Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> * chore: Release railwayapp version 5.57.6 * fix(code): upgrade cloud OpenCode V2 servers to match local clients (#1215) * chore: Release railwayapp version 5.57.7 * fix(iac): validate regions before creating buckets (#1206) Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com> Co-authored-by: Victor Ramirez <victor@railway.com> * fix(IaC): only match an image as a database if it matches a known name (#1214) * chore: Release railwayapp version 5.57.8 * fix(code): unify agent startup and project selection (#1216) * fix(code): unify agent startup and project selection * fix(code): trust Grok cloud-agent sessions by default * fix(ca): consolidate discovery into the global refresh shortcut * chore: Release railwayapp version 5.57.9 * feat(ca): offer to wake and resume an agent slept under a local client (#1219) An agent slept from outside the TUI (an idle timer, `railway ca sleep` elsewhere) takes its Codex or OpenCode server with it. The local client keeps retrying an address that no longer answers: the code domain does not wake a slept VM, and a woken one has no server until `reconnect` respawns it. On each agent refresh, an agent that went from awake to sleeping while this TUI holds a live local client pane on it raises the confirm modal: wake it and resume? A yes wakes the agent and, once the wake lands, respawns the harness server on the saved port with the saved credentials, so the client's own retries succeed with no restart. Sleeps this TUI asked for are not offered back; offers are capped at two per agent per run. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> * chore: Release railwayapp version 5.57.10 * fix(code): repair OpenCode V2 readiness and saved history fallback (#1220) * chore: Release railwayapp version 5.57.11 * fix(up): show build errors and failure reasons (#1222) * fix(up): show build errors and failure reasons * fix(up): clean up deployment links * chore: Release railwayapp version 5.57.12 * feat(code): make stable OpenCode V2 the default (#1221) * feat(code): default new OpenCode launches to stable V2 * fix(code): use recorded OpenCode protocol when respawning after wake * fix(code): use OpenCode labels and oc agent names * chore: Release railwayapp version 5.58.0 * feat: add IaC partial ownership commands (#1226) * chore: Release railwayapp version 5.59.0 * feat(code): cut OpenCode over to V2 and retire the opencode2 shim (#1229) * feat(code): cut OpenCode over to V2 and retire the opencode2 shim OpenCode V2 is stable and the cloud-agent image's `opencode` is being hard cut over to it, so the CLI now launches the image's native `opencode` directly and drops the Railway-managed `~/.local/bin/opencode2` shim that downloaded `@opencode/cli` from npm at launch. - Fold Agent::OpenCode2 into Agent::OpenCode (slug/binary `opencode`). The V1 "OpenCode 1 - legacy" agent and its auth.json seeding are gone; V1 remains only as Protocol::V1 for reconnecting/upgrading managed servers started before the cutover. - Keep `--opencode2` as a hidden, deprecated alias of `--opencode` (with a stderr note); normalize the `opencode2` slug to `opencode` wherever it is read from prefs, panes, saved connections, VM history or backboard. Protocol::harness() now writes `opencode` for V2 connections. - Move opencode2/auth.rs to opencode/auth.rs and add opencode/import_auth.py, which has V2 create/migrate its own store (opencode api --standalone GET /api/info) and merges staged provider credentials into `credential` without overwriting remote accounts, backing up V1-only storage first. - Runtime seed: verify `opencode --version` is 2.x; on an older image run the official installer (https://opencode.ai/v2/install, fetched first, stdin closed) which upgrades ~/.opencode/bin/opencode in place, then remove the old shim. The Desktop/server bootstrap does the same and upgrades when the image is behind a saved server's release. Resuming a saved conversation on a 1.x VM explains how to upgrade instead. - Put ~/.opencode/bin first on HARNESS_PATH; update README and docs. Tests: drop tests/opencode2.py, retarget the credential import test, rework tests/opencode_desktop.py around in-place installs with a stubbed installer, add seed/alias/history coverage. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * refactor(code): drop OpenCode 1 support and the in-VM V2 auto-upgrade The cutover is a hard one: the image's `opencode` is V2 and the CLI runs it directly. Remove everything that existed only to run, reconnect, or migrate OpenCode 1, and never install or upgrade OpenCode on a VM. - Remove the in-place V2 install from the runtime seed and the server bootstrap. A VM whose `opencode --version` is not 2.x fails fast with: "This cloud agent is on an older image whose OpenCode is not V2. Create a new agent with railway code --opencode --new." The resume guard prints the same message. - Delete Protocol::V1 (and the Protocol enum), legacy_runtime, the V1 reconnect and `upgrade` client action, V1 configure_permissions, the V1 local-client npm install, V1 API paths/event shapes in the bridge and client sessions, the V1 auth.json credential conversion, and the V1 storage backup in import_auth.py. V2 credential import stays. - Remove the --opencode2 flag from `railway code` and `railway ca desktop`. - Keep read-side normalization of stored `opencode2` -> `opencode` (prefs, saved server.json, snapshots, history, panes). A saved OpenCode 1 connection record is dropped on read so the rest of the snapshot archive stays usable; the VM bootstrap refuses to reconnect a server started as OpenCode 1 but still stops it. - Docs and tests follow: V2-only bootstrap tests, old-image and V1 server refusals, V1 snapshot records rendered as plain SSH. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * chore: Release railwayapp version 5.59.1 * fix: correct OpenCode label in agent picker (#1228) * chore: Release railwayapp version 5.59.2 * feat: add `railway trace` for tracing settings and traces (#1230) `railway trace enable|disable|inherit` flip a service's tracing override, or the project default with `--project-default`, through `serviceUpdate` and `projectUpdate`. `status`, `list` and `get` read the settings, the per-service span activity, the trace list and one trace's spans; `list --json` and `get --json` print one object per line like `railway logs --json`. * chore: Release railwayapp version 5.60.0 * feat(database): add MySQL point-in-time recovery (#1144) * feat(database): add MySQL point-in-time recovery MySQL PITR is binlog archiving into a Railway bucket, enabled by the `mysql-pitr` composable overlay. The generic `pitr` tree already drives everything an engine's archive needs from its declared contract, so this is mostly the declaration plus the two ways MySQL genuinely differs from Postgres. Both differences become declarations rather than branches on the engine name: - `supports_ha: false`. The image's archiver refuses to run whenever the Group Replication seed list is set, so MySQL PITR is standalone-only. Every path that would take (or follow) the rolling HA workflow now checks this first -- `progress`/`cancel`/`clear` refuse before any network call, and `enable`/`disable` refuse before the HA mutation -- so the user gets one clear sentence instead of a server error from a workflow that was never going to start. - `probe_kind: None`. The live coverage probe in `pitr status` is pgBackRest shelling into the container, which MySQL has no equivalent tool for. Selecting the probe by declared kind means MySQL renders no coverage section at all, rather than a probe run with another engine's tooling or a fake "unavailable" for something it simply does not implement. The archive variable contract needs no code: `BINLOG_ARCHIVE_` as the declared prefix is enough for enabled-state detection and the enable overlay, and the image-eligibility rules ship with the `mysql-pitr` template record. Notably that template does NOT require a floating major tag, unlike postgres-pitr -- its image matrix only publishes exact minors, so a hardcoded "minor pins are bad" rule would have refused every MySQL image. Restore, backups and schedules ride the volume-instance mutations, which are engine-agnostic server-side and needed no change. * fix(database): MySQL PITR rolls HA clusters too mysql-ha archives from whichever member is the writable primary and backboard's enable-pitr-ha dispatches the Group Replication choreography on the registry's haRolloutKind (railwayapp/mono#38063, mysql-ha#45), so the standalone-only refusal on MySQL cluster roots is stale: enable/disable on a cluster root and progress/cancel/clear drive the rolling workflow, the same as Postgres. The gate stays for engines that declare no rolling workflow; the test now exercises it through a synthetic declaration. * fix(database): update the stale standalone-only assumption CI caught capability_declarations_match_what_ships still asserted MySQL PITR's supports_ha as false, which the previous commit had already flipped to true -- the test was written for the old invariant and never updated alongside it, so CI failed on the assertion. Also fixes the same stale assumption in `railway mysql`'s module doc comment and --help text, which still told users PITR is standalone-only on MySQL and cannot be enabled on an HA cluster. * chore: Release railwayapp version 5.61.0 * fix: avoid GitHub API rate limits during CLI installation (#1232) * fix(herdr): retry failed bootstrap launches and profile removals * ci: validate pull requests targeting stacked branches --------- Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com> Co-authored-by: Noah <hi@nd.mt> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com> Co-authored-by: Jake Runzer <jakerunzer@gmail.com> Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com> Co-authored-by: Arnav Gupta <championswimmer@gmail.com> Co-authored-by: Arnav Gupta <arnav@railway.com> Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Paulo Cabral Sanz <paulo@railway.app> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Víctor Ramírez <hey@futurepastori.com> Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com> Co-authored-by: Fandu <113630375+mrfandu1@users.noreply.github.com> Co-authored-by: Victor Ramirez <victor@railway.com> Co-authored-by: Nebula <nebula@railway.app> Co-authored-by: Brody Over <10548119+brody192@users.noreply.github.com> Co-authored-by: Luca Casonato <hello@lcas.dev>
codyde
added a commit
that referenced
this pull request
Sep 23, 2026
* feat(ca): herdr integration for cloud agents (experiment)
Adds `railway ca herdr {install,new,agents,sync,bootstrap}` so cloud
agents show in herdr 0.9 as saved SSH machines. `install` writes and
links a herdr plugin manifest whose every command calls back into these
verbs; `bootstrap` provisions the VM in app mode and prepares its herdr
server.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* feat(ca): herdr plugin follow-ups: remote keys, wake picker, attention recovery
`install` writes and updates the keybindings itself. Bootstrap links a
plugin on the VM too, so prefix+shift+a and prefix+shift+s (sleep this
agent) work while a machine is selected, and starts the chosen harness in
the /app pane. Local prefix+shift+s wakes. `sync` runs on workspace focus
(debounced), remembers agent status, waits for the ssh relay before
enabling a machine, and reconnects machines herdr parked in Attention by
reading the client log. Picker closes after connect/new/wake, notifies
before sleep and wake, caches project names.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* feat(ca): herdr watcher on the cloud agent subscription, manual sync key
`railway ca herdr watch` subscribes to backboard's cloudAgentInvalidation
per environment (the unpublished /graphql/internal graph the apps use) and
reconciles machines on each event, so a sleep from anywhere disables the
machine before herdr's reconnect can park it in Attention and a wake
re-enables it once the relay answers. One watcher per herdr session,
started by install and the startup hook, nudged by SIGUSR1, no polling.
Also: prefix+shift+y runs sync with a toast; the sweep hook moves to
pane.agent_status_changed (focus events are client-local in herdr 0.9 and
never reach hooks); the client-log reading is removed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* fix(ca): harden the herdr plugin after adversarial review
Watcher pidfile only counts a live `railway ca herdr watch` process, so a
reused pid is never signalled or trusted; nudge happens after the debounce;
the detached watcher logs quietly with rotation and exits when the session
socket stops answering. `sync` claims its window before the relay wait,
removes only machines it recorded itself and never on an empty agent list,
keeps the old status when an Enable/Kick fails so it is retried, and reports
failures in its summary and toast. `new` waits for the relay before
`herdr machine add` like every other path. Bootstrap ends in
BOOTSTRAP-FAILED when a step fails, accepts `[table] # comment` headers, and
feeds the token to curl on stdin. `install --remove` works with herdr
stopped and stops the watcher; known_hosts is appended, never rewritten.
herdr verbs are no longer tracked per invocation.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* feat(ca): two herdr keys, picker-hosted new/sync, upgrade the VM's railway
Bindings shrink to prefix+shift+a (the picker, which now leads with "new
agent" and "sync now") and prefix+shift+s (wake). Bootstrap upgrades the
VM's railway in place when it is older than the one running bootstrap, so
the remote picker works as soon as a release carries `ca herdr`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* fix(ca): warn at install when ssh config disarms the relay's known_hosts
herdr reconnects with strict host-key checking and no config of its own. A
Host block that sends ssh.railway.com to UserKnownHostsFile /dev/null parks
every Railway machine in Attention after the first network blip, which
herdr never retries. `install` now runs `ssh -G` on the relay host and says
so, naming the fix.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* fix(ca): build the herdr watcher on Windows
Socket liveness, the SIGUSR1 nudge and the ps/kill pid checks are unix
only; elsewhere they compile to no-ops. The plugin declares linux and macos,
so on Windows the binary builds and the herdr verbs simply do nothing extra.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* test(ca): unix-only for tests that spawn the fake herdr script
The fake is a shebang script, which Windows cannot execute.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vvwnaf9R6RjfPTrRo7MyDu
* fix: harden Herdr reconciliation and bootstrap retries (#1187)
* fix(run): cover execution-search variables in the host env filter (#1175)
The local-spawn filter listed shell hooks, runtime hooks and the loader
prefixes, but not the names that relocate where code is searched for.
PATH picks which binary a bare command name resolves to and HOME picks
which startup file an interactive shell sources, so both choose what the
child runs without naming a command themselves.
Add that category, plus the GIT_CONFIG prefix (GIT_CONFIG_KEY_<n> sets
core.pager and core.sshCommand with no file involved, and the <n> makes
the family open-ended) and the JVM, gem, lua and R startup hooks that
were missing next to the Python and Perl ones.
Dropping is still the right primitive: the spawn merges what survives
onto the inherited environment, so a dropped name leaves the machine's
own value in place for the child.
* chore: Release railwayapp version 5.49.6
* feat(ca): add OpenCode desktop and local client connections (#1173)
* feat(ca): add OpenCode desktop support
* feat(ca): connect OpenCode Desktop through agent HTTPS
* feat(ca): configure OpenCode Desktop automatically
* fix(ca): protect saved OpenCode Desktop credentials
* feat(ca): configure OpenCode Beta desktop stores
* feat(ca): isolate OpenCode editions and launch the latest OpenCode2 Beta
* feat(code): add authenticated OpenCode remote servers
* feat(code): launch local OpenCode clients and discover remote servers
* fix(opencode): highlight shared connection details
* fix(ca): handle OpenCode Desktop restart failures after saving
* fix(ca): shorten OpenCode Desktop setup output
* fix(ca): carry OpenCode2 provider sign-ins to cloud agents
* fix(ca): stream setup scripts over SSH stdin
* feat(ca): give OpenCode agents recognizable lowercase names
* fix(ca): hide empty session hints in the agent menu
* fix(ca): show session loading on the agent status icon
* chore: Release railwayapp version 5.50.0
* fix: honor update opt-out and unify CLI/skill upgrades (#1178)
* fix: honor update opt-out and sync skills after CLI upgrades
* feat: unify CLI and managed skill update reporting
* chore: Release railwayapp version 5.50.1
* fix(release): unblock i686 Windows GNU builds (#1182)
* fix(release): use MinGW import libraries for i686 Windows builds
* fix(release): pin modern MinGW image and align CI build steps
* test(upgrade): serialize executable fixture launches on Unix
* chore: Release railwayapp version 5.50.2
* Add OpenCode Desktop auto-configuration and railway code get-config (#1183)
* feat: automatically configure detected OpenCode desktop clients
* refactor: update OpenCode connection output without desktop restarts
* refactor: present OpenCode results after clearing setup output
* fix: show one OpenCode result panel after canceling connection
* feat: replay the last cloud agent connection with railway code get-config
* chore: Release railwayapp version 5.51.0
* fix: show saved code configuration in one results panel (#1185)
* chore: Release railwayapp version 5.51.1
* fix: harden Herdr reconciliation and bootstrap retries
Contribute scoped state, serialized reconciliation, reconnect recovery, and SSH identity handoff to #1176. Cover stale intent and interrupted onboarding with regression tests.
* fix(ca): preserve agent identity across uncertain VM observations (#1186)
* fix(ca): preserve agent identity across uncertain VM observations
* fix(ca): harden operation watches against stale responses
* chore: Release railwayapp version 5.51.2
* feat: add Codex remote clients, Desktop setup, and named config replay (#1188)
* feat: add Codex remote clients and named connection config replay
* fix: isolate remote Codex client settings and keep Desktop setup in background
* fix: show Railway reconnect commands instead of raw client launch commands
* fix: align Codex naming and saved-config results with OpenCode
* chore: Release railwayapp version 5.52.0
* feat(code): opt in to code endpoints and restore from checkpoints (#1192)
* feat(code): use the shared cloud agent code endpoint
* feat(code): request optional endpoints and support checkpoint restores
* chore: Release railwayapp version 5.52.1
* feat(ca): embed native clients and default harness launches to fresh VMs (#1193)
* feat(ca): embed native clients and default harness launches to fresh VMs
* feat(ca): browse Claude and Grok threads and open OpenCode home
* fix(ca): keep native thread titles and sidebar refreshes stable
* fix(ca): refresh cached conversations on demand
* feat(ca): delete native conversations optimistically
* fix(ca): preserve native client terminal colors
* feat(ca): add bootstrap workflows and improve terminal navigation (#1195)
* feat(ca): save and launch from named environment bootstraps
* feat(ca): store bootstrap defaults locally and reuse internal APIs
* feat(ca): create project bootstraps from the launcher
* feat(ca): polish bootstrap flows and terminal navigation
* test(ca): wait for complete PTY output on Windows
* test(ca): canonicalize linked project fixture on macOS
---------
Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com>
* chore: Release railwayapp version 5.53.0
* Add public domains to sandbox create and fork (#1197)
* Add public domains to sandbox create and fork
* Keep sandbox usage documentation in the docs site
* chore: Release railwayapp version 5.54.0
* fix(test): avoid reverse DNS in Codex bootstrap fixture (#1202)
* Focus cloud-agent and sandbox command help (#1201)
* Focus cloud-agent and sandbox command help
* Complete bootstrap command help and examples
---------
Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com>
* chore: Release railwayapp version 5.54.1
* Run the local Railway TUI inside the CA flow (#1200)
* Local Railway TUI over the public cloud-agent gate
* Fix Codex hyperlinks and full-access permissions in CA
* chore: Release railwayapp version 5.55.0
* feat(agent): add Pi support to setup tooling (#912)
* chore: Release railwayapp version 5.56.0
* Commonify TUI theming: one shared Theme for every ratatui screen (#1204)
* Extract shared Theme into src/tui_theme.rs
Moves the `railway ca` TUI's Theme struct (previously
commands/cloud_agent/tui/theme.rs) to a top-level module so every ratatui
screen can share it instead of each redefining its own colors. Adds a
`danger` field (used by delete confirmations, 5xx/p99, Cancel buttons — the
one role none of the existing fields covered) and a shared, persisted theme
preference at ~/.railway/tui-prefs.json, migrated once from cloud-agent's
older agent-prefs.json so existing users don't see their theme reset.
cloud_agent::tui::theme re-exports the new module so its existing consumers
are unaffected; this step has no visible behavior change for `railway ca`.
* Thread the shared Theme through the scale TUI
Replaces railway scale's module-level Color constants with theme.<field>
lookups from the shared preference, and adds a 't' key to cycle themes
(previewing immediately and persisting the choice for every screen).
Smallest of the remaining screens, so this proves the App-threading +
cycle-key + shared-persistence pattern end-to-end before the larger ones.
* Thread the shared Theme through the volume file browser
Same pattern as the scale TUI: module-level Color constants become
theme.<field> lookups, and 't' cycles the shared theme. Exercises the new
`danger` field for the first time — delete confirmations now read as
destructive (danger) while upload/download overwrite confirmations stay
cautionary (pending/warning), a distinction the old single Yellow border
couldn't make.
* Thread the shared Theme through the metrics TUI
Largest of the remaining screens, and the reason Theme grows a `series: &[Color]`
field: metrics charts draw several simultaneous lines (CPU/memory limits aside,
which reuse `accent`) — egress/ingress, p50/p90/p95/p95/p99, 2xx/3xx/4xx/5xx —
that must stay visually distinct from each other, which the existing semantic
slots (accent/running/pending/danger) don't have enough of on their own. Each
theme now defines an ordered 4-color palette picked by index instead of the
metrics TUI naming a `CPU_COLOR`/`P90_COLOR`/`STATUS_5XX` constant per metric.
Both MetricsApp (single service) and ProjectApp (--all) get their own theme
field, cycled with 'v' (metrics already uses 't' for the time range).
* Thread the shared Theme through the templates picker
The picker keeps its independent light/dark terminal-background detection
(TerminalTheme, detect_terminal_theme, OSC background query) — that answers
a different question ("is the user's terminal light or dark") than "which
brand theme is selected," and stays exactly as it was. The PickerApp field
that held it is renamed `bg` so it doesn't collide with the new `theme`
field carrying the shared brand Theme.
Only the picker's few brand-colored spots (the search input's border, the
loading spinner, the verified badge, health-score coloring) move to the
theme. Since every plain key types into the search box here, cycling the
theme uses Ctrl+T instead of the bare letter the other screens use.
* Thread the shared Theme through the dev/run TUI
Last of the six screens. Leaves the colored::Color -> ratatui::Color
passthrough mapper (convert_color) untouched — it renders subprocess log
output colors, not chrome, and was never in scope for this pass. Everything
else (tab bar, log selection highlight, info pane, help bar) now reads from
the shared theme, cycled with 't' like the scale TUI and volume browser.
* chore: Release railwayapp version 5.57.0
* fix(code): fall back to installed client when the latest check is rate-limited (#1205)
* fix(code): fall back to installed client when the latest check is rate-limited
`railway code --railway` resolves the newest railway-agent-tui on every launch
by calling the unauthenticated GitHub API (api.github.com/.../agent-releases/
releases/latest), which is capped at 60 requests/hr per IP. When that quota is
exhausted — a burst of launches, a shared NAT, other tooling on the same IP —
the call returns 403 and `.error_for_status()` aborted the whole command, even
though a perfectly good client was already installed under
~/.railway/runtimes/railway-tui/<version>/.
Scope a fallback to the release-resolution step: when fetching the latest
release metadata fails (rate limit, network, offline), use the newest installed
client that still runs instead of blocking the launch. The fallback is
deliberately limited to this network step — a later failure such as a download
checksum mismatch is a real integrity problem and still surfaces, never papered
over with an older binary.
Adds tests for the rate-limited fallback (picks the newest runnable installed
version, ignores non-version and empty dirs) and for the no-installed-client
case (still errors, with a clear message). The existing corrupt-download test
continues to assert a hard failure.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* feat(code): retry the latest check with a GitHub token only when rate-limited
Extends the installed-client fallback: when the anonymous latest-release check
hits GitHub's rate limit (403/429 with x-ratelimit-remaining: 0), retry once
authenticated before giving up, raising the cap from 60/hr per IP to 5000/hr for
the token's account.
The token is used ONLY to get past the rate limit — never attached to the normal
request — so ordinary launches stay anonymous. It is resolved from GH_TOKEN, then
GITHUB_TOKEN (matching the gh CLI's own precedence), then the signed-in gh CLI
(`gh auth token`), best-effort and skipped silently when gh isn't installed or
logged in. With no token, behaviour is unchanged: the error propagates and the
newest installed client is used.
Adds a GH_TOKEN > GITHUB_TOKEN > blank-falls-through precedence test; the
rate-limit integration tests now serve GitHub's 403 + x-ratelimit-remaining: 0
shape so the outcome is identical whether or not the test host has a token.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* chore: Release railwayapp version 5.57.1
* Create fresh Railway agents by default and add new-VM project selection (#1203)
* Default code to fresh Railway agents and add creation project selector
* Keep new-agent picker spacing stable across focus changes
* Make Option+N configure and create a fresh cloud agent
* docs(code): explain client update fallback (#1207)
* chore: Release railwayapp version 5.57.2
* fix(database): authenticate the generic HA switchover against the node's health server (#1177)
* fix(database): send the node's HEALTH_API_PASSWORD on the generic HA switchover
Clusters on the generic node contract (redis-ha, mysql-ha, mongo-ha) gate
POST /switchover with HTTP Basic once the node carries HEALTH_API_PASSWORD.
The switchover command now resolves that credential inside the target
container (HEALTH_API_USERNAME defaults to railway) and hands curl
-u user:pass, or nothing when the node has no password, so one command
spans a cluster mid-rollout and no secret enters the exec payload or a log.
* test(cluster_probe): run the sh-shim switchover tests on unix only
The emitted command only ever runs inside a Linux container; Windows has no
sh to check it against, so the shim-backed tests are gated like the Patroni
ones. The prelude-shape test still runs everywhere.
* fix(database): keep the HA switchover credential out of curl's argv (#1180)
The generic HA switchover authenticates against the node's health
server inside its container, and #1177 had the credential reach curl as
`-u "$HEALTH_API_USER:$HEALTH_API_PW"`. Argv is public inside the
container: /proc/<pid>/cmdline (ps) shows every process's arguments to
every other process in the PID namespace for as long as the request
runs, and under the ONE PASSWORD design that argument is the engine's
root password.
Hand curl the credential as a one-line config document on stdin instead
(`user = "user:pass"` piped into `curl -K -`), escaped for curl's config
parser by a pure-POSIX shell function (`\` and `"` backslashed, newline
to `\n`). printf is a builtin of the shells the data images ship (dash,
bash), so no process argv carries the secret at any point. Twin of the
Patroni prelude's change.
HEALTH_API_PASSWORD/HEALTH_API_USERNAME resolution, the bare request for
an open node, and the error surfacing of the node's status and body are
unchanged. The sh-shim tests now capture curl's stdin as well as its
argv and assert the password is in the document and in no argument; a
new test round-trips a password with every character curl's parser
treats specially.
* fix(postgres): keep the Patroni switchover credential out of curl's argv (#1179)
The switchover authenticates against Patroni's REST API inside the
member's container, and since #1171 the credential reached curl as
`-u "$PATRONI_REST_USER:$PATRONI_REST_PW"`. Argv is public inside the
container: /proc/<pid>/cmdline (ps) shows every process's arguments to
every other process in the PID namespace for as long as the request
runs, and under the ONE PASSWORD design that argument is the superuser
password.
Hand curl the credential as a one-line config document on stdin instead
(`user = "user:pass"` piped into `curl -K -`), escaped for curl's config
parser by a pure-POSIX shell function (`\` and `"` backslashed, newline
to `\n`). printf is a builtin of the shells the data images ship (dash,
bash), so no process argv carries the secret at any point.
Resolution precedence, the bare request for a member with no password,
and the error surfacing of Patroni's status and body are unchanged. The
sh-shim tests now capture curl's stdin as well as its argv and assert
the password is in the document and in no argument; a new test round-
trips a password with every character curl's parser treats specially.
* chore: Release railwayapp version 5.57.3
* chore: Release railwayapp version 5.57.4
* fix(postgres): wait out Patroni's switchover poll before declaring failure (#1208)
Patroni holds POST /switchover open for up to ~20s while it polls the
failover result. The CLI's 8s curl budget cut that short, so a completed
handoff still surfaced as exit 28 / HTTP_STATUS:000. Raise the budget
above Patroni's window and point operators at `ha status` when a wait
still times out.
Co-authored-by: Cursor <cursoragent@cursor.com>
* chore: Release railwayapp version 5.57.5
* rendergraphasrailway now seeds the resource ident pool (names) with… (#1209)
Task: Seed pull codegen ident pool with source aliases
Attempt 2 of 3
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
* railway config migrate treated platform configFile values like… (#1210)
Task: Normalize leading slash in migrate configFile paths
Attempt 3 of 3
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
* Set ok=false and exit non-zero on failed apply (#1211)
* src/iac/engine.rs: after applychangeset, ok is now recomputed via…
Task: Set ok=false and exit non-zero on failed apply
Attempt 6 of 6
* The local branch had reset to master, so I restored the PR #1211…
Task: Set ok=false and exit non-zero on failed apply
Attempt 7 of 9
Follow-up: The operator re-ran this task. Review the current state of the branch, address anything outstanding, make the gate pass, then call report again.
---------
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
* Classify imported databases by provenance, match by name (#1212)
* Imported services are now classified as database only when the…
Task: Classify imported databases by provenance, match by name
Attempt 2 of 3
* Fixed the remaining defect in diffgraphs (src/iac/changeset.rs): a…
Task: Classify imported databases by provenance, match by name
Attempt 5 of 9
Follow-up: The operator re-ran this task. Review the current state of the branch, address anything outstanding, make the gate pass, then call report again.
* The Installer (POSIX sh) failure is install.sh hitting GitHub's…
Task: Classify imported databases by provenance, match by name
Attempt 7 of 9
---------
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
* Replaced the "Last resort" comment above the named-partial export in… (#1213)
Task: Reword migrate's 'last resort' partial comment
Attempt 3 of 3
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
* chore: Release railwayapp version 5.57.6
* fix(code): upgrade cloud OpenCode V2 servers to match local clients (#1215)
* chore: Release railwayapp version 5.57.7
* fix(iac): validate regions before creating buckets (#1206)
Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com>
Co-authored-by: Victor Ramirez <victor@railway.com>
* fix(IaC): only match an image as a database if it matches a known name (#1214)
* chore: Release railwayapp version 5.57.8
* fix(code): unify agent startup and project selection (#1216)
* fix(code): unify agent startup and project selection
* fix(code): trust Grok cloud-agent sessions by default
* fix(ca): consolidate discovery into the global refresh shortcut
* chore: Release railwayapp version 5.57.9
* feat(ca): offer to wake and resume an agent slept under a local client (#1219)
An agent slept from outside the TUI (an idle timer, `railway ca sleep`
elsewhere) takes its Codex or OpenCode server with it. The local client
keeps retrying an address that no longer answers: the code domain does
not wake a slept VM, and a woken one has no server until `reconnect`
respawns it.
On each agent refresh, an agent that went from awake to sleeping while
this TUI holds a live local client pane on it raises the confirm modal:
wake it and resume? A yes wakes the agent and, once the wake lands,
respawns the harness server on the saved port with the saved
credentials, so the client's own retries succeed with no restart.
Sleeps this TUI asked for are not offered back; offers are capped at
two per agent per run.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* chore: Release railwayapp version 5.57.10
* fix(code): repair OpenCode V2 readiness and saved history fallback (#1220)
* chore: Release railwayapp version 5.57.11
* fix(up): show build errors and failure reasons (#1222)
* fix(up): show build errors and failure reasons
* fix(up): clean up deployment links
* chore: Release railwayapp version 5.57.12
* feat(code): make stable OpenCode V2 the default (#1221)
* feat(code): default new OpenCode launches to stable V2
* fix(code): use recorded OpenCode protocol when respawning after wake
* fix(code): use OpenCode labels and oc agent names
* chore: Release railwayapp version 5.58.0
* feat: add IaC partial ownership commands (#1226)
* chore: Release railwayapp version 5.59.0
* feat(code): cut OpenCode over to V2 and retire the opencode2 shim (#1229)
* feat(code): cut OpenCode over to V2 and retire the opencode2 shim
OpenCode V2 is stable and the cloud-agent image's `opencode` is being hard
cut over to it, so the CLI now launches the image's native `opencode`
directly and drops the Railway-managed `~/.local/bin/opencode2` shim that
downloaded `@opencode/cli` from npm at launch.
- Fold Agent::OpenCode2 into Agent::OpenCode (slug/binary `opencode`).
The V1 "OpenCode 1 - legacy" agent and its auth.json seeding are gone;
V1 remains only as Protocol::V1 for reconnecting/upgrading managed
servers started before the cutover.
- Keep `--opencode2` as a hidden, deprecated alias of `--opencode` (with
a stderr note); normalize the `opencode2` slug to `opencode` wherever it
is read from prefs, panes, saved connections, VM history or backboard.
Protocol::harness() now writes `opencode` for V2 connections.
- Move opencode2/auth.rs to opencode/auth.rs and add
opencode/import_auth.py, which has V2 create/migrate its own store
(opencode api --standalone GET /api/info) and merges staged provider
credentials into `credential` without overwriting remote accounts,
backing up V1-only storage first.
- Runtime seed: verify `opencode --version` is 2.x; on an older image run
the official installer (https://opencode.ai/v2/install, fetched first,
stdin closed) which upgrades ~/.opencode/bin/opencode in place, then
remove the old shim. The Desktop/server bootstrap does the same and
upgrades when the image is behind a saved server's release. Resuming a
saved conversation on a 1.x VM explains how to upgrade instead.
- Put ~/.opencode/bin first on HARNESS_PATH; update README and docs.
Tests: drop tests/opencode2.py, retarget the credential import test,
rework tests/opencode_desktop.py around in-place installs with a stubbed
installer, add seed/alias/history coverage.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* refactor(code): drop OpenCode 1 support and the in-VM V2 auto-upgrade
The cutover is a hard one: the image's `opencode` is V2 and the CLI runs
it directly. Remove everything that existed only to run, reconnect, or
migrate OpenCode 1, and never install or upgrade OpenCode on a VM.
- Remove the in-place V2 install from the runtime seed and the server
bootstrap. A VM whose `opencode --version` is not 2.x fails fast with:
"This cloud agent is on an older image whose OpenCode is not V2.
Create a new agent with railway code --opencode --new." The resume
guard prints the same message.
- Delete Protocol::V1 (and the Protocol enum), legacy_runtime, the V1
reconnect and `upgrade` client action, V1 configure_permissions, the
V1 local-client npm install, V1 API paths/event shapes in the bridge
and client sessions, the V1 auth.json credential conversion, and the
V1 storage backup in import_auth.py. V2 credential import stays.
- Remove the --opencode2 flag from `railway code` and `railway ca
desktop`.
- Keep read-side normalization of stored `opencode2` -> `opencode`
(prefs, saved server.json, snapshots, history, panes). A saved
OpenCode 1 connection record is dropped on read so the rest of the
snapshot archive stays usable; the VM bootstrap refuses to reconnect
a server started as OpenCode 1 but still stops it.
- Docs and tests follow: V2-only bootstrap tests, old-image and V1
server refusals, V1 snapshot records rendered as plain SSH.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* chore: Release railwayapp version 5.59.1
* fix: correct OpenCode label in agent picker (#1228)
* chore: Release railwayapp version 5.59.2
* feat: add `railway trace` for tracing settings and traces (#1230)
`railway trace enable|disable|inherit` flip a service's tracing override, or
the project default with `--project-default`, through `serviceUpdate` and
`projectUpdate`. `status`, `list` and `get` read the settings, the per-service
span activity, the trace list and one trace's spans; `list --json` and
`get --json` print one object per line like `railway logs --json`.
* chore: Release railwayapp version 5.60.0
* feat(database): add MySQL point-in-time recovery (#1144)
* feat(database): add MySQL point-in-time recovery
MySQL PITR is binlog archiving into a Railway bucket, enabled by the
`mysql-pitr` composable overlay. The generic `pitr` tree already drives
everything an engine's archive needs from its declared contract, so this is
mostly the declaration plus the two ways MySQL genuinely differs from Postgres.
Both differences become declarations rather than branches on the engine name:
- `supports_ha: false`. The image's archiver refuses to run whenever the
Group Replication seed list is set, so MySQL PITR is standalone-only.
Every path that would take (or follow) the rolling HA workflow now checks
this first -- `progress`/`cancel`/`clear` refuse before any network call,
and `enable`/`disable` refuse before the HA mutation -- so the user gets
one clear sentence instead of a server error from a workflow that was
never going to start.
- `probe_kind: None`. The live coverage probe in `pitr status` is pgBackRest
shelling into the container, which MySQL has no equivalent tool for.
Selecting the probe by declared kind means MySQL renders no coverage
section at all, rather than a probe run with another engine's tooling or a
fake "unavailable" for something it simply does not implement.
The archive variable contract needs no code: `BINLOG_ARCHIVE_` as the declared
prefix is enough for enabled-state detection and the enable overlay, and the
image-eligibility rules ship with the `mysql-pitr` template record. Notably
that template does NOT require a floating major tag, unlike postgres-pitr --
its image matrix only publishes exact minors, so a hardcoded "minor pins are
bad" rule would have refused every MySQL image.
Restore, backups and schedules ride the volume-instance mutations, which are
engine-agnostic server-side and needed no change.
* fix(database): MySQL PITR rolls HA clusters too
mysql-ha archives from whichever member is the writable primary and
backboard's enable-pitr-ha dispatches the Group Replication choreography
on the registry's haRolloutKind (railwayapp/mono#38063, mysql-ha#45), so
the standalone-only refusal on MySQL cluster roots is stale: enable/disable
on a cluster root and progress/cancel/clear drive the rolling workflow, the
same as Postgres. The gate stays for engines that declare no rolling
workflow; the test now exercises it through a synthetic declaration.
* fix(database): update the stale standalone-only assumption CI caught
capability_declarations_match_what_ships still asserted MySQL PITR's
supports_ha as false, which the previous commit had already flipped to
true -- the test was written for the old invariant and never updated
alongside it, so CI failed on the assertion.
Also fixes the same stale assumption in `railway mysql`'s module doc
comment and --help text, which still told users PITR is
standalone-only on MySQL and cannot be enabled on an HA cluster.
* chore: Release railwayapp version 5.61.0
* fix: avoid GitHub API rate limits during CLI installation (#1232)
* fix(herdr): retry failed bootstrap launches and profile removals
* ci: validate pull requests targeting stacked branches
---------
Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com>
Co-authored-by: Noah <hi@nd.mt>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com>
Co-authored-by: Jake Runzer <jakerunzer@gmail.com>
Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com>
Co-authored-by: Arnav Gupta <championswimmer@gmail.com>
Co-authored-by: Arnav Gupta <arnav@railway.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Paulo Cabral Sanz <paulo@railway.app>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Víctor Ramírez <hey@futurepastori.com>
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
Co-authored-by: Fandu <113630375+mrfandu1@users.noreply.github.com>
Co-authored-by: Victor Ramirez <victor@railway.com>
Co-authored-by: Nebula <nebula@railway.app>
Co-authored-by: Brody Over <10548119+brody192@users.noreply.github.com>
Co-authored-by: Luca Casonato <hello@lcas.dev>
---------
Signed-off-by: Fandu <113630375+mrfandu1@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: Cody De Arkland <cody@railway.com>
Co-authored-by: Noah <hi@nd.mt>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Cody De Arkland <17350652+codyde@users.noreply.github.com>
Co-authored-by: Jake Runzer <jakerunzer@gmail.com>
Co-authored-by: Cody De Arkland <codyde@users.noreply.github.com>
Co-authored-by: Arnav Gupta <championswimmer@gmail.com>
Co-authored-by: Arnav Gupta <arnav@railway.com>
Co-authored-by: Paulo Cabral Sanz <paulo@railway.app>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Víctor Ramírez <hey@futurepastori.com>
Co-authored-by: Víctor Ramírez <4147268+futurepastori@users.noreply.github.com>
Co-authored-by: Fandu <113630375+mrfandu1@users.noreply.github.com>
Co-authored-by: Victor Ramirez <victor@railway.com>
Co-authored-by: Nebula <nebula@railway.app>
Co-authored-by: Brody Over <10548119+brody192@users.noreply.github.com>
Co-authored-by: Luca Casonato <hello@lcas.dev>
Co-authored-by: Cody De Arkland <codydearkland@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
railway <engine> ha switchoverfor clusters on the generic node contract (redis-ha, mysql-ha, mongo-ha) now authenticates against the node's health server, which gatesPOST /switchoverwith HTTP Basic once the node carriesHEALTH_API_PASSWORD(railwayapp-templates/redis-ha#63, railwayapp-templates/mysql-ha#51, railwayapp-templates/mongo-ha#2).Changes
controllers/cluster_probe.rs: the switchover command resolves the credential inside the target container (HEALTH_API_PASSWORD,HEALTH_API_USERNAMEdefaulting torailway) and hands curl-u user:pass, or nothing when the node has no password — an open node ignores the header and an enforcing one requires it, so one command spans a cluster mid-rollout and no secret enters the exec payload, this process, or a log line. Same shape as the Patroni switchover's credential prelude (fix(postgres): authenticate the CLI's switchover against Patroni's REST API #1171).shwith acurlshim: no password → no-uat all; password →railway:<pw>; an explicit username wins; the prelude text is pinned.Verification
cargo fmt --all -- --checkcargo test cluster_probecargo clippy --all-targets --all-features