Repository navigation
feat(gaia): add developer-only coding-agent MCP handoffs - #3911
Conversation
Request changesThis adds an opt-in developer mode that lets a GAIA user hand a reported problem to their own Claude Code or Codex app over a local, consent-gated MCP bridge, with managed git worktrees for the fix. The consent design is the strongest part of it: sharing can't be remembered, can't be auto-approved in bypass mode, and revocation really does cut off an already-connected client. I found no runtime bugs — the two things to fix are cheap. A new doc file describes a capability this PR doesn't build. A second copy of the engineering skill was added under the docs tree. It isn't the one the agent loads, nothing links to it, and its text tells the model it can offer "bounded live sharing" with a grant that keeps covering later events. Snapshots that expire are the only thing implemented — the spec page in this very PR lists live grants as a future extension. Delete the stray copy, or point the docs at the real one. The new CLI test runs the installed Five smaller suggestions are in the technical details — a worktree branch named after Codex even on the Claude path, a credential-redaction gap for JSON-shaped secrets, and a few convention/doc-duplication items. Real-world evidenceNo
The verdict rests on static review plus those two artifacts. Nothing I saw contradicts the change working; neither finding above is disproved or supported by the evidence. 🔍 Technical details🟡 Important1. Orphan skill copy contradicts both the shipped skill and this PR's own spec ( Two files now carry
The docs copy instructs: "Ask permission for a snapshot or bounded live sharing of this task. A live grant covers later events only within its approved categories, recipient and lifetime."
2. [str(Path(sys.executable).parent / "gaia"), "engineering", *args],The root with the repo's 🟢 Minor3. Worktree branch says
4. Credential redaction misses JSON-shaped secrets ( The 5. Cache-path branch compares a resolved path to an unresolved one (
6. Every other package under 7. The same 10-line paragraph was pasted into three npm docs ( Byte-identical text in all three. CLAUDE.md gives each a distinct job — README is integrator-facing, SPEC is the technical reference, SKILL is the AI-assistant playbook. Worth differentiating: SPEC should carry the flag/env contract and the frozen-binary limitation, SKILL should say what an assistant may and may not claim about a handoff. (Not counted as a nit — one pattern note: evidence appends are bounded in practice to roughly 15 snapshots by Strengths
|
…ring-mcp # Conflicts: # hub/agents/gaia/npm/CHANGELOG.md # hub/agents/gaia/python/gaia_agent/stdio.py # src/gaia/apps/webui/src/components/ChatView.tsx # src/gaia/apps/webui/src/components/PermissionPrompt.tsx # src/gaia/apps/webui/src/stores/notificationStore.ts
…ering CLI GaiaAgent listed EngineeringToolsMixin ahead of ProjectMapMixin, which broke the test that pins ProjectMapMixin as the first base so Agent's no-op task-start hook can't shadow it. The two mixins share no methods, so the engineering mixin now sits second. The engineering CLI test ran the installed `gaia` script, which on a machine with several worktrees imports whichever one was pip-installed last. It now runs `python -m gaia.cli` with this checkout's src on PYTHONPATH, like the other CLI subprocess tests. tests/unit/engineering gains the __init__.py every sibling test package has, so its test_cli/test_store modules can't collide with the same basenames under tests/unit/connectors. Removes docs/spec/skills/gaia-harness-engineering/SKILL.md: nothing loaded or linked it, and it described live sharing grants that aren't implemented.
…s for GAIA - Redaction missed `"api_key": "..."`, the shape most logs and config dumps take, because the key had to be followed directly by `:` or `=`. An optional closing quote now matches it too. - Worktree branches were named `codex/engineering-<job>` even when Claude Code did the work; they are now `gaia/engineering-<job>`. - The default-profile check compared a resolved root against an unresolved home path, so a symlinked ~/.gaia put worktrees under the wrong cache directory. Both sides are resolved now. - The context cursor check uses isinstance (pylint C0123) and still rejects bools. - The npm README, SPEC and SKILL carried one pasted paragraph. README now says what the mode is for, SPEC carries the flag, tool, consent and packaging contract, and SKILL says what an assistant may and may not claim about a handoff.
The spec linked the draft skill copy under docs/spec/skills, which was removed as an orphan. It now names the shipped location, which the host loads only in developer mode.
|
The branch now merges cleanly with current
Still open: the maintainer pilot, the agent eval, and the note about the evidence-size ceiling, which the review didn't count as an item. 🔍 Technical detailsConflicts (merge 0c45842)
Other
Results
|
|
Merged 🔍 Technical detailsThe failing step was The merge was clean. The same stale-base failure hit #3687, #3911, #3982 and #3616 identically. |
# Conflicts: # hub/agents/gaia/python/gaia_agent/agent.py
|
Not merging this unilaterally, for a reason the PR itself states: the body asks "@kovtcharov-amd: ready for architecture and manual pilot review", and the test plan's "Maintainer manual TUI/candidate-worktree pilot and human review" box is unchecked. That is an author-designated human gate, and an agent squash-merging past it would defeat the point of asking. It is also the right gate to have. This feature hands an approved snapshot of a user's source tree to a third-party coding agent (Codex / Claude Code) over a local stdio MCP server. The boundary being crossed is data egress, and the controls — recipient-binding, snapshot expiry, code approval tied to the displayed diagnosis revision, no "always allow" — are exactly the kind of thing that needs a person to exercise once rather than read about. All seven review findings are addressed — I checked each against the head commits, not just the author's summary:
Nothing is outstanding from review. What a human needs to do before merge: run the pilot, and decide whether the two acknowledged gaps ship as-is — preview metadata and results are reported by the coding app, not verified by GAIA, and there is no evidence-size ceiling. 🔍 Technical detailsLayering reads clean. The merge-conflict resolution that matters. Merge state: CI: one real failure, Eval: attempted, but the Claude-based runner hit its weekly quota. Codex end-to-end passed; live Claude inference is unverified. Since this adds tools to the flagship's surface, the eval is in scope under CLAUDE.md and should run before merge. |
# Conflicts: # docs/docs.json # hub/agents/gaia/npm/CHANGELOG.md # setup.py # src/gaia/agents/base/tools.py # src/gaia/cli.py # tests/unit/test_tool_decorator.py # tui/internal/cli/root.go
kovtcharov-amd
left a comment
There was a problem hiding this comment.
Rebased the branch onto current main and pushed — conflicts had grown since the last review pass because main kept moving (7 files now, up from the 5 noted a couple of days ago, including a real clash where two different features added a keyword argument to the same function signature on the same line). Both features are preserved in the merge; local tests, lint, and the Go build all pass, and fresh CI is running against the merged branch.
The one new thing I found: the shared context this feature writes to disk is only protected on Linux/macOS. On Windows, the file-permission lockdown it uses doesn't actually restrict other local accounts from reading it — this repo already hit and fixed the identical bug for the daemon's launch secret, but this feature doesn't reuse that fix. Worth closing before this reaches general availability, since confidentiality of the shared snapshot is half of the feature's pitch.
Everything else already raised in this thread — the developer-mode gate, the consent model, the subsystem layering — checks out on an independent read. Not requesting changes; the maintainer's already-stated plan (pilot + eval before merge) stands.
🔍 Technical details
Merge: pushed c2b6297ea to kovtcharov/gaia:codex/harness-engineering-mcp (merge of origin/main). Conflicts resolved, all additive — no contradictory logic:
docs/docs.json,hub/agents/gaia/npm/CHANGELOG.md,setup.py: both sides' entries kept.src/gaia/agents/base/tools.py:@tool()keeps bothpreflight(main) andregistry(this PR) kwargs;_SUPPORTED_TOOL_KWARGSlists both.src/gaia/cli.py: both theengineeringdispatch branch and theuse_chatgptremoved-provider guard kept.tests/unit/test_tool_decorator.py: expected-error message updated to list both kwargs.tui/internal/cli/root.go:developerModeandfullAccessFlagvars coexist; no leftover references to the pre-rename name.
Verified locally: pytest tests/unit/test_tool_decorator.py tests/unit/engineering tests/mcp/test_engineering_mcp.py (41 passed, 1 pre-existing platform gap — see below); pytest tests/unit/engineering/test_cli.py tests/unit/test_amd_gaia_urls.py tests/unit/test_starter_skills.py (208 passed, 15 skipped); black/isort/targeted pylint clean on touched files; python util/lint.py --tool-descriptions clean (88 flagship schemas within budget); go build ./... and go vet ./internal/cli/... clean in tui/.
Windows ACL gap: private_directory() (src/gaia/engineering/store.py:41-42) does path.mkdir(mode=0o700) / path.chmod(0o700) only. tests/unit/engineering/test_store.py::test_atomic_concurrent_feedback_preserves_all_updates asserts st_mode & 0o777 == 0o600 with no platform skip, and fails as-is on Windows in this environment — but tests/unit/ only runs on ubuntu-latest/macos-latest in CI (.github/workflows/test_unit.yml), so this has never been exercised on Windows there. The fix pattern already exists in-repo: _lock_down_windows_acl in src/gaia/daemon/sidecars/manager.py:115 builds a real NTFS DACL restricted to the current user, with a docstring citing issue #2250 for exactly this "POSIX mode bits are a no-op on Windows" failure mode. It's currently private to the sidecar manager; worth factoring out (or replicating) for gaia.engineering.store.
CI: the previous "Test GAIA CLI on Windows (Full Integration)" failure was HttpError: API rate limit exceeded for installation fetching the FedericoCarboni/setup-ffmpeg@v3 action, not a test failure — confirmed from the job log. That workflow (test_gaia_cli_windows.yml) doesn't run tests/unit/engineering/ at all (its own comment notes tests/unit/ is Ubuntu-only), so it wouldn't have caught the ACL gap above either way.
Gate verification: EngineeringService.__init__ (src/gaia/engineering/service.py:29-36) raises PermissionError unless developer_mode=True (passed by whoever launched gaia mcp engineering --developer-mode ...) or GAIA_DEVELOPER_MODE=1 was already set in that process's environment before any MCP client connected. The stdio protocol gives the connecting Claude Code / Codex client no channel to set this itself, so the gate can't be spoofed client-side.
Layering: src/gaia/engineering/ doesn't touch src/gaia/connectors/, which is correct — connectors govern GAIA-as-client grants to external services (Google/GitHub OAuth, MCP servers); here GAIA is the MCP server a local coding app connects to, an orthogonal direction. No inconsistency with the existing grant model.
…xclusion test_cli_docs_drift.py (added on main since this branch forked) now walks the parser to the leaf and requires each gaia engineering subcommand, and the Global Options exclusion list, to name the literal command.
|
Pushed one more fix after that review: One correction to my last comment: I said |
private_directory() only did POSIX chmod(0o700), a no-op on NTFS, so a shared engineering snapshot or grant was readable by any other local account on a shared Windows box. Reuses the same DACL-lockdown pattern already shipped for the daemon's launch secret (amd#2250).
Environment variables like DB_PASSWORD or AWS_SECRET_ACCESS_KEY carry a prefix, and the snapshot scrubber required a word boundary before the credential keyword, so those values were shipped in the clear to the third-party coding agent. Prefixed and suffixed names now match, and auth tokens and passwd are covered too.
Prose like "secretary: alice" and code like parser.add_argument("--api-key") are left untouched.
|
Ran the maintainer pilot that was still unchecked on the test plan — the full host flow plus a real stdio MCP round-trip from a simulated coding app. The handoff works and the consent model holds: developer mode cannot be switched on by an ambient environment variable, a bad pairing credential stops the server from starting at all, forged job and evidence ids are refused, One thing needed fixing before this lands, now pushed as 🔍 Technical detailsCause. Measured on the pre-fix branch, against the shipped function: After the fix: 11/11 of those shapes redacted, and 0/6 false positives across What the coding app actually received over MCP, before and after — same input file, real Pilot transcript.
Tests. Noticed, not fixed here — |
|
CI triage for the three red lanes, now resolved. Lint was ours and is fixed in Windows and macOS smoke are pre-existing on 🔍 Technical details
So the lint failure was genuinely introduced here, not inherited. Smoke lanes — and on this PR (run 36579718567, head ccdcdf8): Same test, same assertion, same counts. Note that main's |
Developers can now report friction while using GAIA and hand an explicitly approved snapshot to their existing Codex or Claude Code app. Developer mode loads the skill and prepares a cached source repository; the coding app diagnoses the problem, requests code scope, and iterates in a dedicated worktree with separately launched previews. Ordinary sessions do not expose this capability.
The usage guide and architecture scope are included. Autonomous fixes are a later phase; stable auto-update is tracked in #3898. @kovtcharov-amd: ready for architecture and manual pilot review.
Test plan
🔍 Technical details
The bridge provides recipient-bound, expiring snapshots through an explicitly enabled local stdio server. Code approval is tied to the displayed diagnosis revision. Preview metadata and results are reported by the coding app, not supervised or independently verified by GAIA. Packaged build provenance, live grants, automatic trace capture and automatic contribution orchestration remain future work.
Run
gaia-tui --developer-mode --dev, then follow the guide.Synthetic live WebUI consent evidence: