feat(ws1): non-Claude in-process invokers + FastAPI HTTP API (WS1, code only) #44

Merged
amoussa1229 merged 3 commits from feat/ws1-inprocess-models-http-api into main 2026-06-23 19:48:19 +00:00
amoussa1229 commented 2026-06-23 01:20:30 +00:00 (Migrated from github.com)

Summary

  • agent_team/invoker_multi.py (new): unified in-process invoker for non-Claude models; multi_invoke(prompt, *, model) dispatches to GPT-4.1 (cross_reviewer), DeepSeek (fast_coder), or Gemini (scanner) via the orchestrator's models.py; bind_multi_invoker() wires the review loop seam; all model imports are lazy (no import-time side effects)
  • nodes/review_loop_llm.py: default_plan_reviewer swapped from make_run_py_invoker() (subprocess) to make_cross_reviewer_invoker() (in-process GPT-4.1, WS1); subprocess path retained as make_run_py_invoker() for opt-in use
  • nodes/builders_llm.py: new make_fast_coder_invoker() in-process DeepSeek path; old subprocess path renamed to subprocess_build() (opt-in fallback); default_build() now delegates to make_fast_coder_invoker()
  • agent_team/api.py (new): FastAPI HTTP API with bearer-token auth (AGENT_TEAM_API_TOKEN env, constant-time comparison); endpoints POST /tasks, GET /tasks/{thread_id}, POST /orchestrator/invoke; default bind host 127.0.0.1; serve() helper for attended deploy; code only — NOT started
  • run-team.py: _cmd_serve calls bind_multi_invoker() alongside bind_subscription_invoker() so non-Claude seams go live at daemon startup
  • tests/test_ws1_invoker_multi_api.py (new): 21 tests; tests/test_builders_llm.py: updated 2 tests to reflect renamed subprocess path

CI status

ruff check clean · pytest -q — 1065 passed

Security notes (inline adversarial review)

  • Token auth: _get_token() raises RuntimeError on missing/empty AGENT_TEAM_API_TOKEN so the server cannot start without a secret. Token comparison uses hmac.compare_digest (constant-time).
  • Host binding: default 127.0.0.1 hard-coded in the make_app signature; the serve() call also sets it explicitly, so a mis-typed call to serve(host="0.0.0.0") is the only escape path and requires a deliberate code change.
  • Subprocess in /orchestrator/invoke: uses list-form argv ([sys.executable, str(run_py), request.prompt]) — no shell=True, so request.prompt cannot be shell-interpolated. However, the prompt text IS passed as a subprocess argument; a hostile prompt could attempt argument injection if run.py does further shell expansion. Mitigation: run.py receives the prompt as sys.argv[1] (a single string, not shell-parsed). Flag for GPT-4.1 cross-review.
  • No secrets in code: AGENT_TEAM_API_TOKEN read from env at request time, never hardcoded or logged.
  • invoker_multi.py: no authentication or authz on the in-process model calls — they run with the same credentials as the daemon process. Acceptable for the 127.0.0.1-bound, VPN-only surface.

OUTSTANDING (attended — do NOT merge until done)

  • GPT-4.1 cross-family review — mandatory before merge; no Claude self-review counts. Specifically review the orchestrator/invoke subprocess argument-passing and the bearer-token timing guarantee.
  • /sh-security-review — formal security sign-off on the new HTTP API surface (bearer-token auth, 127.0.0.1 bind, subprocess prompt-passing).
  • Live deploy steps (attended, on R720 box):
    • pip install fastapi uvicorn on the box (or add to the box's requirements)
    • Set AGENT_TEAM_API_TOKEN in ~/secrev.env (strong random token, never committed)
    • Restart the orchestrator systemd unit
    • Smoke-test: curl -H "Authorization: Bearer <token>" http://127.0.0.1:8765/tasks — expect 405 (method not allowed), confirming the API is up
    • Verify VPN-only access (the API must NOT be reachable from outside the VPN interface)

Generated by Claude Code

## Summary - **`agent_team/invoker_multi.py`** (new): unified in-process invoker for non-Claude models; `multi_invoke(prompt, *, model)` dispatches to GPT-4.1 (`cross_reviewer`), DeepSeek (`fast_coder`), or Gemini (`scanner`) via the orchestrator's `models.py`; `bind_multi_invoker()` wires the review loop seam; all model imports are lazy (no import-time side effects) - **`nodes/review_loop_llm.py`**: `default_plan_reviewer` swapped from `make_run_py_invoker()` (subprocess) to `make_cross_reviewer_invoker()` (in-process GPT-4.1, WS1); subprocess path retained as `make_run_py_invoker()` for opt-in use - **`nodes/builders_llm.py`**: new `make_fast_coder_invoker()` in-process DeepSeek path; old subprocess path renamed to `subprocess_build()` (opt-in fallback); `default_build()` now delegates to `make_fast_coder_invoker()` - **`agent_team/api.py`** (new): FastAPI HTTP API with bearer-token auth (`AGENT_TEAM_API_TOKEN` env, constant-time comparison); endpoints `POST /tasks`, `GET /tasks/{thread_id}`, `POST /orchestrator/invoke`; default bind host `127.0.0.1`; `serve()` helper for attended deploy; code only — NOT started - **`run-team.py`**: `_cmd_serve` calls `bind_multi_invoker()` alongside `bind_subscription_invoker()` so non-Claude seams go live at daemon startup - **`tests/test_ws1_invoker_multi_api.py`** (new): 21 tests; **`tests/test_builders_llm.py`**: updated 2 tests to reflect renamed subprocess path ## CI status `ruff check` clean · `pytest -q` — **1065 passed** ## Security notes (inline adversarial review) - **Token auth**: `_get_token()` raises `RuntimeError` on missing/empty `AGENT_TEAM_API_TOKEN` so the server cannot start without a secret. Token comparison uses `hmac.compare_digest` (constant-time). - **Host binding**: default `127.0.0.1` hard-coded in the `make_app` signature; the `serve()` call also sets it explicitly, so a mis-typed call to `serve(host="0.0.0.0")` is the only escape path and requires a deliberate code change. - **Subprocess in `/orchestrator/invoke`**: uses list-form argv (`[sys.executable, str(run_py), request.prompt]`) — no `shell=True`, so `request.prompt` cannot be shell-interpolated. However, the prompt text IS passed as a subprocess argument; a hostile prompt could attempt argument injection if `run.py` does further shell expansion. Mitigation: `run.py` receives the prompt as `sys.argv[1]` (a single string, not shell-parsed). Flag for GPT-4.1 cross-review. - **No secrets in code**: `AGENT_TEAM_API_TOKEN` read from env at request time, never hardcoded or logged. - **`invoker_multi.py`**: no authentication or authz on the in-process model calls — they run with the same credentials as the daemon process. Acceptable for the 127.0.0.1-bound, VPN-only surface. ## OUTSTANDING (attended — do NOT merge until done) - [x] **GPT-4.1 cross-family review** — mandatory before merge; no Claude self-review counts. Specifically review the `orchestrator/invoke` subprocess argument-passing and the bearer-token timing guarantee. - [ ] **`/sh-security-review`** — formal security sign-off on the new HTTP API surface (bearer-token auth, 127.0.0.1 bind, subprocess prompt-passing). - [x] **Live deploy steps (attended, on R720 box)**: - `pip install fastapi uvicorn` on the box (or add to the box's requirements) - Set `AGENT_TEAM_API_TOKEN` in `~/secrev.env` (strong random token, never committed) - Restart the `orchestrator` systemd unit - Smoke-test: `curl -H "Authorization: Bearer <token>" http://127.0.0.1:8765/tasks` — expect 405 (method not allowed), confirming the API is up - Verify VPN-only access (the API must NOT be reachable from outside the VPN interface) --- _Generated by [Claude Code](https://claude.ai/code/session_01QYp761G9HojkmLqLZrASVi)_
This repo is archived. You cannot comment on pull requests.
No description provided.