2026-07-03 11:46:00 -04:00
|
|
|
.PHONY: all format format-check lint test tests integration_tests help run dev \
|
|
|
|
|
install-hooks triage-sync triage-reconcile triage-render triage-check
|
2026-02-04 18:30:38 -08:00
|
|
|
|
|
|
|
|
# Default target executed when no arguments are given to make.
|
|
|
|
|
all: help
|
|
|
|
|
|
|
|
|
|
######################
|
|
|
|
|
# DEVELOPMENT
|
|
|
|
|
######################
|
|
|
|
|
|
|
|
|
|
dev:
|
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
|
|
|
uv run langgraph dev
|
2026-02-04 18:30:38 -08:00
|
|
|
|
|
|
|
|
run:
|
feat: add reviewer graph + eval target wiring (#1241)
* feat: add reviewer graph + eval target wiring
- New `reviewer` graph (`agent/reviewer.py`) registered in langgraph.json
alongside the main `agent` graph. Reuses the same sandbox lifecycle,
GH proxy auth, and middleware primitives from `agent.server`, but with
a narrower tool set, a reviewer-specific system prompt, no
commit/push, and the `task` (subagent) tool stripped via
`_ToolExclusionMiddleware` so review stays in one context.
- New `github_comment` tool: agents call it once per issue with
`(file, line, body, severity)` and the eval scores those calls
against golden comments.
- `ensure_no_empty_msg` middleware (the no_op nudge) is intentionally
*not* on the reviewer's stack — that middleware exists to enforce the
main agent's "always finalize via Slack/Linear/PR" contract, which
the reviewer doesn't have. The main agent's behavior is unchanged.
- `evals/reviewer/target.py`: send PR info as a user message, extract
every `github_comment` tool call (multiple expected per review) into
the run output.
- `evals/reviewer/judge.py`: per-example evaluator now returns a list
of metrics under `{"results": [...]}` so LangSmith averages each
numeric key (f1/precision/recall/tp/fp/fn) across the experiment in
the UI. Dropped the broken `aggregate_pr` summary evaluator that
reached for an attribute that doesn't exist on `RunTree`.
- `evals/reviewer/run_eval.py`: `--limit` now slices the dataset via
`client.list_examples(limit=N)` since `aevaluate` doesn't accept
`max_examples`.
- Makefile: `dev` and `run` targets now use `uv run` so they work
without an activated venv.
* resolve comments
---------
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-05-06 10:15:58 -07:00
|
|
|
uv run uvicorn agent.webapp:app --reload --port 8000
|
2026-02-04 18:30:38 -08:00
|
|
|
|
|
|
|
|
install:
|
|
|
|
|
uv pip install -e .
|
|
|
|
|
|
|
|
|
|
######################
|
|
|
|
|
# TESTING
|
|
|
|
|
######################
|
|
|
|
|
|
|
|
|
|
TEST_FILE ?= tests/
|
|
|
|
|
|
|
|
|
|
test tests:
|
2026-02-23 19:11:41 -08:00
|
|
|
@if [ -d "$(TEST_FILE)" ] || [ -f "$(TEST_FILE)" ]; then \
|
|
|
|
|
uv run pytest -vvv $(TEST_FILE); \
|
|
|
|
|
else \
|
|
|
|
|
echo "Skipping tests: path not found: $(TEST_FILE)"; \
|
|
|
|
|
fi
|
2026-02-04 18:30:38 -08:00
|
|
|
|
|
|
|
|
integration_tests:
|
2026-02-23 19:11:41 -08:00
|
|
|
@if [ -d "tests/integration_tests/" ] || [ -f "tests/integration_tests/" ]; then \
|
|
|
|
|
uv run pytest -vvv tests/integration_tests/; \
|
|
|
|
|
else \
|
|
|
|
|
echo "Skipping integration tests: path not found: tests/integration_tests/"; \
|
|
|
|
|
fi
|
2026-02-04 18:30:38 -08:00
|
|
|
|
|
|
|
|
######################
|
|
|
|
|
# LINTING AND FORMATTING
|
|
|
|
|
######################
|
|
|
|
|
|
|
|
|
|
PYTHON_FILES=.
|
|
|
|
|
|
|
|
|
|
lint:
|
|
|
|
|
uv run ruff check $(PYTHON_FILES)
|
|
|
|
|
uv run ruff format $(PYTHON_FILES) --diff
|
|
|
|
|
|
|
|
|
|
format:
|
|
|
|
|
uv run ruff format $(PYTHON_FILES)
|
|
|
|
|
uv run ruff check --fix $(PYTHON_FILES)
|
|
|
|
|
|
2026-02-23 19:11:41 -08:00
|
|
|
format-check:
|
|
|
|
|
uv run ruff format $(PYTHON_FILES) --check
|
|
|
|
|
|
2026-07-03 11:46:00 -04:00
|
|
|
######################
|
|
|
|
|
# UPSTREAM SYNC / CHERRY-PICK TRIAGE
|
|
|
|
|
######################
|
|
|
|
|
|
|
|
|
|
install-hooks:
|
|
|
|
|
bash scripts/install-hooks.sh
|
|
|
|
|
|
|
|
|
|
triage-sync:
|
|
|
|
|
python3 scripts/triage.py sync
|
|
|
|
|
|
|
|
|
|
triage-reconcile:
|
|
|
|
|
python3 scripts/triage.py reconcile
|
|
|
|
|
|
|
|
|
|
triage-render:
|
|
|
|
|
python3 scripts/triage.py generate
|
|
|
|
|
|
|
|
|
|
triage-check:
|
|
|
|
|
python3 scripts/triage.py generate --check
|
|
|
|
|
|
2026-02-04 18:30:38 -08:00
|
|
|
######################
|
|
|
|
|
# HELP
|
|
|
|
|
######################
|
|
|
|
|
|
|
|
|
|
help:
|
|
|
|
|
@echo '----'
|
|
|
|
|
@echo 'dev - run LangGraph dev server'
|
|
|
|
|
@echo 'run - run webhook server'
|
|
|
|
|
@echo 'install - install dependencies'
|
|
|
|
|
@echo 'format - run code formatters'
|
|
|
|
|
@echo 'lint - run linters'
|
|
|
|
|
@echo 'test - run unit tests'
|
|
|
|
|
@echo 'integration_tests - run integration tests'
|
2026-07-03 11:46:00 -04:00
|
|
|
@echo 'install-hooks - install cherry-pick triage git hooks (per clone)'
|
|
|
|
|
@echo 'triage-sync - fetch upstream + add new dev..upstream/main commits as untriaged'
|
|
|
|
|
@echo 'triage-reconcile - drain cherry-pick journal into the triage ledger'
|
|
|
|
|
@echo 'triage-render - regenerate docs/upstream-sync/triage.md'
|
|
|
|
|
@echo 'triage-check - fail if triage.md is stale vs triage.jsonl (CI)'
|