open-swe/agent/reviewer_eval_store.py
Johannes du Plessis 0927f2dd9c
feat: Run reviewer eval in a GitHub Action; dashboard becomes read-only (#1556)
* Run reviewer eval in a GitHub Action; make dashboard a read-only progress view

The dashboard launched the eval as a subprocess inside the serving deployment
worker, so a container recycle killed long runs and discarded results that had
already completed server-side. Move the harness to a workflow_dispatch Action
(run on prod). run_eval now publishes status/progress/log-tail to the LangGraph
store record the dashboard reads, so /admin/evals stays a live view; a killed
Action surfaces as failed via the stale-heartbeat reconcile.

* reviewer_eval workflow: pass inputs via env, no shell interpolation

Addresses the reviewer finding: workflow_dispatch string inputs were
interpolated into the run: block (limit unquoted), allowing shell injection in
a job holding LANGSMITH/ANTHROPIC keys. Pass inputs through env and reference
quoted "$VARS"; validate limit is numeric and build its flag in bash.

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
2026-06-16 19:38:36 -07:00

24 lines
871 B
Python

"""Shared constants for the reviewer-eval status record.
Kept deliberately light (no dashboard/server imports) so the eval harness in the
GitHub Action can publish progress to the store without importing the FastAPI
dashboard. Both ``agent.dashboard.eval_jobs`` (reader) and
``evals.reviewer.store_reporter`` (writer) import from here.
"""
from __future__ import annotations
import re
EVALS_NAMESPACE: list[str] = ["evals"]
REVIEWER_EVAL_KEY = "reviewer"
DEFAULT_EVAL_PROJECT = "open-swe-evals"
_LOG_TAIL_CHARS = 12000
_EXPERIMENT_URL_RE = re.compile(r"https://\S*smith\.langchain\.com/\S+")
# The running Action refreshes the heartbeat this often; a record is only
# reconciled as failed once its heartbeat is older than the stale threshold, so
# a brief dashboard/Action lag doesn't kill a live run.
_HEARTBEAT_INTERVAL_SECONDS = 10
_HEARTBEAT_STALE_SECONDS = 60