Reworks the P1 sim so the four §7.1 exit criteria are demonstrated against the ACTUAL mechanic, not a model (resolves the verifier's "sim models the ledger, not the LangGraph integration" finding). - New tests/sim/test_p1_graph_integration.py drives the real agent_team.graph StateGraph (interrupt/Command(resume)) + the real langgraph SqliteSaver checkpointer + the committed pending_questions compare-and-set, proving: (a) suspend survives a simulated restart (drop saver/conn, rebuild over the same checkpoint DB) and resumes; (b) duplicate answer loses the CAS and the graph never double-advances; (c) a post-deadline answer loses to expire and the task is not resumed; (d) two concurrent tasks resume to the correct thread, with a turn-guarded no-double-apply check. - graph.py: derive a STABLE question_id from uuid5(thread_id, turn). The clarifier node replays on resume, so the prior fresh-uuid id changed between the delivered/ledgered question and the qa_history entry — breaking the §3.3.1 identity contract. Now the delivered id == ledger key == history entry (unit-tested in test_graph.py). - harness._connect() now uses the committed schema.connect() (WAL + busy_timeout) instead of a raw sqlite3.connect, so concurrent responders genuinely serialize; the criterion-(d) concurrency test no longer swallows OperationalError (it asserts zero errors + exactly one CAS winner). - requirements.txt: pin langgraph-checkpoint-sqlite==3.1.0 (design D9 durable checkpointer), now exercised by the integration test. Full suite: 564 passed; ruff + format clean. |
||
|---|---|---|
| .. | ||
| conftest.py | ||
| harness.py | ||
| test_p1_exit_criteria.py | ||
| test_p1_graph_integration.py | ||