Commit graph

42 commits

Author SHA1 Message Date
Adam Moussa
63b6f11bf6 Add tests for the Slack interactions endpoint 2026-05-29 14:33:26 -04:00
Adam Moussa
817e1f6a74 Add synthetic export fixture and run classification quality gate in CI 2026-05-29 14:30:44 -04:00
Adam Moussa
1a126c8179 Enable test execution in CI 2026-05-29 14:30:44 -04:00
Adam Moussa
dcbb0606e6 Add pytest config and conftest for test discovery 2026-05-29 14:30:44 -04:00
Adam Moussa
54ef5400b1
Merge pull request #13 from Sea-Haven-Industries/feature/phase-6-docs
Some checks are pending
Deploy / deploy (push) Waiting to run
Phase 6: docs and re-enable CD on merge
2026-05-29 14:03:59 -04:00
Adam Moussa
6ddcd56f6c
Merge pull request #11 from Sea-Haven-Industries/feature/phase-5-grafana
Phase 5: self-hosted Grafana (EC2, ALB, dashboards-as-code)
2026-05-29 13:58:30 -04:00
Adam Moussa
c41836b319
Merge pull request #10 from Sea-Haven-Industries/feature/phase-4-slack
Phase 4: Slack post + interactions Lambdas (drill-down modals)
2026-05-29 13:55:24 -04:00
Adam Moussa
4cc58f5683
Merge pull request #9 from Sea-Haven-Industries/feature/phase-3-analytics
Some checks are pending
Deploy / deploy (push) Waiting to run
Phase 3: analytics dataset (Glue projection table + Athena)
2026-05-29 13:49:47 -04:00
Adam Moussa
0e96cb59d7
Merge pull request #8 from Sea-Haven-Industries/feature/phase-2-classifier
Phase 2: two-axis classifier Lambda
2026-05-29 13:43:59 -04:00
Adam Moussa
9ee3d84fe0 Enable QEMU for arm64 Lambda bundling in CI/CD 2026-05-29 13:37:01 -04:00
Adam Moussa
0ca738ca8c Enable QEMU for arm64 Lambda bundling in CI/CD 2026-05-29 13:37:00 -04:00
Adam Moussa
c4b9bd4305 Enable QEMU for arm64 Lambda bundling in CI/CD 2026-05-29 13:36:59 -04:00
Adam Moussa
5106853cb9 Enable QEMU for arm64 Lambda bundling in CI/CD 2026-05-29 13:36:58 -04:00
Adam Moussa
d33b82bd92 Enable QEMU for arm64 Lambda bundling in CI/CD 2026-05-29 13:36:56 -04:00
Adam Moussa
9346e29763
Merge pull request #7 from Sea-Haven-Industries/feature/phase-1-ingestion
Phase 1: add drop-folder ingestion and scoped uploader IAM user (Phase 1)
2026-05-29 13:24:07 -04:00
Adam Moussa
f86b4ba1c5
Merge pull request #6 from Sea-Haven-Industries/feature/phase-0-scaffold
Phase 0: scaffold apm-wo-analysis repository
2026-05-29 13:14:37 -04:00
Adam Moussa
1a231131bd Revert "Temporarily disable CD deploy during stack merges"
This reverts commit 611979e44e.
2026-05-29 13:07:02 -04:00
Adam Moussa
fee0fe9a2c Merge phase-0 to carry the CD-disable into phase-6 for an explicit revert 2026-05-29 13:06:57 -04:00
Adam Moussa
dec02170e6 Mark Phase 6 Confluence/Slack docs complete 2026-05-29 12:59:22 -04:00
Adam Moussa
dcda61e728 Add operational runbook (Phase 6)
docs/RUNBOOK.md: incident runbook for a missing daily analysis (detection →
context → triage → resolution-by-cause → post-incident), plus operational
procedures — export upload (direct + drop-folder agent), Grafana OS/app/plugin
patching cadence (clean-replacement preferred), dashboard-JSON redeploy flow +
gotchas, config + grafana.db/EBS backup-restore (DLM snapshot), and a common-
failures quick index. Mirrors to Confluence.
2026-05-29 12:31:14 -04:00
Adam Moussa
60e0b878e3 Refresh README to full operational doc (Phase 6)
Rewrite to the Sea Haven operational template with real resource names from both
stacks: AWS Resources + Lambda Functions tables, Configuration (Secrets/SSM/env/
context), Operations (verify, logs, classifier DLQ, reprocess, Grafana admin),
Documentation, and Notes/Gotchas incl. the deploy-time lessons + a known-debt
list. Corrects stale bits: ALB is internet-facing + office-IP-restricted (not
internal); classifier uses partition projection (no runtime Glue registration);
summary/details JSON live under meta/ not analytics/; grafana uses an instance
role. Adds the resources missing from the old README (classifier DLQ, slack-
interactions Lambda, HTTP API, meta/ + grafana-config/ prefixes, DLM backup,
encrypted volume).
2026-05-29 12:27:05 -04:00
Adam Moussa
611979e44e Temporarily disable CD deploy during stack merges
Gate the deploy job with if:${{ false }} so merging the Phase 0-5 stack into
main does not fire cdk deploy --all on every merge. Both stacks are already
deployed manually and validated in prod. Re-enable at the start of Phase 6 by
reverting this commit.
2026-05-29 11:43:48 -04:00
Adam Moussa
eae4d67e00 Apply cross-review findings (Phase 2/4/5 hardening)
From the cross_reviewer (GPT-4.1) per-PR passes, now that the orchestrator is
back up:

Phase 2 (classifier):
- Process ALL S3 records, not just event["Records"][0] — batched notifications
  no longer silently dropped (the review's only BLOCK).
- Derive the partition dt from the S3 event time, not the Lambda wall-clock —
  stable across retries / the midnight boundary.
- Add an SQS dead-letter queue so a failed run surfaces instead of dropping a
  day's data after Lambda's retries.

Phase 4 (Slack):
- Stage throttling (rate 10 / burst 20) on the public /slack/interactions HTTP
  API. (AWS WAF doesn't attach to apigwv2 HTTP APIs; stage throttling is the
  mechanism.)

Phase 5 (Grafana):
- Explicit encrypted=True on the gp3 root volume.

Tests: synth assertions for the DLQ, stage throttling, and the encrypted volume.
60/60 pass; cdk synth green for both stacks. Deferred NITs (print->logging, sig-
failure source-IP logging, S3 versioning, CIDR-maintenance runbook) -> Phase 6.

NOTE: like the earlier deploy fixes these sit on phase-5 but span phases — the
classifier/DLQ to #8, throttling to #10, encryption to #11 — reconcile at merge.
The encrypted-volume change needs the deferred clean instance replacement to
take effect (can't encrypt a live volume in place).
2026-05-29 11:29:14 -04:00
Adam Moussa
c970635be1 Fix WO table filter: drop custom allValue so :singlequote expands All
The 5 multi-select filter vars had allValue='All'. Grafana does NOT apply the
:singlequote format to a custom allValue, so 'All' was injected bare into
site IN (ALL) -> Athena read ALL as a column ('Column ALL cannot be resolved').
Removing the custom allValue lets :singlequote expand the All selection to the
real quoted value list, so the IN clause is valid SQL.
2026-05-29 11:13:10 -04:00
Adam Moussa
cb01b892bc Fix Grafana panels: rawSQL not rawSql (Athena plugin query key)
All 13 panel/variable queries keyed the SQL as rawSql (lowercase); the
grafana-athena-datasource plugin reads rawSQL (capital SQL). With the wrong key
the plugin saw an empty query, so no Athena query ever fired — variables had no
options and every panel showed a clean 'No data' (no error). This was the root
cause of the empty dashboard; data/datasource/permissions were all fine.
2026-05-29 10:55:02 -04:00
Adam Moussa
d884de8fcc Fix Grafana config-sync: --exact-timestamps for same-size updates
aws s3 sync skips same-size files on download unless --exact-timestamps is set,
so a dashboard edit that doesn't change file size (e.g. refresh 2->1, or a query
tweak) never propagated to the instance. Add --exact-timestamps to all four
sync invocations (boot + 15-min timer).
2026-05-28 19:06:06 -04:00
Adam Moussa
dce53fdc86 Fix Grafana template vars: refresh on dashboard load, not time-range change
All 6 query variables had refresh=2 (on time-range change) with no cached value,
so a plain dashboard load never populated them — $dt resolved to empty and every
panel filtered WHERE dt='' (no data). Set refresh=1 (on dashboard load).
2026-05-28 19:01:24 -04:00
Adam Moussa
3c8b7704f6 Fix Grafana Athena auth: use default credential chain, not ec2_iam_role
Grafana rejected the datasource with 'trying to use non-allowed auth method
ec2_iam_role: Failed to create client' — the plugin's allowed_auth_providers
defaults to default,keys,credentials and excludes ec2_iam_role. Switch authType
to 'default' (AWS SDK default chain), which on EC2 resolves to the instance role
via IMDS (still no static keys) and is allowed out of the box.
2026-05-28 18:59:09 -04:00
Adam Moussa
68f3acded8 Fix dashboard rendering: barchart panel type + metadata off the table prefix
Two issues found loading the deployed dashboard:

1. Panels used type "bar-chart" (hyphenated); Grafana's core panel is "barchart"
   — hence "plugin bar-chart required". Fixed both panels.

2. ALL panels showed "no data" because Athena failed with HIVE_BAD_DATA:
   the classifier wrote summary.json/details.json INTO analytics/dt=*/ — the
   same prefix the Glue table scans — so Athena tried to read the JSON as
   Parquet and every query failed. Move the metadata to a separate meta/dt=*/
   prefix: classifier writes there (grant_read_write meta/*), the Slack Lambdas
   read there (read_meta_json, grant_read meta/*), and analytics/ holds only
   Parquet. Verified: the category GROUP BY query now succeeds against Athena.
2026-05-28 18:49:41 -04:00
Adam Moussa
ea54cb1e60 Fix Grafana bootstrap: grafana-cli --homepath + valid plugin version
Found on first boot (cloud-init errored, grafana-server never started):
- grafana-cli needs --homepath=/usr/share/grafana or it can't find config
  defaults; under set -e that aborted the whole bootstrap.
- pinned plugin version 2.18.2 doesn't exist (conflated with awswrangler's
  version) — grafana-athena-datasource latest is 3.2.0.

Verified by running the corrected bootstrap on the instance via SSM: plugin
installs, grafana-server active, /api/health 200, ALB target healthy.

NOTE: the running instance was repaired in-place (the user-data change updated
the launch template but did not replace the instance). The committed user-data
is now correct, so a fresh launch boots clean — a one-time clean instance
replacement should validate that before prod sign-off.
2026-05-28 18:43:08 -04:00
Adam Moussa
f193754c27 Fix deploy-time failures found in prod testing
Two issues only a real deploy/run surfaced (synth + offline tests passed):

1. Classifier exceeded Lambda's 250 MB unzipped limit (bundled awswrangler +
   pandas + pyarrow + numpy). Move them to the AWS-managed SDK-for-pandas layer
   (AWSSDKPandas-Python312-Arm64:27, awswrangler 3.16.1, pre-stripped to fit);
   bundle only openpyxl. Drop the unused anthropic SDK — _call_haiku uses stdlib
   urllib. Function package now ~890 KB.

2. Slack rejected the daily post with invalid_blocks: every category drill
   button shared action_id "drill_category". Qualify it as "drill_category:<cat>"
   for uniqueness; the interactions handler now matches on the prefix. Add a
   regression test asserting all daily-summary action_ids are unique.

Verified in prod: classifier writes Parquet + summary.json + details.json;
slack-post posts the daily summary + 3rd-escalation alert; the interactions
endpoint (apm-wo.seahaven.com) returns 401 on a bad signature. 58/58 tests pass.

NOTE: these fixes sit on the phase-5 branch but logically belong to earlier
phases — the layer fix to #8 (classifier), the Slack fix to #10 — and must be
moved/cherry-picked there before those PRs merge independently. See cleanup.
2026-05-28 18:29:16 -04:00
Adam Moussa
c04c239774 Update README status for Phase 5 Grafana stack 2026-05-28 18:06:10 -04:00
Adam Moussa
4870784fbd Add self-hosted Grafana stack: EC2, ALB, dashboards-as-code (Phase 5)
The one non-serverless piece — Grafana OSS on a t4g.small (AL2023, ARM64) in the
imported seahaven-vpc, fronted by an internet-facing ALB locked by SG to the
office CIDRs (no Client VPN exists, so "VPN-only" = office-IP restriction, the
syslog-server pattern). Instance in private subnets, reachable only from the ALB
SG, administered via SSM Session Manager (no SSH/key pair).

grafana_stack.py: ALB (HTTPS, *.seahaven.com cert, open=False so the SG office
rules aren't undone by an auto 0.0.0.0/0), instance role (Athena query + Glue
read + S3 analytics/athena-results, no static keys), Route53 grafana.seahaven.com
alias, gp3 root volume RETAINed, daily DLM snapshot of the tagged instance, and a
BucketDeployment that uploads grafana/ to the S3 config prefix.

grafana_userdata.sh: install Grafana OSS, pin the Athena datasource plugin, write
grafana.ini (root_url grafana.seahaven.com, kiosk embedding), sync provisioning +
dashboards from S3 on boot, and a systemd timer re-syncs every 15 min so repo
edits land without an instance rebuild.

Dashboard (grafana-author agent, grafana/dashboards/apm-work-orders.json, uid
apm-wo so the Slack 📊 button resolves): 7 panels — category distribution,
escalation summary, action/routine, escalations-by-site, trend time-series over
dt (the new capability), filterable WO table (5 template vars, escalation row
coloring, CSV export, no APM links), and the mismatch panel. Datasource uid
"athena" pinned in the provisioning yaml.

Tests: tests/test_grafana_synth.py — ALB admits only the office CIDRs on 443
(caught and fixed a default 0.0.0.0/0 listener rule), instance only-from-ALB,
no static keys, scoped instance role + SSM, gp3+retained root volume, daily DLM
backup, grafana.seahaven.com alias. 57/57 tests pass; full cdk synth green.
2026-05-28 18:05:20 -04:00
Adam Moussa
cfb9fb85d2 Update README for Phase 4 Slack surfaces 2026-05-28 17:49:57 -04:00
Adam Moussa
c3a2c936d1 Add Slack post + interactions Lambdas with drill-down modals (Phase 4)
Two push surfaces (no App Home) + interactive drill-down, per CLAUDE.md.

Block Kit (blockkit.py, pure/offline): build_daily_summary (header, vs-yesterday
deltas, escalation breakdown with 3rd highlighted, action/routine, top sites,
mismatch callout, category drill buttons + 📊 Open dashboard link, footer),
build_escalation_alert (one @here, returns None on zero-3rd — suppression), and
build_wo_modal (views.open payload, capped under Slack's 100-block limit).

Lambdas: slack_post/handler.py (classifier-invoked: read today/yesterday
summary.json, post daily summary, conditionally post the batched alert from
details.json) and slack_post/interactions.py (API Gateway: verify Slack
signature, filter details.json, views.open the WO modal within the 3s trigger_id
window). slackio.py centralizes Secrets Manager creds, the SSM dashboard URL,
signature verification, and analytics/ reads — keeping blockkit pure.

Classifier: emit analytics/dt=*/details.json (per-WO index for the modals) and
async-invoke slack-post after the snapshot write (best-effort; a Slack failure
never fails classification).

CDK: slack-post + interactions Lambdas (Docker-bundled slack_sdk), HTTP API on
apm-wo.seahaven.com (wildcard ACM cert + Route53 alias; signature-verified, so
the route is unauthenticated by design), SSM /apm-wo-analysis/grafana-base-url,
and scoped IAM (read analytics/, read the Slack secret + dashboard param;
classifier granted lambda:InvokeFunction on slack-post). Slack creds live in one
Secrets Manager secret apm-wo-analysis/slack-credentials {botToken, signingSecret,
channelId}; cdk.json gains cert/zone/domain context.

WO drill-downs link to Grafana only — no APM deep-links (per decision).

Deliverables for test time: slack/manifest.yaml (app manifest, interactivity
request_url = apm-wo.seahaven.com).

Tests: tests/test_blockkit.py (30 offline cases — deltas, zero-3rd None, <100
blocks under large inputs, modal truncation/overflow, dashboard URL) and Phase 4
assertions in test_pipeline_synth.py (both Lambdas, the API route/domain/alias,
and no broad/write IAM on the Slack roles). 49/49 tests pass; cdk synth green.
2026-05-28 17:48:51 -04:00
Adam Moussa
23420d841d Update README status for Phase 3 analytics dataset 2026-05-28 17:25:27 -04:00
Adam Moussa
6ee218eea3 Add analytics dataset: projection table + Athena workgroup (Phase 3)
CDK-define the apm_wo_snapshots Glue table with partition projection over
analytics/ (dt as projected date partition, 2026-01-01..NOW). Projection means
no crawler, no MSCK REPAIR, and — critically — the classifier needs no Glue
catalog access at all.

pipeline_stack.py: glue.CfnTable (Parquet SerDe, 17-column schema mirroring the
classifier's snapshot incl. contractor_description and the two boolean flags) +
athena.CfnWorkGroup `apm-wo-analysis` (enforced result location, SSE-S3) + a
30-day lifecycle rule on athena-results/ (disposable query output in a RETAIN
bucket). Trim the classifier role: drop the entire Glue policy statement.

handler.py: stop registering the table at runtime — drop database=/table= from
to_parquet so the classifier writes pure Parquet; partition projection handles
the rest. Keeps overwrite_partitions for idempotent same-day re-uploads.

test_pipeline_synth.py: offline synth assertions (bundling skipped) — projection
properties, column schema/types, workgroup result enforcement, and that no IAM
policy grants glue:* to the classifier.

cdk synth green; 11/11 tests pass.
2026-05-28 17:24:36 -04:00
Adam Moussa
0272db1716 Update README for Phase 2 classifier
Reflect the implemented two-axis classifier + handler, the 5.2% 'Other'
smoke-test result, and the phased build status (Phase 2 in review).
2026-05-28 17:13:11 -04:00
Adam Moussa
eb724580e0 Add two-axis classifier Lambda and CDK wiring (Phase 2)
Implement the core classification engine and wire it into the pipeline stack.

classify.py: two-axis classifier — HTML-strip, comment-intent regex buckets
(escalations → status inquiry, most-specific first), Hold Reason / WO Status
structured state, comment-vs-state mismatch detector, and a Claude Haiku
fallback (Secrets Manager key) reserved for ambiguous free-text. Exports
ESCALATION_CATEGORIES / ACTION_NEEDED_CATEGORIES.

handler.py: S3-triggered handler — parse xlsx/csv, classify each non-blank
row, write a per-WO Parquet snapshot to analytics/dt=YYYY-MM-DD/ (registers
the Glue partition via awswrangler) and a summary.json for slack-post (Phase 4).

pipeline_stack.py: Glue database, ARM64 Python 3.12 classifier Lambda
(Docker-bundled deps), S3 raw/ notification (.xlsx/.csv), and least-privilege
IAM (read raw/, read-write analytics/, scoped Glue catalog, read Anthropic key).

Smoke-tested against the real export: 347 rows, "Other" at 5.2% (target ~9%),
18 mismatches flagged. 7/7 unit + smoke tests pass; cdk synth green.
2026-05-28 17:09:52 -04:00
Adam Moussa
7befc8f0d7 Add drop-folder ingestion and scoped uploader IAM user
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.

Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.

The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
2026-05-28 16:32:56 -04:00
Adam Moussa
58b91bda70 Scaffold apm-wo-analysis repository
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.

Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.

Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.

Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.
2026-05-28 16:13:13 -04:00
Adam Moussa
f24ded21a7 Add project README 2026-05-28 16:13:01 -04:00