The classifier DLQ (apm-wo-analysis-classifier-dlq) had no alarm, so a
message landing there after Lambda exhausts async retries — a genuinely
dropped classifier run — surfaced nowhere.
Add a CloudWatch alarm on AWS/SQS ApproximateNumberOfMessagesVisible
(Maximum, threshold >0, 300s period, 1 evaluation period, notBreaching)
that pages the shared site-alerts SNS topic. Mirrors the
workorder-email-processor-dlq-messages alarm in procurement-ingest's
wo_stack and the payments-payroll-batch-dlq-messages alarm convention.
ALARM-only, no OK action, per the CloudWatch-alarm preference.
Org pin moved from 2.253.1 (bundles vulnerable fast-uri 3.1.0, 2 high
GHSAs) to 2.257.0 (bundles patched 3.1.2). Removes the blanket
dependabot ignore per the new handbook pinning policy (exact pins kept
current by Dependabot; blanket ignores banned).
* Add dependency-review caller workflow
Add a pull_request-triggered caller that invokes the org-level
callable-dependency-review workflow to scan dependency changes and
fail on high-severity advisories.
* chore: retrigger checks
* chore: retrigger dep review (post-fix)
Rewrite to the Sea Haven operational template with real resource names from both
stacks: AWS Resources + Lambda Functions tables, Configuration (Secrets/SSM/env/
context), Operations (verify, logs, classifier DLQ, reprocess, Grafana admin),
Documentation, and Notes/Gotchas incl. the deploy-time lessons + a known-debt
list. Corrects stale bits: ALB is internet-facing + office-IP-restricted (not
internal); classifier uses partition projection (no runtime Glue registration);
summary/details JSON live under meta/ not analytics/; grafana uses an instance
role. Adds the resources missing from the old README (classifier DLQ, slack-
interactions Lambda, HTTP API, meta/ + grafana-config/ prefixes, DLM backup,
encrypted volume).
Gate the deploy job with if:${{ false }} so merging the Phase 0-5 stack into
main does not fire cdk deploy --all on every merge. Both stacks are already
deployed manually and validated in prod. Re-enable at the start of Phase 6 by
reverting this commit.
From the cross_reviewer (GPT-4.1) per-PR passes, now that the orchestrator is
back up:
Phase 2 (classifier):
- Process ALL S3 records, not just event["Records"][0] — batched notifications
no longer silently dropped (the review's only BLOCK).
- Derive the partition dt from the S3 event time, not the Lambda wall-clock —
stable across retries / the midnight boundary.
- Add an SQS dead-letter queue so a failed run surfaces instead of dropping a
day's data after Lambda's retries.
Phase 4 (Slack):
- Stage throttling (rate 10 / burst 20) on the public /slack/interactions HTTP
API. (AWS WAF doesn't attach to apigwv2 HTTP APIs; stage throttling is the
mechanism.)
Phase 5 (Grafana):
- Explicit encrypted=True on the gp3 root volume.
Tests: synth assertions for the DLQ, stage throttling, and the encrypted volume.
60/60 pass; cdk synth green for both stacks. Deferred NITs (print->logging, sig-
failure source-IP logging, S3 versioning, CIDR-maintenance runbook) -> Phase 6.
NOTE: like the earlier deploy fixes these sit on phase-5 but span phases — the
classifier/DLQ to #8, throttling to #10, encryption to #11 — reconcile at merge.
The encrypted-volume change needs the deferred clean instance replacement to
take effect (can't encrypt a live volume in place).
The 5 multi-select filter vars had allValue='All'. Grafana does NOT apply the
:singlequote format to a custom allValue, so 'All' was injected bare into
site IN (ALL) -> Athena read ALL as a column ('Column ALL cannot be resolved').
Removing the custom allValue lets :singlequote expand the All selection to the
real quoted value list, so the IN clause is valid SQL.
All 13 panel/variable queries keyed the SQL as rawSql (lowercase); the
grafana-athena-datasource plugin reads rawSQL (capital SQL). With the wrong key
the plugin saw an empty query, so no Athena query ever fired — variables had no
options and every panel showed a clean 'No data' (no error). This was the root
cause of the empty dashboard; data/datasource/permissions were all fine.
aws s3 sync skips same-size files on download unless --exact-timestamps is set,
so a dashboard edit that doesn't change file size (e.g. refresh 2->1, or a query
tweak) never propagated to the instance. Add --exact-timestamps to all four
sync invocations (boot + 15-min timer).
All 6 query variables had refresh=2 (on time-range change) with no cached value,
so a plain dashboard load never populated them — $dt resolved to empty and every
panel filtered WHERE dt='' (no data). Set refresh=1 (on dashboard load).
Grafana rejected the datasource with 'trying to use non-allowed auth method
ec2_iam_role: Failed to create client' — the plugin's allowed_auth_providers
defaults to default,keys,credentials and excludes ec2_iam_role. Switch authType
to 'default' (AWS SDK default chain), which on EC2 resolves to the instance role
via IMDS (still no static keys) and is allowed out of the box.
Two issues found loading the deployed dashboard:
1. Panels used type "bar-chart" (hyphenated); Grafana's core panel is "barchart"
— hence "plugin bar-chart required". Fixed both panels.
2. ALL panels showed "no data" because Athena failed with HIVE_BAD_DATA:
the classifier wrote summary.json/details.json INTO analytics/dt=*/ — the
same prefix the Glue table scans — so Athena tried to read the JSON as
Parquet and every query failed. Move the metadata to a separate meta/dt=*/
prefix: classifier writes there (grant_read_write meta/*), the Slack Lambdas
read there (read_meta_json, grant_read meta/*), and analytics/ holds only
Parquet. Verified: the category GROUP BY query now succeeds against Athena.
Found on first boot (cloud-init errored, grafana-server never started):
- grafana-cli needs --homepath=/usr/share/grafana or it can't find config
defaults; under set -e that aborted the whole bootstrap.
- pinned plugin version 2.18.2 doesn't exist (conflated with awswrangler's
version) — grafana-athena-datasource latest is 3.2.0.
Verified by running the corrected bootstrap on the instance via SSM: plugin
installs, grafana-server active, /api/health 200, ALB target healthy.
NOTE: the running instance was repaired in-place (the user-data change updated
the launch template but did not replace the instance). The committed user-data
is now correct, so a fresh launch boots clean — a one-time clean instance
replacement should validate that before prod sign-off.
Two issues only a real deploy/run surfaced (synth + offline tests passed):
1. Classifier exceeded Lambda's 250 MB unzipped limit (bundled awswrangler +
pandas + pyarrow + numpy). Move them to the AWS-managed SDK-for-pandas layer
(AWSSDKPandas-Python312-Arm64:27, awswrangler 3.16.1, pre-stripped to fit);
bundle only openpyxl. Drop the unused anthropic SDK — _call_haiku uses stdlib
urllib. Function package now ~890 KB.
2. Slack rejected the daily post with invalid_blocks: every category drill
button shared action_id "drill_category". Qualify it as "drill_category:<cat>"
for uniqueness; the interactions handler now matches on the prefix. Add a
regression test asserting all daily-summary action_ids are unique.
Verified in prod: classifier writes Parquet + summary.json + details.json;
slack-post posts the daily summary + 3rd-escalation alert; the interactions
endpoint (apm-wo.seahaven.com) returns 401 on a bad signature. 58/58 tests pass.
NOTE: these fixes sit on the phase-5 branch but logically belong to earlier
phases — the layer fix to #8 (classifier), the Slack fix to #10 — and must be
moved/cherry-picked there before those PRs merge independently. See cleanup.