The classifier DLQ (apm-wo-analysis-classifier-dlq) had no alarm, so a
message landing there after Lambda exhausts async retries — a genuinely
dropped classifier run — surfaced nowhere.
Add a CloudWatch alarm on AWS/SQS ApproximateNumberOfMessagesVisible
(Maximum, threshold >0, 300s period, 1 evaluation period, notBreaching)
that pages the shared site-alerts SNS topic. Mirrors the
workorder-email-processor-dlq-messages alarm in procurement-ingest's
wo_stack and the payments-payroll-batch-dlq-messages alarm convention.
ALARM-only, no OK action, per the CloudWatch-alarm preference.
Org pin moved from 2.253.1 (bundles vulnerable fast-uri 3.1.0, 2 high
GHSAs) to 2.257.0 (bundles patched 3.1.2). Removes the blanket
dependabot ignore per the new handbook pinning policy (exact pins kept
current by Dependabot; blanket ignores banned).
* Add dependency-review caller workflow
Add a pull_request-triggered caller that invokes the org-level
callable-dependency-review workflow to scan dependency changes and
fail on high-severity advisories.
* chore: retrigger checks
* chore: retrigger dep review (post-fix)
Rewrite to the Sea Haven operational template with real resource names from both
stacks: AWS Resources + Lambda Functions tables, Configuration (Secrets/SSM/env/
context), Operations (verify, logs, classifier DLQ, reprocess, Grafana admin),
Documentation, and Notes/Gotchas incl. the deploy-time lessons + a known-debt
list. Corrects stale bits: ALB is internet-facing + office-IP-restricted (not
internal); classifier uses partition projection (no runtime Glue registration);
summary/details JSON live under meta/ not analytics/; grafana uses an instance
role. Adds the resources missing from the old README (classifier DLQ, slack-
interactions Lambda, HTTP API, meta/ + grafana-config/ prefixes, DLM backup,
encrypted volume).
Gate the deploy job with if:${{ false }} so merging the Phase 0-5 stack into
main does not fire cdk deploy --all on every merge. Both stacks are already
deployed manually and validated in prod. Re-enable at the start of Phase 6 by
reverting this commit.
From the cross_reviewer (GPT-4.1) per-PR passes, now that the orchestrator is
back up:
Phase 2 (classifier):
- Process ALL S3 records, not just event["Records"][0] — batched notifications
no longer silently dropped (the review's only BLOCK).
- Derive the partition dt from the S3 event time, not the Lambda wall-clock —
stable across retries / the midnight boundary.
- Add an SQS dead-letter queue so a failed run surfaces instead of dropping a
day's data after Lambda's retries.
Phase 4 (Slack):
- Stage throttling (rate 10 / burst 20) on the public /slack/interactions HTTP
API. (AWS WAF doesn't attach to apigwv2 HTTP APIs; stage throttling is the
mechanism.)
Phase 5 (Grafana):
- Explicit encrypted=True on the gp3 root volume.
Tests: synth assertions for the DLQ, stage throttling, and the encrypted volume.
60/60 pass; cdk synth green for both stacks. Deferred NITs (print->logging, sig-
failure source-IP logging, S3 versioning, CIDR-maintenance runbook) -> Phase 6.
NOTE: like the earlier deploy fixes these sit on phase-5 but span phases — the
classifier/DLQ to #8, throttling to #10, encryption to #11 — reconcile at merge.
The encrypted-volume change needs the deferred clean instance replacement to
take effect (can't encrypt a live volume in place).
The 5 multi-select filter vars had allValue='All'. Grafana does NOT apply the
:singlequote format to a custom allValue, so 'All' was injected bare into
site IN (ALL) -> Athena read ALL as a column ('Column ALL cannot be resolved').
Removing the custom allValue lets :singlequote expand the All selection to the
real quoted value list, so the IN clause is valid SQL.
All 13 panel/variable queries keyed the SQL as rawSql (lowercase); the
grafana-athena-datasource plugin reads rawSQL (capital SQL). With the wrong key
the plugin saw an empty query, so no Athena query ever fired — variables had no
options and every panel showed a clean 'No data' (no error). This was the root
cause of the empty dashboard; data/datasource/permissions were all fine.
aws s3 sync skips same-size files on download unless --exact-timestamps is set,
so a dashboard edit that doesn't change file size (e.g. refresh 2->1, or a query
tweak) never propagated to the instance. Add --exact-timestamps to all four
sync invocations (boot + 15-min timer).
All 6 query variables had refresh=2 (on time-range change) with no cached value,
so a plain dashboard load never populated them — $dt resolved to empty and every
panel filtered WHERE dt='' (no data). Set refresh=1 (on dashboard load).
Grafana rejected the datasource with 'trying to use non-allowed auth method
ec2_iam_role: Failed to create client' — the plugin's allowed_auth_providers
defaults to default,keys,credentials and excludes ec2_iam_role. Switch authType
to 'default' (AWS SDK default chain), which on EC2 resolves to the instance role
via IMDS (still no static keys) and is allowed out of the box.