Two issues found loading the deployed dashboard:
1. Panels used type "bar-chart" (hyphenated); Grafana's core panel is "barchart"
— hence "plugin bar-chart required". Fixed both panels.
2. ALL panels showed "no data" because Athena failed with HIVE_BAD_DATA:
the classifier wrote summary.json/details.json INTO analytics/dt=*/ — the
same prefix the Glue table scans — so Athena tried to read the JSON as
Parquet and every query failed. Move the metadata to a separate meta/dt=*/
prefix: classifier writes there (grant_read_write meta/*), the Slack Lambdas
read there (read_meta_json, grant_read meta/*), and analytics/ holds only
Parquet. Verified: the category GROUP BY query now succeeds against Athena.
Two issues only a real deploy/run surfaced (synth + offline tests passed):
1. Classifier exceeded Lambda's 250 MB unzipped limit (bundled awswrangler +
pandas + pyarrow + numpy). Move them to the AWS-managed SDK-for-pandas layer
(AWSSDKPandas-Python312-Arm64:27, awswrangler 3.16.1, pre-stripped to fit);
bundle only openpyxl. Drop the unused anthropic SDK — _call_haiku uses stdlib
urllib. Function package now ~890 KB.
2. Slack rejected the daily post with invalid_blocks: every category drill
button shared action_id "drill_category". Qualify it as "drill_category:<cat>"
for uniqueness; the interactions handler now matches on the prefix. Add a
regression test asserting all daily-summary action_ids are unique.
Verified in prod: classifier writes Parquet + summary.json + details.json;
slack-post posts the daily summary + 3rd-escalation alert; the interactions
endpoint (apm-wo.seahaven.com) returns 401 on a bad signature. 58/58 tests pass.
NOTE: these fixes sit on the phase-5 branch but logically belong to earlier
phases — the layer fix to #8 (classifier), the Slack fix to #10 — and must be
moved/cherry-picked there before those PRs merge independently. See cleanup.
Two push surfaces (no App Home) + interactive drill-down, per CLAUDE.md.
Block Kit (blockkit.py, pure/offline): build_daily_summary (header, vs-yesterday
deltas, escalation breakdown with 3rd highlighted, action/routine, top sites,
mismatch callout, category drill buttons + 📊 Open dashboard link, footer),
build_escalation_alert (one @here, returns None on zero-3rd — suppression), and
build_wo_modal (views.open payload, capped under Slack's 100-block limit).
Lambdas: slack_post/handler.py (classifier-invoked: read today/yesterday
summary.json, post daily summary, conditionally post the batched alert from
details.json) and slack_post/interactions.py (API Gateway: verify Slack
signature, filter details.json, views.open the WO modal within the 3s trigger_id
window). slackio.py centralizes Secrets Manager creds, the SSM dashboard URL,
signature verification, and analytics/ reads — keeping blockkit pure.
Classifier: emit analytics/dt=*/details.json (per-WO index for the modals) and
async-invoke slack-post after the snapshot write (best-effort; a Slack failure
never fails classification).
CDK: slack-post + interactions Lambdas (Docker-bundled slack_sdk), HTTP API on
apm-wo.seahaven.com (wildcard ACM cert + Route53 alias; signature-verified, so
the route is unauthenticated by design), SSM /apm-wo-analysis/grafana-base-url,
and scoped IAM (read analytics/, read the Slack secret + dashboard param;
classifier granted lambda:InvokeFunction on slack-post). Slack creds live in one
Secrets Manager secret apm-wo-analysis/slack-credentials {botToken, signingSecret,
channelId}; cdk.json gains cert/zone/domain context.
WO drill-downs link to Grafana only — no APM deep-links (per decision).
Deliverables for test time: slack/manifest.yaml (app manifest, interactivity
request_url = apm-wo.seahaven.com).
Tests: tests/test_blockkit.py (30 offline cases — deltas, zero-3rd None, <100
blocks under large inputs, modal truncation/overflow, dashboard URL) and Phase 4
assertions in test_pipeline_synth.py (both Lambdas, the API route/domain/alias,
and no broad/write IAM on the Slack roles). 49/49 tests pass; cdk synth green.
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.
Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.
Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.
Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.