Scaffold apm-wo-analysis repository
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.
Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.
Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.
Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.
2026-05-28 16:10:20 -04:00
|
|
|
"""Pipeline stack: S3, classifier + slack-post Lambdas, Glue, Athena, IAM.
|
|
|
|
|
|
|
|
|
|
Scaffold — the exports bucket (Phase 1 of docs/BUILD.md) is included so the
|
|
|
|
|
stack synthesizes to something real. The classifier Lambda (Phase 2), Glue
|
|
|
|
|
database + Athena workgroup with partition projection (Phase 3), and the
|
|
|
|
|
slack-post Lambda + IAM (Phase 4) are added in their respective phases.
|
|
|
|
|
|
|
|
|
|
Lambda defaults when added: Python 3.12, ARM64, explicit LogGroup with 60-day
|
|
|
|
|
retention. No DynamoDB — this is an S3 + Athena analytics workload (see CLAUDE.md).
|
|
|
|
|
"""
|
|
|
|
|
|
|
|
|
|
from aws_cdk import (
|
|
|
|
|
Duration,
|
|
|
|
|
RemovalPolicy,
|
|
|
|
|
Stack,
|
|
|
|
|
)
|
Add drop-folder ingestion and scoped uploader IAM user
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.
Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.
The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
2026-05-28 16:32:56 -04:00
|
|
|
from aws_cdk import (
|
|
|
|
|
aws_iam as iam,
|
|
|
|
|
)
|
Scaffold apm-wo-analysis repository
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.
Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.
Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.
Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.
2026-05-28 16:10:20 -04:00
|
|
|
from aws_cdk import (
|
|
|
|
|
aws_s3 as s3,
|
|
|
|
|
)
|
|
|
|
|
from constructs import Construct
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
class PipelineStack(Stack):
|
|
|
|
|
def __init__(self, scope: Construct, construct_id: str, **kwargs) -> None:
|
|
|
|
|
super().__init__(scope, construct_id, **kwargs)
|
|
|
|
|
|
|
|
|
|
# Phase 1 — single exports bucket.
|
|
|
|
|
# Prefixes: raw/ (incoming), analytics/ (per-WO snapshots), athena-results/.
|
|
|
|
|
self.exports_bucket = s3.Bucket(
|
|
|
|
|
self,
|
|
|
|
|
"Exports",
|
|
|
|
|
bucket_name=f"apm-wo-analysis-exports-{self.account}",
|
|
|
|
|
encryption=s3.BucketEncryption.S3_MANAGED,
|
|
|
|
|
block_public_access=s3.BlockPublicAccess.BLOCK_ALL,
|
|
|
|
|
enforce_ssl=True,
|
|
|
|
|
removal_policy=RemovalPolicy.RETAIN,
|
|
|
|
|
lifecycle_rules=[
|
|
|
|
|
s3.LifecycleRule(
|
|
|
|
|
id="expire-raw-exports",
|
|
|
|
|
prefix="raw/",
|
|
|
|
|
expiration=Duration.days(90),
|
|
|
|
|
)
|
|
|
|
|
],
|
|
|
|
|
)
|
|
|
|
|
|
Add drop-folder ingestion and scoped uploader IAM user
Complete Phase 1 ingestion. Add a least-privilege IAM user
(apm-wo-drop-uploader) to the pipeline stack, scoped to s3:PutObject
on the raw/ prefix only — the local launchd uploader authenticates as
this user via a dedicated profile, so a laptop credential leak cannot
read, list, or touch the analytics data.
Replace the scaffold uploader stub with the hardened stampli-pattern
script (lockfile, logging, timestamped archive, notifications, settle
delay) and align names to the convention (~/apm-wo-drop, ~/.local/bin,
com.seahaven.apm-wo-uploader). The plist sets PATH/HOME because launchd
runs with a stripped environment and otherwise cannot find aws.
The exports bucket already shipped in the Phase 0 scaffold, so the code
delta here is the uploader identity and tooling.
2026-05-28 16:32:56 -04:00
|
|
|
# Phase 1 — least-privilege identity for the local drop-folder uploader.
|
|
|
|
|
# Scoped to s3:PutObject on raw/* only. The access key is created
|
|
|
|
|
# out-of-band (aws iam create-access-key) and stored in the local
|
|
|
|
|
# ~/.aws/credentials profile `apm-wo-drop` — never in CloudFormation.
|
|
|
|
|
self.drop_uploader = iam.User(
|
|
|
|
|
self, "DropUploader", user_name="apm-wo-drop-uploader"
|
|
|
|
|
)
|
|
|
|
|
self.drop_uploader.add_to_policy(
|
|
|
|
|
iam.PolicyStatement(
|
|
|
|
|
sid="PutRawExportsOnly",
|
|
|
|
|
actions=["s3:PutObject"],
|
|
|
|
|
resources=[self.exports_bucket.arn_for_objects("raw/*")],
|
|
|
|
|
)
|
|
|
|
|
)
|
|
|
|
|
|
Scaffold apm-wo-analysis repository
Stand up the Phase 0 CDK scaffold for the daily APM work-order
analysis pipeline: two-stack CDK app (pipeline + grafana), classifier
and slack-post Lambda packages, dashboards-as-code, the local
drop-folder uploader, and a classifier smoke-test placeholder.
Wire CI/CD to the org reusable workflows: ci.yaml -> ci-python-sam
(ruff + cdk synth) and deploy.yaml -> cd-cdk (OIDC, cdk deploy --all).
Pin aws-cdk-lib==2.253.1; Lambdas target Python 3.12 / arm64.
Rewrite .gitignore to the org Python-CDK standard so the source-of-
truth files (CLAUDE.md, docs/, .claude/agents) are tracked while build
artifacts (.venv, cdk.out, caches) stay ignored.
Domain logic, stack resources, and dashboards are stubbed and filled
in across Phases 1-5 (docs/BUILD.md). cdk synth is green for both
stacks; ruff check/format pass.
2026-05-28 16:10:20 -04:00
|
|
|
# Phase 2 — classifier Lambda, S3-triggered on the raw/ prefix. TODO
|
|
|
|
|
# Phase 3 — Glue database `apm_wo_analysis` + Athena workgroup
|
|
|
|
|
# (partition projection on dt; no crawler). TODO
|
|
|
|
|
# Phase 4 — slack-post Lambda + scoped IAM. TODO
|