From 891dc0831b5ba101912b7940964f28e5e95227ab Mon Sep 17 00:00:00 2001 From: Adam Moussa <166072409+amoussa1229@users.noreply.github.com> Date: Fri, 5 Jun 2026 15:54:26 -0400 Subject: [PATCH] docs: document persistent data volume, restore-on-boot, cached AMI, alarms (#25) Storage architecture changed 2026-06-05: state moved off the root volume onto a standalone RETAIN data volume with automatic S3 restore on empty boot. --- README.md | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 1db4f1a..4ebfc99 100644 --- a/README.md +++ b/README.md @@ -4,10 +4,13 @@ Self-hosted Forgejo git server for archiving GitHub repos and mirroring active o ## Architecture -- **EC2**: t4g.small (arm64), Amazon Linux 2023, 50GB gp3 EBS — look up instance ID with: +- **EC2**: t4g.small (arm64), Amazon Linux 2023, 20 GiB gp3 root (no state lives on it) — look up instance ID with: ``` aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text ``` + The AMI is cached in the committed `cdk.context.json` (`cachedInContext: true`) so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (`cdk context --reset && cdk synth`). +- **Data volume**: standalone 50 GiB gp3 encrypted EBS volume (`RemovalPolicy.RETAIN`) mounted at `/var/lib/forgejo` — sqlite database, repositories, and logs all live here and **survive instance replacement and stack deletion**. UserData waits for the volume attachment, mounts the existing filesystem (a `blkid` guard prevents formatting a disk that has one), and tags it `forgejo-backup=true` for DLM snapshots. +- **Restore-on-boot**: if the data volume has no database on boot (first boot or total volume loss), UserData automatically downloads the latest S3 dump and restores it before starting the service — volume loss self-heals to ≤24h-old state with no manual steps. - **Network**: Private subnet (us-east-1a), behind `seahaven-com` ALB for SSL termination - **DNS**: `forgejo.seahaven.com` — Route53 alias record pointing to the `seahaven-com` ALB (not a direct A record) - **TLS**: Wildcard cert on ALB, HTTP internally on port 3000 @@ -29,7 +32,7 @@ All backups follow a 3-2-1 strategy: 3 copies, 2 storage types, 1 offsite provid | Copy | Location | Type | Retention | |------|----------|------|-----------| -| Live | EBS volume (us-east-1) | Block | N/A | +| Live | Standalone EBS data volume, RETAIN (us-east-1) | Block | N/A | | Near-site | S3 replica (us-west-2) | Object | Archive: indefinite, noncurrent versions: 90d | | Offsite | GCS `forgejo-backups-offsite-seahaven` (GCP us-central1) | Object | 2-year locked retention | @@ -60,6 +63,10 @@ The `forgejo-backup-verification` Lambda runs daily at 08:00 UTC and checks: On the 1st of each month at 09:00 UTC, it runs a restore test: downloads the latest dump, extracts the archive, and runs SQLite integrity checks. +Two CloudWatch alarms watch the verification Lambda (both notify the `site-alerts` SNS topic, ALARM action only): +- `forgejo-backup-verification-errors` — Errors ≥ 1 in an hour, missing data = not breaching ("when it runs, did it fail") +- `forgejo-backup-verification-not-running` — Invocations < 1 over 24h, missing data = breaching ("did it run at all") + ### Manual backup ```bash