docs: document persistent data volume, restore-on-boot, cached AMI, alarms (#25)
Some checks failed
Deploy / deploy (push) Has been cancelled

Storage architecture changed 2026-06-05: state moved off the root
volume onto a standalone RETAIN data volume with automatic S3 restore
on empty boot.
This commit is contained in:
Adam Moussa 2026-06-05 15:54:26 -04:00 • committed by GitHub
parent b75b4cc130
commit 891dc0831b
No known key found for this signature in database
GPG key ID: B5690EEEBB952194

View file

@ -4,10 +4,13 @@ Self-hosted Forgejo git server for archiving GitHub repos and mirroring active o
## Architecture
- **EC2**: t4g.small (arm64), Amazon Linux 2023, 50GB gp3 EBS — look up instance ID with:
- **EC2**: t4g.small (arm64), Amazon Linux 2023, 20 GiB gp3 root (no state lives on it) — look up instance ID with:
```
aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text
```
The AMI is cached in the committed `cdk.context.json` (`cachedInContext: true`) so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (`cdk context --reset <ami key> && cdk synth`).
- **Data volume**: standalone 50 GiB gp3 encrypted EBS volume (`RemovalPolicy.RETAIN`) mounted at `/var/lib/forgejo` — sqlite database, repositories, and logs all live here and **survive instance replacement and stack deletion**. UserData waits for the volume attachment, mounts the existing filesystem (a `blkid` guard prevents formatting a disk that has one), and tags it `forgejo-backup=true` for DLM snapshots.
- **Restore-on-boot**: if the data volume has no database on boot (first boot or total volume loss), UserData automatically downloads the latest S3 dump and restores it before starting the service — volume loss self-heals to ≤24h-old state with no manual steps.
- **Network**: Private subnet (us-east-1a), behind `seahaven-com` ALB for SSL termination
- **DNS**: `forgejo.seahaven.com` — Route53 alias record pointing to the `seahaven-com` ALB (not a direct A record)
- **TLS**: Wildcard cert on ALB, HTTP internally on port 3000
@ -29,7 +32,7 @@ All backups follow a 3-2-1 strategy: 3 copies, 2 storage types, 1 offsite provid
| Copy | Location | Type | Retention |
|------|----------|------|-----------|
| Live | EBS volume (us-east-1) | Block | N/A |
| Live | Standalone EBS data volume, RETAIN (us-east-1) | Block | N/A |
| Near-site | S3 replica (us-west-2) | Object | Archive: indefinite, noncurrent versions: 90d |
| Offsite | GCS `forgejo-backups-offsite-seahaven` (GCP us-central1) | Object | 2-year locked retention |
@ -60,6 +63,10 @@ The `forgejo-backup-verification` Lambda runs daily at 08:00 UTC and checks:
On the 1st of each month at 09:00 UTC, it runs a restore test: downloads the latest dump, extracts the archive, and runs SQLite integrity checks.
Two CloudWatch alarms watch the verification Lambda (both notify the `site-alerts` SNS topic, ALARM action only):
- `forgejo-backup-verification-errors` — Errors ≥ 1 in an hour, missing data = not breaching ("when it runs, did it fail")
- `forgejo-backup-verification-not-running` — Invocations < 1 over 24h, missing data = breaching ("did it run at all")
### Manual backup
```bash