forgejo/README.md
Adam Moussa 5ed1db788e
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5)
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite

Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
  Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
  with 2-year irreversible retention lock

Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.

* Enable QEMU in CI for arm64 Lambda Docker builds

* Commit cdk.context.json for CI synth without AWS credentials

Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.

* Fix GCP project ID to sea-haven-backups

* Address code review findings for backup verification

Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README

Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing

* Fix backup strategy bug findings

* Handle SQL text dumps separately from binary SQLite in restore test

Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.

* Fix GCS backup check: align staleness cutoff and add size validation

GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.

* Rename SECRET_ARN env vars to SECRET_NAME to match actual values

* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule

* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt

* Fix restore runbook: trailing-dot cp idiom and Glacier restore step

* Rename GCS service account to match read-only permissions

* Add 4 GiB ephemeral storage to verification Lambda

Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.

* Replace hardcoded instance ID in README with CloudFormation lookup

The instance ID changes on every instance replacement (version
upgrades, stack updates). Using a dynamic query prevents stale
references and removes a manual update step from the deploy process.

* Read backup S3 prefix from SSM parameter at runtime

Adds /forgejo/backup-s3-prefix SSM parameter (value: archive)
and updates the backup script to fetch it instead of hardcoding
the prefix. Eliminates the manual post-deploy sed step.

* Address cross-review findings for backup verification

- Add size guard before downloading dump in restore test (3.5 GB cap)
- Use paginator for list_objects_v2 in S3 checks and restore test
- Remove unnecessary overrideLogicalId on GcsTransferCredentials secret
- Pass explicit { mode: "daily" } to daily EventBridge rule target
- Add fallback for SSM parameter fetch in backup script
- Export replica bucket ARN/name from replica stack, consume via props

* Add CloudWatch alarm for backup verification Lambda errors

Fires on any Lambda error and on missing data (missed schedule).
Catches silent failures where the Slack notification never fires.

* Add .env to .gitignore

Required by org CI conventions check.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00

8.4 KiB

forgejo

Self-hosted Forgejo git server for archiving GitHub repos and mirroring active ones. Runs on a single EC2 instance within the Sea Haven VPC, fronted by the seahaven-com ALB for HTTPS.

Architecture

  • EC2: t4g.small (arm64), Amazon Linux 2023, 50GB gp3 EBS — look up instance ID with:
    aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text
    
  • Network: Private subnet (us-east-1a), behind seahaven-com ALB for SSL termination
  • DNS: forgejo.seahaven.com — Route53 alias record pointing to the seahaven-com ALB (not a direct A record)
  • TLS: Wildcard cert on ALB, HTTP internally on port 3000
  • Backup: Nightly forgejo dump to S3 + EBS snapshots via DLM (see 3-2-1 Backup Strategy)
  • Admin access: SSM Session Manager (no SSH port exposed)
  • CI/CD: GitHub Actions with OIDC role githubdeploy-forgejo

Ports

Port Protocol Source Purpose
443 HTTPS ALB (public) Web UI + HTTP git clone
3000 HTTP ALB → instance Internal traffic from ALB
2222 SSH VPC + VPN Git SSH operations

3-2-1 Backup Strategy

All backups follow a 3-2-1 strategy: 3 copies, 2 storage types, 1 offsite provider.

Copy Location Type Retention
Live EBS volume (us-east-1) Block N/A
Near-site S3 replica (us-west-2) Object Archive: indefinite, noncurrent versions: 90d
Offsite GCS forgejo-backups-offsite-seahaven (GCP us-central1) Object 2-year locked retention

Daily data flow:

Time (UTC) Event
05:00 forgejo dump → s3://forgejo-backups-328440206208/archive/{date}/
~05:01 S3 CRR replicates to forgejo-backups-replica-328440206208 (us-west-2)
06:00 DLM EBS snapshot (30-day retention)
08:00 Verification Lambda checks all 3 locations, posts to Slack
10:00 GCS Storage Transfer pulls from S3 to GCS offsite

S3 source lifecycle: Standard 30d → Glacier (no expiration).

Immutability layers:

  • S3 Versioning on both source and replica buckets
  • S3 Object Lock (Governance, 90d) on the replica bucket
  • GCS Bucket Lock (2yr, irreversible) on the offsite bucket

Verification

The forgejo-backup-verification Lambda runs daily at 08:00 UTC and checks:

  1. S3 source has a recent dump under archive/
  2. S3 replica has replicated the latest dump
  3. GCS offsite has received the latest transfer
  4. EBS snapshots exist within the last 48 hours

On the 1st of each month at 09:00 UTC, it runs a restore test: downloads the latest dump, extracts the archive, and runs SQLite integrity checks.

Manual backup

sudo /usr/local/bin/forgejo-backup.sh

Restore from S3

For backups older than 30 days (Glacier), restore the object first:

aws s3api restore-object --bucket forgejo-backups-328440206208 \
  --key "archive/<date>/forgejo-<date>.tar.gz" \
  --restore-request '{"Days":7,"GlacierJobParameters":{"Tier":"Standard"}}'
# Wait ~3-5 hours for restore to complete, then:

Download and restore:

aws s3 cp s3://forgejo-backups-328440206208/archive/<date>/forgejo-<date>.tar.gz /tmp/
systemctl stop forgejo
mkdir -p /tmp/forgejo-restore && tar -xzf /tmp/forgejo-<date>.tar.gz -C /tmp/forgejo-restore
cd /tmp/forgejo-restore
cp app.ini /etc/forgejo/app.ini
cp gitea-db.sqlite3 /var/lib/forgejo/data/forgejo.db
rm -rf /var/lib/forgejo/data/repositories
cp -a repos /var/lib/forgejo/data/repositories
cp -a data/. /var/lib/forgejo/data/
[ -d lfs ] && cp -a lfs/. /var/lib/forgejo/data/lfs/
[ -d custom ] && cp -a custom/. /var/lib/forgejo/custom/
chown -R forgejo:forgejo /var/lib/forgejo /etc/forgejo/app.ini
systemctl start forgejo
rm -rf /tmp/forgejo-restore /tmp/forgejo-<date>.tar.gz

Restore from GCS (disaster recovery)

gcloud config set project sea-haven-backups
gsutil cp gs://forgejo-backups-offsite-seahaven/archive/<date>/forgejo-<date>.tar.gz /tmp/
# Then follow the same restore steps as S3 above

Autodiscovery

An hourly cron job checks the Sea-Haven-Industries GitHub org for new repositories and mirrors them into Forgejo automatically.

  • Active repos are created as mirrors (ongoing sync).
  • Archived repos are created as static one-time imports.
  • Script: /usr/local/bin/forgejo-autodiscover.sh
  • Log: /var/log/forgejo-autodiscover.log

Token Refresh

A daily cron at 4:30 UTC reads the GitHub PAT from Secrets Manager (forgejo/github-pat) and updates the git remote URL on every mirror repository so credentials stay current.

  • Script: /usr/local/bin/forgejo-refresh-tokens.sh

PAT Rotation

The GitHub personal access token used for mirroring is a fine-grained PAT scoped to Sea-Haven-Industries with Contents: Read-only permissions and a 1-year expiration. It is stored in Secrets Manager at forgejo/github-pat.

To rotate:

  1. Create a new fine-grained PAT on GitHub with the same scope.
  2. Update the secret value in Secrets Manager (forgejo/github-pat).
  3. The daily token-refresh cron will pick it up automatically.

To force immediate propagation:

sudo /usr/local/bin/forgejo-refresh-tokens.sh

Secrets Manager

Secret Purpose
forgejo/admin-password Forgejo admin user password
forgejo/api-token Forgejo API token (used by autodiscovery and token refresh scripts)
forgejo/github-pat GitHub fine-grained PAT for mirroring
forgejo/gcs-sa-key GCP service account key for offsite backup verification
forgejo/gcs-transfer-credentials AWS IAM credentials for GCS Storage Transfer Service
forgejo/slack-webhook Slack webhook URL for backup verification alerts

First-time setup

After the stack deploys, connect via SSM and create the admin user:

INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
aws ssm start-session --target "$INSTANCE_ID"

sudo -u forgejo /usr/local/bin/forgejo admin user create \
  --admin \
  --username adam \
  --password '<password>' \
  --email adam@seahavenind.com \
  --config /etc/forgejo/app.ini

Admin password is stored in Secrets Manager at forgejo/admin-password.

Access the web UI at https://forgejo.seahaven.com.

Migrating repos from GitHub

Archived repos (one-time import)

In the Forgejo web UI: New Migration → GitHub → paste the GitHub repo URL. Use a GitHub personal access token for private repos. These are full imports (code, issues, PRs, releases).

Active repos (mirror sync)

Same migration flow, but check This Repository Will Be A Mirror. Forgejo polls GitHub hourly (DEFAULT_INTERVAL = 1h in app.ini) and keeps the mirror in sync.

GCP Offsite Setup (one-time)

Run the setup script to create the GCS offsite bucket, service account, and store credentials:

./scripts/gcp-setup.sh

This creates the sea-haven-backups GCP project with a locked-retention GCS bucket. After running, configure the Storage Transfer job in the GCP Console using the AWS credentials from forgejo/gcs-transfer-credentials.

Deployment

npm install
npx cdk deploy --all

This deploys two stacks:

  • forgejo-replica (us-west-2) — S3 replica bucket with Object Lock
  • forgejo (us-east-1) — main stack with Forgejo instance, CRR, and verification Lambda

CI/CD is handled by GitHub Actions — PRs run CI, merges to main deploy via the reusable CDK workflow.

Post-deploy: store Slack webhook

Store the Slack webhook URL for backup verification alerts:

aws secretsmanager create-secret --name forgejo/slack-webhook \
  --secret-string "https://hooks.slack.com/services/YOUR/WEBHOOK/URL" \
  --region us-east-1

Updating Forgejo

Update the FORGEJO_VERSION constant in lib/forgejo-stack.ts and deploy. This replaces the instance, so ensure the latest EBS snapshot is available for data recovery if needed. Alternatively, update in-place via SSM:

INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
aws ssm start-session --target "$INSTANCE_ID"

sudo systemctl stop forgejo
sudo curl -Lo /usr/local/bin/forgejo "https://codeberg.org/forgejo/forgejo/releases/download/v<NEW_VERSION>/forgejo-<NEW_VERSION>-linux-arm64"
sudo chmod +x /usr/local/bin/forgejo
sudo systemctl start forgejo