* fix(terraform): allow HCP refresh of the log group and parameter The scoped apply role can create those resources, but CloudWatch and SSM list them on a wildcard ARN. DescribeLogGroups and DescribeParameters need that resource. * fix(terraform): let the plan role read bucket website config The S3 provider refreshes GetBucketWebsite. The plan role was denied on the three Forgejo buckets. * fix(terraform): let the plan role read backup object metadata HeadObject on the Lambda zip is s3:GetObject. The plan role only had the bucket ARNs. * fix(terraform): let the plan role read object tags and retention The S3 provider refreshes tagging, ACL, attributes, and Object Lock on the Lambda zip. * fix(terraform): scope plan object reads and restore on the data volume The plan role only needs object reads for the verification zip. Unpacking a dump in /tmp fills the root volume. * fix(terraform): accept dumps that already contain data/forgejo.db Today's archive has no gitea-db.sqlite3 at the root. Copy that file only when it is present. * docs(terraform): keep restore runbook on one bucket and fail closed Glacier and the download use the prod bucket. GCS unpacks on the data volume. Neither path deletes live repos until the dump has a database. * docs(terraform): keep optional restore copies from aborting under set -e if/fi matches user_data.sh. The file comment now says forgejo-services versions also go through the org-account role. * fix(terraform): restore the dump app.ini with the database INTERNAL_TOKEN, JWT_SECRET, and LFS_JWT_SECRET live in that file. A restore that keeps the generated file cannot decrypt the dumped secrets. |
||
|---|---|---|
| .github/workflows | ||
| bin | ||
| lambda/backup-verification | ||
| lib | ||
| scripts | ||
| terraform | ||
| .gitignore | ||
| AGENTS.md | ||
| cdk.context.json | ||
| cdk.json | ||
| package-lock.json | ||
| package.json | ||
| README.md | ||
| tsconfig.json | ||
forgejo
Self-hosted Forgejo git server for archiving GitHub repos and mirroring active ones. The live host is still the mgmt CDK stack until cutover. The replacement is HCP Terraform in seahaven-prod.
HCP Terraform (PLAT-80)
| HCP workspace | forgejo-prod (project seahaven-prod, working directory terraform/) |
| VCS triggers | trigger-patterns = ["terraform/**/*", "lambda/**/*"] |
| Account | seahaven-prod 011934824531 |
| Plan/apply roles | hcptf-forgejo-plan / hcptf-forgejo |
| Apply | Manual. Auto-apply stays off until after DNS cutover and one nightly dump in the new bucket. |
A lambda/-only merge must still queue a run, so the trigger patterns include lambda/**/* as well as terraform/**/*. Docs-only commits do not apply.
The instance attaches to the After Hours VPC (10.70.0.0/16) through workspace variables existing_vpc_id and existing_public_subnet_ids (afterhours-shift-manager outputs vpc_id and public_subnet_ids). The first subnet is the instance availability zone. AMI id is pinned in terraform/variables.tf (ami_id). Do not switch it to most_recent.
DNS stays in the mgmt zone Z06652411XKH89KTZD3XA. Terraform does not own forgejo.seahaven.com. Cutover is an alias flip to the new ALB. Until enable_https is true, the ALB listens on port 80 so a restore can be proved against alb_dns_name without moving the public name. Create the acm_validation_records output in the mgmt zone before setting enable_https.
enable_schedules stays false until cutover so the not-running alarm does not page site-alerts.
Secret values are not in Terraform. IAM uses name-prefix ARNs because hcptf-bootstrap-plan cannot DescribeSecret. These names must exist in seahaven-prod before the instance boots: forgejo/admin-password, forgejo/api-token, forgejo/github-pat, forgejo/gcs-sa-key, forgejo/slack-webhook, forgejo/gcs-transfer-credentials. The GCS transfer user is forgejo-gcs-transfer. Create its access key by hand and store it in forgejo/gcs-transfer-credentials. Do not put the key in Terraform.
First apply uses hcptf-bootstrap / hcptf-bootstrap-plan after scripts/create-hcptf-bootstrap-roles.sh --account prod --allow-workspace forgejo-prod in seahaven-org-baseline. That apply creates IAM and errors on the instance. Retarget TFC_AWS_APPLY_ROLE_ARN and TFC_AWS_PLAN_ROLE_ARN to hcptf-forgejo and hcptf-forgejo-plan, drop --allow-workspace, then apply again. Do not use a project variable set.
Backup bucket names are forgejo-backups-011934824531 and forgejo-backups-replica-011934824531 (us-west-2, Object Lock governance 90 days). The sections below describe the live mgmt host until that cutover.
Architecture
- EC2: t4g.small (arm64), Amazon Linux 2023, 20 GiB gp3 root (no state lives on it) — look up instance ID with:
The AMI is cached in the committedaws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output textcdk.context.json(cachedInContext: true) so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (cdk context --reset <ami key> && cdk synth). - Data volume: standalone 50 GiB gp3 encrypted EBS volume (
RemovalPolicy.RETAIN) mounted at/var/lib/forgejo— sqlite database, repositories, and logs all live here and survive instance replacement and stack deletion. UserData waits for the volume attachment, mounts the existing filesystem (ablkidguard prevents formatting a disk that has one), and tags itforgejo-backup=truefor DLM snapshots. - Restore-on-boot: if the data volume has no database on boot (first boot or total volume loss), UserData automatically downloads the latest S3 dump and restores it before starting the service — volume loss self-heals to ≤24h-old state with no manual steps.
- Network: Private subnet (us-east-1a), behind
seahaven-comALB for SSL termination - DNS:
forgejo.seahaven.com— Route53 alias record pointing to theseahaven-comALB (not a direct A record) - TLS: Wildcard cert on ALB, HTTP internally on port 3000
- Backup: Nightly
forgejo dumpto S3 + EBS snapshots via DLM (see 3-2-1 Backup Strategy) - Admin access: SSM Session Manager (no SSH port exposed)
- CI/CD: Pull requests call the org CI workflows (CDK synth, plus Terraform fmt/validate). CDK deploy on push is frozen.
workflow_dispatchcan still patch the live mgmt host. The replacement workspace isforgejo-prod.
Ports
| Port | Protocol | Source | Purpose |
|---|---|---|---|
| 443 | HTTPS | ALB (public) | Web UI + HTTP git clone |
| 3000 | HTTP | ALB → instance | Internal traffic from ALB |
| 2222 | SSH | VPC + VPN | Git SSH operations |
3-2-1 Backup Strategy
All backups follow a 3-2-1 strategy: 3 copies, 2 storage types, 1 offsite provider.
| Copy | Location | Type | Retention |
|---|---|---|---|
| Live | Standalone EBS data volume, RETAIN (us-east-1) | Block | N/A |
| Near-site | S3 replica (us-west-2) | Object | Archive: indefinite, noncurrent versions: 90d |
| Offsite | GCS forgejo-backups-offsite-seahaven (GCP us-central1) |
Object | 2-year locked retention |
Daily data flow:
| Time (UTC) | Event |
|---|---|
| 05:00 | forgejo dump → s3://forgejo-backups-328440206208/archive/{date}/ |
| ~05:01 | S3 CRR replicates to forgejo-backups-replica-328440206208 (us-west-2) |
| 06:00 | DLM EBS snapshot (30-day retention) |
| 08:00 | Verification Lambda checks all 3 locations, posts to Slack |
| 10:00 | GCS Storage Transfer pulls from S3 to GCS offsite |
S3 source lifecycle: Standard 30d → Glacier (no expiration).
Immutability layers:
- S3 Versioning on both source and replica buckets
- S3 Object Lock (Governance, 90d) on the replica bucket
- GCS Bucket Lock (2yr, irreversible) on the offsite bucket
Verification
The forgejo-backup-verification Lambda runs daily at 08:00 UTC and checks:
- S3 source has a recent dump under
archive/ - S3 replica has replicated the latest dump
- GCS offsite has received the latest transfer
- EBS snapshots exist within the last 48 hours
On the 1st of each month at 09:00 UTC, it runs a restore test: downloads the latest dump, extracts the archive, and runs SQLite integrity checks.
Two CloudWatch alarms watch the verification Lambda (both notify the site-alerts SNS topic, ALARM action only):
forgejo-backup-verification-errors— Errors ≥ 1 in an hour, missing data = not breaching ("when it runs, did it fail")forgejo-backup-verification-not-running— Invocations < 1 over 24h, missing data = breaching ("did it run at all")
Manual backup
sudo /usr/local/bin/forgejo-backup.sh
Restore from S3
For backups older than 30 days (Glacier), restore the object first:
aws s3api restore-object --bucket forgejo-backups-011934824531 \
--key "archive/<date>/forgejo-<date>.tar.gz" \
--restore-request '{"Days":7,"GlacierJobParameters":{"Tier":"Standard"}}'
# Wait ~3-5 hours for restore to complete, then:
Download and restore. The copy does not start until the dump contains a database, so a failed download leaves the live data alone.
set -euo pipefail
rm -rf /var/lib/forgejo/.restore && mkdir -p /var/lib/forgejo/.restore
aws s3 cp s3://forgejo-backups-011934824531/archive/<date>/forgejo-<date>.tar.gz - --no-progress \
| tar -xz -C /var/lib/forgejo/.restore
if [ ! -s /var/lib/forgejo/.restore/gitea-db.sqlite3 ] && [ ! -s /var/lib/forgejo/.restore/data/forgejo.db ]; then
echo "Dump has no database" >&2
exit 1
fi
systemctl stop forgejo
if [ -d /var/lib/forgejo/.restore/data ]; then
cp -a /var/lib/forgejo/.restore/data/. /var/lib/forgejo/data/
fi
rm -rf /var/lib/forgejo/data/repositories
mkdir -p /var/lib/forgejo/data/repositories
if [ -d /var/lib/forgejo/.restore/repos ]; then
cp -a /var/lib/forgejo/.restore/repos/. /var/lib/forgejo/data/repositories/
fi
if [ -f /var/lib/forgejo/.restore/gitea-db.sqlite3 ]; then
cp /var/lib/forgejo/.restore/gitea-db.sqlite3 /var/lib/forgejo/data/forgejo.db
fi
if [ -d /var/lib/forgejo/.restore/lfs ]; then
mkdir -p /var/lib/forgejo/data/lfs
cp -a /var/lib/forgejo/.restore/lfs/. /var/lib/forgejo/data/lfs/
fi
if [ -d /var/lib/forgejo/.restore/custom ]; then
cp -a /var/lib/forgejo/.restore/custom/. /var/lib/forgejo/custom/
fi
if [ ! -s /var/lib/forgejo/data/forgejo.db ]; then
echo "Restore did not produce /var/lib/forgejo/data/forgejo.db" >&2
exit 1
fi
if [ -f /var/lib/forgejo/.restore/app.ini ]; then
cp /var/lib/forgejo/.restore/app.ini /etc/forgejo/app.ini
chown root:forgejo /etc/forgejo/app.ini
chmod 660 /etc/forgejo/app.ini
fi
chown -R forgejo:forgejo /var/lib/forgejo
systemctl start forgejo
rm -rf /var/lib/forgejo/.restore
Restore from GCS (disaster recovery)
This path does not read S3. It unpacks the offsite object on the data volume and uses the same copy order as the S3 restore.
set -euo pipefail
gcloud config set project sea-haven-backups
rm -rf /var/lib/forgejo/.restore && mkdir -p /var/lib/forgejo/.restore
gsutil cp gs://forgejo-backups-offsite-seahaven/archive/<date>/forgejo-<date>.tar.gz - \
| tar -xz -C /var/lib/forgejo/.restore
if [ ! -s /var/lib/forgejo/.restore/gitea-db.sqlite3 ] && [ ! -s /var/lib/forgejo/.restore/data/forgejo.db ]; then
echo "Dump has no database" >&2
exit 1
fi
systemctl stop forgejo
if [ -d /var/lib/forgejo/.restore/data ]; then
cp -a /var/lib/forgejo/.restore/data/. /var/lib/forgejo/data/
fi
rm -rf /var/lib/forgejo/data/repositories
mkdir -p /var/lib/forgejo/data/repositories
if [ -d /var/lib/forgejo/.restore/repos ]; then
cp -a /var/lib/forgejo/.restore/repos/. /var/lib/forgejo/data/repositories/
fi
if [ -f /var/lib/forgejo/.restore/gitea-db.sqlite3 ]; then
cp /var/lib/forgejo/.restore/gitea-db.sqlite3 /var/lib/forgejo/data/forgejo.db
fi
if [ -d /var/lib/forgejo/.restore/lfs ]; then
mkdir -p /var/lib/forgejo/data/lfs
cp -a /var/lib/forgejo/.restore/lfs/. /var/lib/forgejo/data/lfs/
fi
if [ -d /var/lib/forgejo/.restore/custom ]; then
cp -a /var/lib/forgejo/.restore/custom/. /var/lib/forgejo/custom/
fi
if [ ! -s /var/lib/forgejo/data/forgejo.db ]; then
echo "Restore did not produce /var/lib/forgejo/data/forgejo.db" >&2
exit 1
fi
if [ -f /var/lib/forgejo/.restore/app.ini ]; then
cp /var/lib/forgejo/.restore/app.ini /etc/forgejo/app.ini
chown root:forgejo /etc/forgejo/app.ini
chmod 660 /etc/forgejo/app.ini
fi
chown -R forgejo:forgejo /var/lib/forgejo
systemctl start forgejo
rm -rf /var/lib/forgejo/.restore
Autodiscovery
An hourly cron job checks the Sea-Haven-Industries GitHub org for new repositories and mirrors them into Forgejo automatically.
- Active repos are created as mirrors (ongoing sync).
- Archived repos are created as static one-time imports.
- Script:
/usr/local/bin/forgejo-autodiscover.sh - Log:
/var/log/forgejo-autodiscover.log
Token Refresh
A daily cron at 4:30 UTC reads the GitHub PAT from Secrets Manager (forgejo/github-pat) and updates the git remote URL on every mirror repository so credentials stay current.
- Script:
/usr/local/bin/forgejo-refresh-tokens.sh
PAT Rotation
The GitHub personal access token used for mirroring is a fine-grained PAT scoped to Sea-Haven-Industries with Contents: Read-only permissions and a 1-year expiration. It is stored in Secrets Manager at forgejo/github-pat.
To rotate:
- Create a new fine-grained PAT on GitHub with the same scope.
- Update the secret value in Secrets Manager (
forgejo/github-pat). - The daily token-refresh cron will pick it up automatically.
To force immediate propagation:
sudo /usr/local/bin/forgejo-refresh-tokens.sh
Secrets Manager
| Secret | Purpose |
|---|---|
forgejo/admin-password |
Forgejo admin user password |
forgejo/api-token |
Forgejo API token (used by autodiscovery and token refresh scripts) |
forgejo/github-pat |
GitHub fine-grained PAT for mirroring |
forgejo/gcs-sa-key |
GCP service account key for offsite backup verification |
forgejo/gcs-transfer-credentials |
AWS IAM credentials for GCS Storage Transfer Service |
forgejo/slack-webhook |
Slack webhook URL for backup verification alerts |
First-time setup
After the stack deploys, connect via SSM and create the admin user:
INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
aws ssm start-session --target "$INSTANCE_ID"
sudo -u forgejo /usr/local/bin/forgejo admin user create \
--admin \
--username adam \
--password '<password>' \
--email adam@seahavenind.com \
--config /etc/forgejo/app.ini
Admin password is stored in Secrets Manager at forgejo/admin-password.
Access the web UI at https://forgejo.seahaven.com.
Migrating repos from GitHub
Archived repos (one-time import)
In the Forgejo web UI: New Migration → GitHub → paste the GitHub repo URL. Use a GitHub personal access token for private repos. These are full imports (code, issues, PRs, releases).
Active repos (mirror sync)
Same migration flow, but check This Repository Will Be A Mirror. Forgejo polls GitHub hourly (DEFAULT_INTERVAL = 1h in app.ini) and keeps the mirror in sync.
GCP Offsite Setup (one-time)
Run the setup script to create the GCS offsite bucket, service account, and store credentials:
./scripts/gcp-setup.sh
This creates the sea-haven-backups GCP project with a locked-retention GCS bucket. After running, configure the Storage Transfer job in the GCP Console using the AWS credentials from forgejo/gcs-transfer-credentials.
Infrastructure as Code (CDK)
All AWS infrastructure is defined as an AWS CDK v2 app written in TypeScript. There is no console-managed infrastructure — every resource in the sections above is synthesized from this repo.
Project layout
| Path | Purpose |
|---|---|
bin/app.ts |
CDK app entry point. Instantiates both stacks and wires the dependency between them. |
lib/forgejo-stack.ts |
Main stack (forgejo, us-east-1): EC2 instance, security groups, IAM, ALB target/DNS, secrets, EBS data volume, DLM snapshots, and the backup-verification construct. Holds the FORGEJO_VERSION constant. |
lib/forgejo-replica-stack.ts |
Replica stack (forgejo-replica, us-west-2): the S3 CRR replica bucket with versioning, Object Lock (Governance, 90d), and Glacier lifecycle. |
lib/constructs/backup-verification.ts |
BackupVerification construct — the daily verification Lambda, its schedule, and CloudWatch alarms. |
lambda/backup-verification/ |
Python 3.12 handler (app.py + requirements.txt) bundled via @aws-cdk/aws-lambda-python-alpha. |
scripts/gcp-setup.sh |
One-time GCP offsite bucket/service-account provisioning (see GCP Offsite Setup). |
cdk.json |
CDK configuration — the app command and context feature flags. |
cdk.context.json |
Cached context lookups (VPC + AMI), committed so synth is deterministic. |
Stacks
The app defines two stacks, both pinned to account 328440206208 with explicit kebab-case stackNames:
forgejo-replica(us-west-2) — the S3 replica bucket. Deployed first; exports the replica bucket ARN and name.forgejo(us-east-1) — the main stack. Consumes the replica bucket ARN/name and declares a dependency onforgejo-replica, so CDK always deploys the replica first.
cdk.json
cdk.json is the CDK entry configuration, committed to the repo:
app:tsx bin/app.ts— runs the TypeScript app directly viatsx(a pinneddevDependency), so no separatetscbuild step is needed for synth.watch: include/exclude globs forcdk watch.context: CDK feature flags (e.g.@aws-cdk/core:target-partitions). Cached lookup context (VPC subnets, the AL2023 arm64 AMI) lives separately incdk.context.json—cdk.jsonholds only feature flags.
aws-cdk-lib is pinned to an exact version (no ^/~) per the CDK version policy; Dependabot keeps it current.
Common commands
npm install
npm run synth # cdk synth — emit CloudFormation without deploying
npm run diff # cdk diff — diff local app against deployed stacks
npm run deploy # cdk deploy — deploy (pass -- --all for both stacks)
Deployment
Pull requests call the org reusables from .github/workflows/ci.yaml: CDK synth, and Terraform fmt / init -backend=false / validate on terraform/. A merge to main does not deploy. The CDK workflow (.github/workflows/deploy.yaml) runs only on workflow_dispatch, and that path still targets the live mgmt stacks.
New infrastructure is the terraform/ root, applied from HCP workspace forgejo-prod. See HCP Terraform (PLAT-80). Do not cdk deploy this repo into seahaven-prod.
Post-deploy: store Slack webhook
Store the Slack webhook URL for backup verification alerts:
aws secretsmanager create-secret --name forgejo/slack-webhook \
--secret-string "https://hooks.slack.com/services/YOUR/WEBHOOK/URL" \
--region us-east-1
Updating Forgejo
Update the FORGEJO_VERSION constant in lib/forgejo-stack.ts and deploy. This replaces the instance, so ensure the latest EBS snapshot is available for data recovery if needed. Alternatively, update in-place via SSM:
INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
aws ssm start-session --target "$INSTANCE_ID"
sudo systemctl stop forgejo
sudo curl -Lo /usr/local/bin/forgejo "https://codeberg.org/forgejo/forgejo/releases/download/v<NEW_VERSION>/forgejo-<NEW_VERSION>-linux-arm64"
sudo chmod +x /usr/local/bin/forgejo
sudo systemctl start forgejo