mirror of
https://github.com/Sea-Haven-Industries/forgejo.git
synced 2026-09-30 15:33:11 +00:00
* feat(terraform): migrate forgejo to HCP Terraform Move the prod host onto workspace forgejo-prod in the After Hours VPC and freeze CDK push deploys so cutover can happen without applying into seahaven-prod. * fix(terraform): restore forgejo.db from the dump tarball Boot restore was copying data/ and repos/ and leaving the sqlite file at the archive root, so a volume restore started Forgejo with no database.
296 lines
16 KiB
Markdown
296 lines
16 KiB
Markdown
# forgejo
|
|
|
|

|
|

|
|

|
|

|
|
|
|
Self-hosted Forgejo git server for archiving GitHub repos and mirroring active ones. The live host is still the mgmt CDK stack until cutover. The replacement is HCP Terraform in seahaven-prod.
|
|
|
|
## HCP Terraform (PLAT-80)
|
|
|
|
| | |
|
|
|---|---|
|
|
| HCP workspace | `forgejo-prod` (project `seahaven-prod`, working directory `terraform/`) |
|
|
| VCS triggers | `trigger-patterns = ["terraform/**/*", "lambda/**/*"]` |
|
|
| Account | seahaven-prod `011934824531` |
|
|
| Plan/apply roles | `hcptf-forgejo-plan` / `hcptf-forgejo` |
|
|
| Apply | Manual. Auto-apply stays off until after DNS cutover and one nightly dump in the new bucket. |
|
|
|
|
A `lambda/`-only merge must still queue a run, so the trigger patterns include `lambda/**/*` as well as `terraform/**/*`. Docs-only commits do not apply.
|
|
|
|
The instance attaches to the After Hours VPC (`10.70.0.0/16`) through workspace variables `existing_vpc_id` and `existing_public_subnet_ids` (afterhours-shift-manager outputs `vpc_id` and `public_subnet_ids`). The first subnet is the instance availability zone. AMI id is pinned in `terraform/variables.tf` (`ami_id`). Do not switch it to `most_recent`.
|
|
|
|
DNS stays in the mgmt zone `Z06652411XKH89KTZD3XA`. Terraform does not own `forgejo.seahaven.com`. Cutover is an alias flip to the new ALB. Until `enable_https` is true, the ALB listens on port 80 so a restore can be proved against `alb_dns_name` without moving the public name. Create the `acm_validation_records` output in the mgmt zone before setting `enable_https`.
|
|
|
|
`enable_schedules` stays false until cutover so the not-running alarm does not page `site-alerts`.
|
|
|
|
Secret values are not in Terraform. IAM uses name-prefix ARNs because `hcptf-bootstrap-plan` cannot `DescribeSecret`. These names must exist in seahaven-prod before the instance boots: `forgejo/admin-password`, `forgejo/api-token`, `forgejo/github-pat`, `forgejo/gcs-sa-key`, `forgejo/slack-webhook`, `forgejo/gcs-transfer-credentials`. The GCS transfer user is `forgejo-gcs-transfer`. Create its access key by hand and store it in `forgejo/gcs-transfer-credentials`. Do not put the key in Terraform.
|
|
|
|
First apply uses `hcptf-bootstrap` / `hcptf-bootstrap-plan` after `scripts/create-hcptf-bootstrap-roles.sh --account prod --allow-workspace forgejo-prod` in seahaven-org-baseline. That apply creates IAM and errors on the instance. Retarget `TFC_AWS_APPLY_ROLE_ARN` and `TFC_AWS_PLAN_ROLE_ARN` to `hcptf-forgejo` and `hcptf-forgejo-plan`, drop `--allow-workspace`, then apply again. Do not use a project variable set.
|
|
|
|
Backup bucket names are `forgejo-backups-011934824531` and `forgejo-backups-replica-011934824531` (us-west-2, Object Lock governance 90 days). The sections below describe the live mgmt host until that cutover.
|
|
|
|
## Architecture
|
|
|
|
- **EC2**: t4g.small (arm64), Amazon Linux 2023, 20 GiB gp3 root (no state lives on it) — look up instance ID with:
|
|
```
|
|
aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text
|
|
```
|
|
The AMI is cached in the committed `cdk.context.json` (`cachedInContext: true`) so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (`cdk context --reset <ami key> && cdk synth`).
|
|
- **Data volume**: standalone 50 GiB gp3 encrypted EBS volume (`RemovalPolicy.RETAIN`) mounted at `/var/lib/forgejo` — sqlite database, repositories, and logs all live here and **survive instance replacement and stack deletion**. UserData waits for the volume attachment, mounts the existing filesystem (a `blkid` guard prevents formatting a disk that has one), and tags it `forgejo-backup=true` for DLM snapshots.
|
|
- **Restore-on-boot**: if the data volume has no database on boot (first boot or total volume loss), UserData automatically downloads the latest S3 dump and restores it before starting the service — volume loss self-heals to ≤24h-old state with no manual steps.
|
|
- **Network**: Private subnet (us-east-1a), behind `seahaven-com` ALB for SSL termination
|
|
- **DNS**: `forgejo.seahaven.com` — Route53 alias record pointing to the `seahaven-com` ALB (not a direct A record)
|
|
- **TLS**: Wildcard cert on ALB, HTTP internally on port 3000
|
|
- **Backup**: Nightly `forgejo dump` to S3 + EBS snapshots via DLM (see [3-2-1 Backup Strategy](#3-2-1-backup-strategy))
|
|
- **Admin access**: SSM Session Manager (no SSH port exposed)
|
|
- **CI/CD**: Pull requests call the org CI workflows (CDK synth, plus Terraform fmt/validate). CDK deploy on push is frozen. `workflow_dispatch` can still patch the live mgmt host. The replacement workspace is `forgejo-prod`.
|
|
|
|
### Ports
|
|
|
|
| Port | Protocol | Source | Purpose |
|
|
|------|----------|--------|---------|
|
|
| 443 | HTTPS | ALB (public) | Web UI + HTTP git clone |
|
|
| 3000 | HTTP | ALB → instance | Internal traffic from ALB |
|
|
| 2222 | SSH | VPC + VPN | Git SSH operations |
|
|
|
|
## 3-2-1 Backup Strategy
|
|
|
|
All backups follow a 3-2-1 strategy: 3 copies, 2 storage types, 1 offsite provider.
|
|
|
|
| Copy | Location | Type | Retention |
|
|
|------|----------|------|-----------|
|
|
| Live | Standalone EBS data volume, RETAIN (us-east-1) | Block | N/A |
|
|
| Near-site | S3 replica (us-west-2) | Object | Archive: indefinite, noncurrent versions: 90d |
|
|
| Offsite | GCS `forgejo-backups-offsite-seahaven` (GCP us-central1) | Object | 2-year locked retention |
|
|
|
|
**Daily data flow:**
|
|
|
|
| Time (UTC) | Event |
|
|
|------------|-------|
|
|
| 05:00 | `forgejo dump` → `s3://forgejo-backups-328440206208/archive/{date}/` |
|
|
| ~05:01 | S3 CRR replicates to `forgejo-backups-replica-328440206208` (us-west-2) |
|
|
| 06:00 | DLM EBS snapshot (30-day retention) |
|
|
| 08:00 | Verification Lambda checks all 3 locations, posts to Slack |
|
|
| 10:00 | GCS Storage Transfer pulls from S3 to GCS offsite |
|
|
|
|
**S3 source lifecycle:** Standard 30d → Glacier (no expiration).
|
|
|
|
**Immutability layers:**
|
|
- S3 Versioning on both source and replica buckets
|
|
- S3 Object Lock (Governance, 90d) on the replica bucket
|
|
- GCS Bucket Lock (2yr, irreversible) on the offsite bucket
|
|
|
|
### Verification
|
|
|
|
The `forgejo-backup-verification` Lambda runs daily at 08:00 UTC and checks:
|
|
1. S3 source has a recent dump under `archive/`
|
|
2. S3 replica has replicated the latest dump
|
|
3. GCS offsite has received the latest transfer
|
|
4. EBS snapshots exist within the last 48 hours
|
|
|
|
On the 1st of each month at 09:00 UTC, it runs a restore test: downloads the latest dump, extracts the archive, and runs SQLite integrity checks.
|
|
|
|
Two CloudWatch alarms watch the verification Lambda (both notify the `site-alerts` SNS topic, ALARM action only):
|
|
- `forgejo-backup-verification-errors` — Errors ≥ 1 in an hour, missing data = not breaching ("when it runs, did it fail")
|
|
- `forgejo-backup-verification-not-running` — Invocations < 1 over 24h, missing data = breaching ("did it run at all")
|
|
|
|
### Manual backup
|
|
|
|
```bash
|
|
sudo /usr/local/bin/forgejo-backup.sh
|
|
```
|
|
|
|
### Restore from S3
|
|
|
|
For backups older than 30 days (Glacier), restore the object first:
|
|
|
|
```bash
|
|
aws s3api restore-object --bucket forgejo-backups-328440206208 \
|
|
--key "archive/<date>/forgejo-<date>.tar.gz" \
|
|
--restore-request '{"Days":7,"GlacierJobParameters":{"Tier":"Standard"}}'
|
|
# Wait ~3-5 hours for restore to complete, then:
|
|
```
|
|
|
|
Download and restore:
|
|
|
|
```bash
|
|
aws s3 cp s3://forgejo-backups-328440206208/archive/<date>/forgejo-<date>.tar.gz /tmp/
|
|
systemctl stop forgejo
|
|
mkdir -p /tmp/forgejo-restore && tar -xzf /tmp/forgejo-<date>.tar.gz -C /tmp/forgejo-restore
|
|
cd /tmp/forgejo-restore
|
|
cp app.ini /etc/forgejo/app.ini
|
|
cp gitea-db.sqlite3 /var/lib/forgejo/data/forgejo.db
|
|
rm -rf /var/lib/forgejo/data/repositories
|
|
cp -a repos /var/lib/forgejo/data/repositories
|
|
cp -a data/. /var/lib/forgejo/data/
|
|
[ -d lfs ] && cp -a lfs/. /var/lib/forgejo/data/lfs/
|
|
[ -d custom ] && cp -a custom/. /var/lib/forgejo/custom/
|
|
chown -R forgejo:forgejo /var/lib/forgejo /etc/forgejo/app.ini
|
|
systemctl start forgejo
|
|
rm -rf /tmp/forgejo-restore /tmp/forgejo-<date>.tar.gz
|
|
```
|
|
|
|
### Restore from GCS (disaster recovery)
|
|
|
|
```bash
|
|
gcloud config set project sea-haven-backups
|
|
gsutil cp gs://forgejo-backups-offsite-seahaven/archive/<date>/forgejo-<date>.tar.gz /tmp/
|
|
# Then follow the same restore steps as S3 above
|
|
```
|
|
|
|
## Autodiscovery
|
|
|
|
An hourly cron job checks the `Sea-Haven-Industries` GitHub org for new repositories and mirrors them into Forgejo automatically.
|
|
|
|
- **Active repos** are created as mirrors (ongoing sync).
|
|
- **Archived repos** are created as static one-time imports.
|
|
- **Script**: `/usr/local/bin/forgejo-autodiscover.sh`
|
|
- **Log**: `/var/log/forgejo-autodiscover.log`
|
|
|
|
## Token Refresh
|
|
|
|
A daily cron at 4:30 UTC reads the GitHub PAT from Secrets Manager (`forgejo/github-pat`) and updates the git remote URL on every mirror repository so credentials stay current.
|
|
|
|
- **Script**: `/usr/local/bin/forgejo-refresh-tokens.sh`
|
|
|
|
## PAT Rotation
|
|
|
|
The GitHub personal access token used for mirroring is a fine-grained PAT scoped to `Sea-Haven-Industries` with **Contents: Read-only** permissions and a 1-year expiration. It is stored in Secrets Manager at `forgejo/github-pat`.
|
|
|
|
To rotate:
|
|
|
|
1. Create a new fine-grained PAT on GitHub with the same scope.
|
|
2. Update the secret value in Secrets Manager (`forgejo/github-pat`).
|
|
3. The daily token-refresh cron will pick it up automatically.
|
|
|
|
To force immediate propagation:
|
|
|
|
```bash
|
|
sudo /usr/local/bin/forgejo-refresh-tokens.sh
|
|
```
|
|
|
|
## Secrets Manager
|
|
|
|
| Secret | Purpose |
|
|
|--------|---------|
|
|
| `forgejo/admin-password` | Forgejo admin user password |
|
|
| `forgejo/api-token` | Forgejo API token (used by autodiscovery and token refresh scripts) |
|
|
| `forgejo/github-pat` | GitHub fine-grained PAT for mirroring |
|
|
| `forgejo/gcs-sa-key` | GCP service account key for offsite backup verification |
|
|
| `forgejo/gcs-transfer-credentials` | AWS IAM credentials for GCS Storage Transfer Service |
|
|
| `forgejo/slack-webhook` | Slack webhook URL for backup verification alerts |
|
|
|
|
## First-time setup
|
|
|
|
After the stack deploys, connect via SSM and create the admin user:
|
|
|
|
```bash
|
|
INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
|
|
aws ssm start-session --target "$INSTANCE_ID"
|
|
|
|
sudo -u forgejo /usr/local/bin/forgejo admin user create \
|
|
--admin \
|
|
--username adam \
|
|
--password '<password>' \
|
|
--email adam@seahavenind.com \
|
|
--config /etc/forgejo/app.ini
|
|
```
|
|
|
|
Admin password is stored in Secrets Manager at `forgejo/admin-password`.
|
|
|
|
Access the web UI at `https://forgejo.seahaven.com`.
|
|
|
|
## Migrating repos from GitHub
|
|
|
|
### Archived repos (one-time import)
|
|
|
|
In the Forgejo web UI: **New Migration → GitHub** → paste the GitHub repo URL. Use a GitHub personal access token for private repos. These are full imports (code, issues, PRs, releases).
|
|
|
|
### Active repos (mirror sync)
|
|
|
|
Same migration flow, but check **This Repository Will Be A Mirror**. Forgejo polls GitHub hourly (`DEFAULT_INTERVAL = 1h` in app.ini) and keeps the mirror in sync.
|
|
|
|
## GCP Offsite Setup (one-time)
|
|
|
|
Run the setup script to create the GCS offsite bucket, service account, and store credentials:
|
|
|
|
```bash
|
|
./scripts/gcp-setup.sh
|
|
```
|
|
|
|
This creates the `sea-haven-backups` GCP project with a locked-retention GCS bucket. After running, configure the Storage Transfer job in the GCP Console using the AWS credentials from `forgejo/gcs-transfer-credentials`.
|
|
|
|
## Infrastructure as Code (CDK)
|
|
|
|
All AWS infrastructure is defined as an [AWS CDK](https://docs.aws.amazon.com/cdk/) v2 app written in TypeScript. There is no console-managed infrastructure — every resource in the sections above is synthesized from this repo.
|
|
|
|
### Project layout
|
|
|
|
| Path | Purpose |
|
|
|------|---------|
|
|
| `bin/app.ts` | CDK app entry point. Instantiates both stacks and wires the dependency between them. |
|
|
| `lib/forgejo-stack.ts` | Main stack (`forgejo`, us-east-1): EC2 instance, security groups, IAM, ALB target/DNS, secrets, EBS data volume, DLM snapshots, and the backup-verification construct. Holds the `FORGEJO_VERSION` constant. |
|
|
| `lib/forgejo-replica-stack.ts` | Replica stack (`forgejo-replica`, us-west-2): the S3 CRR replica bucket with versioning, Object Lock (Governance, 90d), and Glacier lifecycle. |
|
|
| `lib/constructs/backup-verification.ts` | `BackupVerification` construct — the daily verification Lambda, its schedule, and CloudWatch alarms. |
|
|
| `lambda/backup-verification/` | Python 3.12 handler (`app.py` + `requirements.txt`) bundled via `@aws-cdk/aws-lambda-python-alpha`. |
|
|
| `scripts/gcp-setup.sh` | One-time GCP offsite bucket/service-account provisioning (see [GCP Offsite Setup](#gcp-offsite-setup-one-time)). |
|
|
| `cdk.json` | CDK configuration — the `app` command and context feature flags. |
|
|
| `cdk.context.json` | Cached context lookups (VPC + AMI), committed so synth is deterministic. |
|
|
|
|
### Stacks
|
|
|
|
The app defines two stacks, both pinned to account `328440206208` with explicit kebab-case `stackName`s:
|
|
|
|
- **`forgejo-replica`** (us-west-2) — the S3 replica bucket. Deployed first; exports the replica bucket ARN and name.
|
|
- **`forgejo`** (us-east-1) — the main stack. Consumes the replica bucket ARN/name and declares a dependency on `forgejo-replica`, so CDK always deploys the replica first.
|
|
|
|
### `cdk.json`
|
|
|
|
`cdk.json` is the CDK entry configuration, committed to the repo:
|
|
|
|
- **`app`**: `tsx bin/app.ts` — runs the TypeScript app directly via `tsx` (a pinned `devDependency`), so no separate `tsc` build step is needed for synth.
|
|
- **`watch`**: include/exclude globs for `cdk watch`.
|
|
- **`context`**: CDK feature flags (e.g. `@aws-cdk/core:target-partitions`). Cached lookup context (VPC subnets, the AL2023 arm64 AMI) lives separately in `cdk.context.json` — `cdk.json` holds only feature flags.
|
|
|
|
`aws-cdk-lib` is pinned to an exact version (no `^`/`~`) per the CDK version policy; Dependabot keeps it current.
|
|
|
|
### Common commands
|
|
|
|
```bash
|
|
npm install
|
|
npm run synth # cdk synth — emit CloudFormation without deploying
|
|
npm run diff # cdk diff — diff local app against deployed stacks
|
|
npm run deploy # cdk deploy — deploy (pass -- --all for both stacks)
|
|
```
|
|
|
|
## Deployment
|
|
|
|
Pull requests call the org reusables from `.github/workflows/ci.yaml`: CDK synth, and Terraform `fmt` / `init -backend=false` / `validate` on `terraform/`. A merge to `main` does not deploy. The CDK workflow (`.github/workflows/deploy.yaml`) runs only on `workflow_dispatch`, and that path still targets the live mgmt stacks.
|
|
|
|
New infrastructure is the `terraform/` root, applied from HCP workspace `forgejo-prod`. See [HCP Terraform (PLAT-80)](#hcp-terraform-plat-80). Do not `cdk deploy` this repo into seahaven-prod.
|
|
|
|
## Post-deploy: store Slack webhook
|
|
|
|
Store the Slack webhook URL for backup verification alerts:
|
|
|
|
```bash
|
|
aws secretsmanager create-secret --name forgejo/slack-webhook \
|
|
--secret-string "https://hooks.slack.com/services/YOUR/WEBHOOK/URL" \
|
|
--region us-east-1
|
|
```
|
|
|
|
## Updating Forgejo
|
|
|
|
Update the `FORGEJO_VERSION` constant in `lib/forgejo-stack.ts` and deploy. This replaces the instance, so ensure the latest EBS snapshot is available for data recovery if needed. Alternatively, update in-place via SSM:
|
|
|
|
```bash
|
|
INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
|
|
aws ssm start-session --target "$INSTANCE_ID"
|
|
|
|
sudo systemctl stop forgejo
|
|
sudo curl -Lo /usr/local/bin/forgejo "https://codeberg.org/forgejo/forgejo/releases/download/v<NEW_VERSION>/forgejo-<NEW_VERSION>-linux-arm64"
|
|
sudo chmod +x /usr/local/bin/forgejo
|
|
sudo systemctl start forgejo
|
|
```
|