forgejo/README.md

329 lines
18 KiB
Markdown
Raw Normal View History

# forgejo
![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
![Python](https://img.shields.io/badge/Python-3776AB?logo=python&logoColor=white)
![AWS CDK](https://img.shields.io/badge/AWS-CDK-FF9900?logo=amazonaws&logoColor=white)
![CI](https://github.com/Sea-Haven-Industries/forgejo/actions/workflows/ci.yaml/badge.svg)
Self-hosted Forgejo git server for archiving GitHub repos and mirroring active ones. The live host is still the mgmt CDK stack until cutover. The replacement is HCP Terraform in seahaven-prod.
## HCP Terraform (PLAT-80)
| | |
|---|---|
| HCP workspace | `forgejo-prod` (project `seahaven-prod`, working directory `terraform/`) |
| VCS triggers | `trigger-patterns = ["terraform/**/*", "lambda/**/*"]` |
| Account | seahaven-prod `011934824531` |
| Plan/apply roles | `hcptf-forgejo-plan` / `hcptf-forgejo` |
| Apply | Manual. Auto-apply stays off until after DNS cutover and one nightly dump in the new bucket. |
A `lambda/`-only merge must still queue a run, so the trigger patterns include `lambda/**/*` as well as `terraform/**/*`. Docs-only commits do not apply.
The instance attaches to the After Hours VPC (`10.70.0.0/16`) through workspace variables `existing_vpc_id` and `existing_public_subnet_ids` (afterhours-shift-manager outputs `vpc_id` and `public_subnet_ids`). The first subnet is the instance availability zone. AMI id is pinned in `terraform/variables.tf` (`ami_id`). Do not switch it to `most_recent`.
DNS stays in the mgmt zone `Z06652411XKH89KTZD3XA`. Terraform does not own `forgejo.seahaven.com`. Cutover is an alias flip to the new ALB. Until `enable_https` is true, the ALB listens on port 80 so a restore can be proved against `alb_dns_name` without moving the public name. Create the `acm_validation_records` output in the mgmt zone before setting `enable_https`.
`enable_schedules` stays false until cutover so the not-running alarm does not page `site-alerts`.
Secret values are not in Terraform. IAM uses name-prefix ARNs because `hcptf-bootstrap-plan` cannot `DescribeSecret`. These names must exist in seahaven-prod before the instance boots: `forgejo/admin-password`, `forgejo/api-token`, `forgejo/github-pat`, `forgejo/gcs-sa-key`, `forgejo/slack-webhook`, `forgejo/gcs-transfer-credentials`. The GCS transfer user is `forgejo-gcs-transfer`. Create its access key by hand and store it in `forgejo/gcs-transfer-credentials`. Do not put the key in Terraform.
First apply uses `hcptf-bootstrap` / `hcptf-bootstrap-plan` after `scripts/create-hcptf-bootstrap-roles.sh --account prod --allow-workspace forgejo-prod` in seahaven-org-baseline. That apply creates IAM and errors on the instance. Retarget `TFC_AWS_APPLY_ROLE_ARN` and `TFC_AWS_PLAN_ROLE_ARN` to `hcptf-forgejo` and `hcptf-forgejo-plan`, drop `--allow-workspace`, then apply again. Do not use a project variable set.
Backup bucket names are `forgejo-backups-011934824531` and `forgejo-backups-replica-011934824531` (us-west-2, Object Lock governance 90 days). The sections below describe the live mgmt host until that cutover.
## Architecture
- **EC2**: t4g.small (arm64), Amazon Linux 2023, 20 GiB gp3 root (no state lives on it) — look up instance ID with:
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
```
aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text
```
The AMI is cached in the committed `cdk.context.json` (`cachedInContext: true`) so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (`cdk context --reset <ami key> && cdk synth`).
- **Data volume**: standalone 50 GiB gp3 encrypted EBS volume (`RemovalPolicy.RETAIN`) mounted at `/var/lib/forgejo` — sqlite database, repositories, and logs all live here and **survive instance replacement and stack deletion**. UserData waits for the volume attachment, mounts the existing filesystem (a `blkid` guard prevents formatting a disk that has one), and tags it `forgejo-backup=true` for DLM snapshots.
- **Restore-on-boot**: if the data volume has no database on boot (first boot or total volume loss), UserData automatically downloads the latest S3 dump and restores it before starting the service — volume loss self-heals to ≤24h-old state with no manual steps.
- **Network**: Private subnet (us-east-1a), behind `seahaven-com` ALB for SSL termination
- **DNS**: `forgejo.seahaven.com` — Route53 alias record pointing to the `seahaven-com` ALB (not a direct A record)
- **TLS**: Wildcard cert on ALB, HTTP internally on port 3000
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
- **Backup**: Nightly `forgejo dump` to S3 + EBS snapshots via DLM (see [3-2-1 Backup Strategy](#3-2-1-backup-strategy))
- **Admin access**: SSM Session Manager (no SSH port exposed)
- **CI/CD**: Pull requests call the org CI workflows (CDK synth, plus Terraform fmt/validate). CDK deploy on push is frozen. `workflow_dispatch` can still patch the live mgmt host. The replacement workspace is `forgejo-prod`.
### Ports
| Port | Protocol | Source | Purpose |
|------|----------|--------|---------|
| 443 | HTTPS | ALB (public) | Web UI + HTTP git clone |
| 3000 | HTTP | ALB → instance | Internal traffic from ALB |
| 2222 | SSH | VPC + VPN | Git SSH operations |
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
## 3-2-1 Backup Strategy
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
All backups follow a 3-2-1 strategy: 3 copies, 2 storage types, 1 offsite provider.
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
| Copy | Location | Type | Retention |
|------|----------|------|-----------|
| Live | Standalone EBS data volume, RETAIN (us-east-1) | Block | N/A |
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
| Near-site | S3 replica (us-west-2) | Object | Archive: indefinite, noncurrent versions: 90d |
| Offsite | GCS `forgejo-backups-offsite-seahaven` (GCP us-central1) | Object | 2-year locked retention |
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
**Daily data flow:**
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
| Time (UTC) | Event |
|------------|-------|
| 05:00 | `forgejo dump` → `s3://forgejo-backups-328440206208/archive/{date}/` |
| ~05:01 | S3 CRR replicates to `forgejo-backups-replica-328440206208` (us-west-2) |
| 06:00 | DLM EBS snapshot (30-day retention) |
| 08:00 | Verification Lambda checks all 3 locations, posts to Slack |
| 10:00 | GCS Storage Transfer pulls from S3 to GCS offsite |
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
**S3 source lifecycle:** Standard 30d → Glacier (no expiration).
**Immutability layers:**
- S3 Versioning on both source and replica buckets
- S3 Object Lock (Governance, 90d) on the replica bucket
- GCS Bucket Lock (2yr, irreversible) on the offsite bucket
### Verification
The `forgejo-backup-verification` Lambda runs daily at 08:00 UTC and checks:
1. S3 source has a recent dump under `archive/`
2. S3 replica has replicated the latest dump
3. GCS offsite has received the latest transfer
4. EBS snapshots exist within the last 48 hours
On the 1st of each month at 09:00 UTC, it runs a restore test: downloads the latest dump, extracts the archive, and runs SQLite integrity checks.
Two CloudWatch alarms watch the verification Lambda (both notify the `site-alerts` SNS topic, ALARM action only):
- `forgejo-backup-verification-errors` — Errors ≥ 1 in an hour, missing data = not breaching ("when it runs, did it fail")
- `forgejo-backup-verification-not-running` — Invocations < 1 over 24h, missing data = breaching ("did it run at all")
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
### Manual backup
```bash
sudo /usr/local/bin/forgejo-backup.sh
```
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
### Restore from S3
For backups older than 30 days (Glacier), restore the object first:
```bash
aws s3api restore-object --bucket forgejo-backups-011934824531 \
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
--key "archive/<date>/forgejo-<date>.tar.gz" \
--restore-request '{"Days":7,"GlacierJobParameters":{"Tier":"Standard"}}'
# Wait ~3-5 hours for restore to complete, then:
```
Download and restore. The copy does not start until the dump contains a database, so a failed download leaves the live data alone.
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
```bash
set -euo pipefail
rm -rf /var/lib/forgejo/.restore && mkdir -p /var/lib/forgejo/.restore
aws s3 cp s3://forgejo-backups-011934824531/archive/<date>/forgejo-<date>.tar.gz - --no-progress \
| tar -xz -C /var/lib/forgejo/.restore
if [ ! -s /var/lib/forgejo/.restore/gitea-db.sqlite3 ] && [ ! -s /var/lib/forgejo/.restore/data/forgejo.db ]; then
echo "Dump has no database" >&2
exit 1
fi
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
systemctl stop forgejo
[ -d /var/lib/forgejo/.restore/data ] && cp -a /var/lib/forgejo/.restore/data/. /var/lib/forgejo/data/
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
rm -rf /var/lib/forgejo/data/repositories
mkdir -p /var/lib/forgejo/data/repositories
[ -d /var/lib/forgejo/.restore/repos ] && cp -a /var/lib/forgejo/.restore/repos/. /var/lib/forgejo/data/repositories/
[ -f /var/lib/forgejo/.restore/gitea-db.sqlite3 ] && cp /var/lib/forgejo/.restore/gitea-db.sqlite3 /var/lib/forgejo/data/forgejo.db
[ -d /var/lib/forgejo/.restore/lfs ] && mkdir -p /var/lib/forgejo/data/lfs && cp -a /var/lib/forgejo/.restore/lfs/. /var/lib/forgejo/data/lfs/
[ -d /var/lib/forgejo/.restore/custom ] && cp -a /var/lib/forgejo/.restore/custom/. /var/lib/forgejo/custom/
if [ ! -s /var/lib/forgejo/data/forgejo.db ]; then
echo "Restore did not produce /var/lib/forgejo/data/forgejo.db" >&2
exit 1
fi
chown -R forgejo:forgejo /var/lib/forgejo
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
systemctl start forgejo
rm -rf /var/lib/forgejo/.restore
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
```
### Restore from GCS (disaster recovery)
This path does not read S3. It unpacks the offsite object on the data volume and uses the same copy order as the S3 restore.
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
```bash
set -euo pipefail
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
gcloud config set project sea-haven-backups
rm -rf /var/lib/forgejo/.restore && mkdir -p /var/lib/forgejo/.restore
gsutil cp gs://forgejo-backups-offsite-seahaven/archive/<date>/forgejo-<date>.tar.gz - \
| tar -xz -C /var/lib/forgejo/.restore
if [ ! -s /var/lib/forgejo/.restore/gitea-db.sqlite3 ] && [ ! -s /var/lib/forgejo/.restore/data/forgejo.db ]; then
echo "Dump has no database" >&2
exit 1
fi
systemctl stop forgejo
[ -d /var/lib/forgejo/.restore/data ] && cp -a /var/lib/forgejo/.restore/data/. /var/lib/forgejo/data/
rm -rf /var/lib/forgejo/data/repositories
mkdir -p /var/lib/forgejo/data/repositories
[ -d /var/lib/forgejo/.restore/repos ] && cp -a /var/lib/forgejo/.restore/repos/. /var/lib/forgejo/data/repositories/
[ -f /var/lib/forgejo/.restore/gitea-db.sqlite3 ] && cp /var/lib/forgejo/.restore/gitea-db.sqlite3 /var/lib/forgejo/data/forgejo.db
[ -d /var/lib/forgejo/.restore/lfs ] && mkdir -p /var/lib/forgejo/data/lfs && cp -a /var/lib/forgejo/.restore/lfs/. /var/lib/forgejo/data/lfs/
[ -d /var/lib/forgejo/.restore/custom ] && cp -a /var/lib/forgejo/.restore/custom/. /var/lib/forgejo/custom/
if [ ! -s /var/lib/forgejo/data/forgejo.db ]; then
echo "Restore did not produce /var/lib/forgejo/data/forgejo.db" >&2
exit 1
fi
chown -R forgejo:forgejo /var/lib/forgejo
systemctl start forgejo
rm -rf /var/lib/forgejo/.restore
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
```
## Autodiscovery
An hourly cron job checks the `Sea-Haven-Industries` GitHub org for new repositories and mirrors them into Forgejo automatically.
- **Active repos** are created as mirrors (ongoing sync).
- **Archived repos** are created as static one-time imports.
- **Script**: `/usr/local/bin/forgejo-autodiscover.sh`
- **Log**: `/var/log/forgejo-autodiscover.log`
## Token Refresh
A daily cron at 4:30 UTC reads the GitHub PAT from Secrets Manager (`forgejo/github-pat`) and updates the git remote URL on every mirror repository so credentials stay current.
- **Script**: `/usr/local/bin/forgejo-refresh-tokens.sh`
## PAT Rotation
The GitHub personal access token used for mirroring is a fine-grained PAT scoped to `Sea-Haven-Industries` with **Contents: Read-only** permissions and a 1-year expiration. It is stored in Secrets Manager at `forgejo/github-pat`.
To rotate:
1. Create a new fine-grained PAT on GitHub with the same scope.
2. Update the secret value in Secrets Manager (`forgejo/github-pat`).
3. The daily token-refresh cron will pick it up automatically.
To force immediate propagation:
```bash
sudo /usr/local/bin/forgejo-refresh-tokens.sh
```
## Secrets Manager
| Secret | Purpose |
|--------|---------|
| `forgejo/admin-password` | Forgejo admin user password |
| `forgejo/api-token` | Forgejo API token (used by autodiscovery and token refresh scripts) |
| `forgejo/github-pat` | GitHub fine-grained PAT for mirroring |
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
| `forgejo/gcs-sa-key` | GCP service account key for offsite backup verification |
| `forgejo/gcs-transfer-credentials` | AWS IAM credentials for GCS Storage Transfer Service |
| `forgejo/slack-webhook` | Slack webhook URL for backup verification alerts |
## First-time setup
After the stack deploys, connect via SSM and create the admin user:
```bash
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
aws ssm start-session --target "$INSTANCE_ID"
sudo -u forgejo /usr/local/bin/forgejo admin user create \
--admin \
--username adam \
--password '<password>' \
--email adam@seahavenind.com \
--config /etc/forgejo/app.ini
```
Admin password is stored in Secrets Manager at `forgejo/admin-password`.
Access the web UI at `https://forgejo.seahaven.com`.
## Migrating repos from GitHub
### Archived repos (one-time import)
In the Forgejo web UI: **New Migration → GitHub** → paste the GitHub repo URL. Use a GitHub personal access token for private repos. These are full imports (code, issues, PRs, releases).
### Active repos (mirror sync)
Same migration flow, but check **This Repository Will Be A Mirror**. Forgejo polls GitHub hourly (`DEFAULT_INTERVAL = 1h` in app.ini) and keeps the mirror in sync.
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
## GCP Offsite Setup (one-time)
Run the setup script to create the GCS offsite bucket, service account, and store credentials:
```bash
./scripts/gcp-setup.sh
```
This creates the `sea-haven-backups` GCP project with a locked-retention GCS bucket. After running, configure the Storage Transfer job in the GCP Console using the AWS credentials from `forgejo/gcs-transfer-credentials`.
## Infrastructure as Code (CDK)
All AWS infrastructure is defined as an [AWS CDK](https://docs.aws.amazon.com/cdk/) v2 app written in TypeScript. There is no console-managed infrastructure — every resource in the sections above is synthesized from this repo.
### Project layout
| Path | Purpose |
|------|---------|
| `bin/app.ts` | CDK app entry point. Instantiates both stacks and wires the dependency between them. |
| `lib/forgejo-stack.ts` | Main stack (`forgejo`, us-east-1): EC2 instance, security groups, IAM, ALB target/DNS, secrets, EBS data volume, DLM snapshots, and the backup-verification construct. Holds the `FORGEJO_VERSION` constant. |
| `lib/forgejo-replica-stack.ts` | Replica stack (`forgejo-replica`, us-west-2): the S3 CRR replica bucket with versioning, Object Lock (Governance, 90d), and Glacier lifecycle. |
| `lib/constructs/backup-verification.ts` | `BackupVerification` construct — the daily verification Lambda, its schedule, and CloudWatch alarms. |
| `lambda/backup-verification/` | Python 3.12 handler (`app.py` + `requirements.txt`) bundled via `@aws-cdk/aws-lambda-python-alpha`. |
| `scripts/gcp-setup.sh` | One-time GCP offsite bucket/service-account provisioning (see [GCP Offsite Setup](#gcp-offsite-setup-one-time)). |
| `cdk.json` | CDK configuration — the `app` command and context feature flags. |
| `cdk.context.json` | Cached context lookups (VPC + AMI), committed so synth is deterministic. |
### Stacks
The app defines two stacks, both pinned to account `328440206208` with explicit kebab-case `stackName`s:
- **`forgejo-replica`** (us-west-2) — the S3 replica bucket. Deployed first; exports the replica bucket ARN and name.
- **`forgejo`** (us-east-1) — the main stack. Consumes the replica bucket ARN/name and declares a dependency on `forgejo-replica`, so CDK always deploys the replica first.
### `cdk.json`
`cdk.json` is the CDK entry configuration, committed to the repo:
- **`app`**: `tsx bin/app.ts` — runs the TypeScript app directly via `tsx` (a pinned `devDependency`), so no separate `tsc` build step is needed for synth.
- **`watch`**: include/exclude globs for `cdk watch`.
- **`context`**: CDK feature flags (e.g. `@aws-cdk/core:target-partitions`). Cached lookup context (VPC subnets, the AL2023 arm64 AMI) lives separately in `cdk.context.json` — `cdk.json` holds only feature flags.
`aws-cdk-lib` is pinned to an exact version (no `^`/`~`) per the CDK version policy; Dependabot keeps it current.
### Common commands
```bash
npm install
npm run synth # cdk synth — emit CloudFormation without deploying
npm run diff # cdk diff — diff local app against deployed stacks
npm run deploy # cdk deploy — deploy (pass -- --all for both stacks)
```
## Deployment
Pull requests call the org reusables from `.github/workflows/ci.yaml`: CDK synth, and Terraform `fmt` / `init -backend=false` / `validate` on `terraform/`. A merge to `main` does not deploy. The CDK workflow (`.github/workflows/deploy.yaml`) runs only on `workflow_dispatch`, and that path still targets the live mgmt stacks.
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
New infrastructure is the `terraform/` root, applied from HCP workspace `forgejo-prod`. See [HCP Terraform (PLAT-80)](#hcp-terraform-plat-80). Do not `cdk deploy` this repo into seahaven-prod.
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
## Post-deploy: store Slack webhook
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
Store the Slack webhook URL for backup verification alerts:
Add 3-2-1 backup strategy (#2) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
```bash
aws secretsmanager create-secret --name forgejo/slack-webhook \
--secret-string "https://hooks.slack.com/services/YOUR/WEBHOOK/URL" \
--region us-east-1
```
## Updating Forgejo
Update the `FORGEJO_VERSION` constant in `lib/forgejo-stack.ts` and deploy. This replaces the instance, so ensure the latest EBS snapshot is available for data recovery if needed. Alternatively, update in-place via SSM:
```bash
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5) * Add 3-2-1 backup strategy with cross-region replication and GCS offsite Implements a fully compliant 3-2-1 backup architecture: - Copy 1 (live): Harden existing EBS snapshots to 30-day retention - Copy 2 (near-site): S3 cross-region replication to us-west-2 with Object Lock (governance 90d) and versioning - Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project with 2-year irreversible retention lock Also adds a verification Lambda that checks all 3 locations daily and runs monthly restore tests with SQLite integrity checks. * Enable QEMU in CI for arm64 Lambda Docker builds * Commit cdk.context.json for CI synth without AWS credentials Vpc.fromLookup requires cached context to synthesize without AWS credentials. Required for CI which runs cdk synth without an OIDC role. * Fix GCP project ID to sea-haven-backups * Address code review findings for backup verification Fix 4 critical issues: - Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without) - Add stack dependency so replica deploys before main stack - Fix DB file extension matching (.sqlite3/.sql instead of .db) - Replace nonexistent `forgejo restore` command with actual restore steps in README Fix 4 moderate issues: - Add timeout=10 to Slack webhook urlopen call - Add filter='data' to tarfile.extract for PEP 706 compliance - Add explicit ValueError for unknown handler mode - Use date-scoped S3/GCS prefix instead of unbounded listing * Fix backup strategy bug findings * Handle SQL text dumps separately from binary SQLite in restore test Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export), not a binary SQLite file. Opening it directly with sqlite3.connect() throws DatabaseError. Now imports the SQL dump into a temp DB first. * Fix GCS backup check: align staleness cutoff and add size validation GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h), making the staleness check unreachable. Also added 1MB minimum file size validation to match the S3 check. * Rename SECRET_ARN env vars to SECRET_NAME to match actual values * Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule * Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt * Fix restore runbook: trailing-dot cp idiom and Glacier restore step * Rename GCS service account to match read-only permissions * Add 4 GiB ephemeral storage to verification Lambda Monthly restore-test downloads and extracts the full dump tarball in /tmp. As the dump grows with LFS data, the default 512 MB will eventually cause ENOSPC failures. * Replace hardcoded instance ID in README with CloudFormation lookup The instance ID changes on every instance replacement (version upgrades, stack updates). Using a dynamic query prevents stale references and removes a manual update step from the deploy process. * Read backup S3 prefix from SSM parameter at runtime Adds /forgejo/backup-s3-prefix SSM parameter (value: archive) and updates the backup script to fetch it instead of hardcoding the prefix. Eliminates the manual post-deploy sed step. * Address cross-review findings for backup verification - Add size guard before downloading dump in restore test (3.5 GB cap) - Use paginator for list_objects_v2 in S3 checks and restore test - Remove unnecessary overrideLogicalId on GcsTransferCredentials secret - Pass explicit { mode: "daily" } to daily EventBridge rule target - Add fallback for SSM parameter fetch in backup script - Export replica bucket ARN/name from replica stack, consume via props * Add CloudWatch alarm for backup verification Lambda errors Fires on any Lambda error and on missing data (missed schedule). Catches silent failures where the Slack notification never fires. * Add .env to .gitignore Required by org CI conventions check. --------- Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
INSTANCE_ID=$(aws cloudformation describe-stacks --stack-name forgejo --query 'Stacks[0].Outputs[?OutputKey==`InstanceId`].OutputValue' --output text)
aws ssm start-session --target "$INSTANCE_ID"
sudo systemctl stop forgejo
sudo curl -Lo /usr/local/bin/forgejo "https://codeberg.org/forgejo/forgejo/releases/download/v<NEW_VERSION>/forgejo-<NEW_VERSION>-linux-arm64"
sudo chmod +x /usr/local/bin/forgejo
sudo systemctl start forgejo
```