Commit graph

9 commits

Author SHA1 Message Date
dependabot[bot]
6cf50af1af
chore(deps-dev): bump typescript from 6.0.3 to 7.0.2 (#43)
Some checks failed
Deploy / deploy (push) Has been cancelled
* chore(deps-dev): bump typescript from 6.0.3 to 7.0.2

Bumps [typescript](https://github.com/microsoft/TypeScript) from 6.0.3 to 7.0.2.
- [Release notes](https://github.com/microsoft/TypeScript/releases)
- [Commits](https://github.com/microsoft/TypeScript/commits)

---
updated-dependencies:
- dependency-name: typescript
  dependency-version: 7.0.2
  dependency-type: direct:development
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>

* fix: regenerate package-lock.json to reflect typescript bump

Update cdk.json to use tsx and regenerate package-lock.json to satisfy CI

* fix: declare and pin tsx runner for CDK app

TypeScript 7's native port is incompatible with ts-node, so cdk.json was
switched to tsx — but it was invoked via `npx`, which fetches an
unpinned copy from the registry at synth/deploy time with no lockfile
entry or integrity check.

Declare tsx as a pinned devDependency (~4.23.0) and invoke the local bin
instead of npx. Update the README's cdk.json section to match. Verified
with `npm run build` (tsc 7.0.2) and `cdk synth` — both green.

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Adam Moussa <adam@seahavenind.com>
2026-07-10 17:41:56 -04:00
Adam Moussa
72f0637fee
Document CDK app and cdk.json in README (#46)
Some checks are pending
Deploy / deploy (push) Waiting to run
The README covered runtime architecture and operations thoroughly but
never described the infrastructure-as-code layer itself. Readers had no
map of the CDK app: which files define the stacks, what cdk.json is, or
how to synth/diff. Add an Infrastructure as Code section covering the
project layout, the two stacks and their dependency, cdk.json, and the
common CDK commands.
2026-07-10 16:07:40 -04:00
Adam Moussa
e32401e554
Repo hygiene: PR labeler + README badges + dependabot (INFRA-56/57/66) (#29)
Some checks are pending
Deploy / deploy (push) Waiting to run
2026-06-11 14:13:26 -04:00
Adam Moussa
891dc0831b
docs: document persistent data volume, restore-on-boot, cached AMI, alarms (#25)
Some checks failed
Deploy / deploy (push) Has been cancelled
Storage architecture changed 2026-06-05: state moved off the root
volume onto a standalone RETAIN data volume with automatic S3 restore
on empty boot.
2026-06-05 15:54:26 -04:00
Adam Moussa
5ed1db788e
Add 3-2-1 backup strategy with cross-region and GCS offsite (#5)
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite

Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
  Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
  with 2-year irreversible retention lock

Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.

* Enable QEMU in CI for arm64 Lambda Docker builds

* Commit cdk.context.json for CI synth without AWS credentials

Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.

* Fix GCP project ID to sea-haven-backups

* Address code review findings for backup verification

Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README

Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing

* Fix backup strategy bug findings

* Handle SQL text dumps separately from binary SQLite in restore test

Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.

* Fix GCS backup check: align staleness cutoff and add size validation

GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.

* Rename SECRET_ARN env vars to SECRET_NAME to match actual values

* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule

* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt

* Fix restore runbook: trailing-dot cp idiom and Glacier restore step

* Rename GCS service account to match read-only permissions

* Add 4 GiB ephemeral storage to verification Lambda

Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.

* Replace hardcoded instance ID in README with CloudFormation lookup

The instance ID changes on every instance replacement (version
upgrades, stack updates). Using a dynamic query prevents stale
references and removes a manual update step from the deploy process.

* Read backup S3 prefix from SSM parameter at runtime

Adds /forgejo/backup-s3-prefix SSM parameter (value: archive)
and updates the backup script to fetch it instead of hardcoding
the prefix. Eliminates the manual post-deploy sed step.

* Address cross-review findings for backup verification

- Add size guard before downloading dump in restore test (3.5 GB cap)
- Use paginator for list_objects_v2 in S3 checks and restore test
- Remove unnecessary overrideLogicalId on GcsTransferCredentials secret
- Pass explicit { mode: "daily" } to daily EventBridge rule target
- Add fallback for SSM parameter fetch in backup script
- Export replica bucket ARN/name from replica stack, consume via props

* Add CloudWatch alarm for backup verification Lambda errors

Fires on any Lambda error and on missing data (missed schedule).
Catches silent failures where the Slack notification never fires.

* Add .env to .gitignore

Required by org CI conventions check.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-15 17:59:31 -04:00
Adam Moussa
cfda99927b
Add 3-2-1 backup strategy (#2)
* Add 3-2-1 backup strategy with cross-region replication and GCS offsite

Implements a fully compliant 3-2-1 backup architecture:
- Copy 1 (live): Harden existing EBS snapshots to 30-day retention
- Copy 2 (near-site): S3 cross-region replication to us-west-2 with
  Object Lock (governance 90d) and versioning
- Copy 3 (offsite): GCS bucket in dedicated seahaven-backups GCP project
  with 2-year irreversible retention lock

Also adds a verification Lambda that checks all 3 locations daily and
runs monthly restore tests with SQLite integrity checks.

* Enable QEMU in CI for arm64 Lambda Docker builds

* Commit cdk.context.json for CI synth without AWS credentials

Vpc.fromLookup requires cached context to synthesize without
AWS credentials. Required for CI which runs cdk synth without
an OIDC role.

* Fix GCP project ID to sea-haven-backups

* Address code review findings for backup verification

Fix 4 critical issues:
- Add filter/priority/deleteMarkerReplication to S3 CRR rule (deploy would fail without)
- Add stack dependency so replica deploys before main stack
- Fix DB file extension matching (.sqlite3/.sql instead of .db)
- Replace nonexistent `forgejo restore` command with actual restore steps in README

Fix 4 moderate issues:
- Add timeout=10 to Slack webhook urlopen call
- Add filter='data' to tarfile.extract for PEP 706 compliance
- Add explicit ValueError for unknown handler mode
- Use date-scoped S3/GCS prefix instead of unbounded listing

* Fix backup strategy bug findings

* Handle SQL text dumps separately from binary SQLite in restore test

Forgejo dump produces gitea-db.sql as a text SQL dump (XORM export),
not a binary SQLite file. Opening it directly with sqlite3.connect()
throws DatabaseError. Now imports the SQL dump into a temp DB first.

* Fix GCS backup check: align staleness cutoff and add size validation

GCS check used a 72h cutoff but only listed 2 days of prefixes (~48h),
making the staleness check unreachable. Also added 1MB minimum file
size validation to match the S3 check.

* Rename SECRET_ARN env vars to SECRET_NAME to match actual values

* Fix EBS snapshot state check, drop unused GCS write grant and dead lifecycle rule

* Fix restore runbook, DLM snapshot tagging, README cleanup, and gsutil prompt

* Fix restore runbook: trailing-dot cp idiom and Glacier restore step

* Rename GCS service account to match read-only permissions

* Add 4 GiB ephemeral storage to verification Lambda

Monthly restore-test downloads and extracts the full dump tarball
in /tmp. As the dump grows with LFS data, the default 512 MB will
eventually cause ENOSPC failures.

---------

Co-authored-by: Cursor Agent <cursoragent@cursor.com>
2026-05-14 18:08:06 -04:00
Adam Moussa
6ccfc1c506
Update README with backup, autodiscovery, and token management docs (#1)
Some checks failed
Deploy / deploy (push) Has been cancelled
Add documentation for S3 backups with Glacier lifecycle, hourly
autodiscovery of new GitHub org repos, daily PAT token refresh,
PAT rotation procedure, and Secrets Manager secret inventory.
2026-05-11 19:03:42 -04:00
Adam Moussa
0a4119880e Update README for ALB-backed HTTPS setup 2026-05-11 18:13:54 -04:00
Adam Moussa
a500d69716 Add Forgejo CDK stack
EC2 (t4g.small, arm64) in private subnet with VPN-only access,
DLM nightly snapshots, and Route53 DNS at forgejo.seahaven.com.
2026-05-11 17:52:37 -04:00