docs(infra): record the prod HCP path after management decommission (PLAT-77) (#58)

The management stack is deleted, so the docs now describe the live share and the CDK deploy workflow is removed.
This commit is contained in:
Adam Moussa 2026-09-30 00:03:47 +00:00 • committed by GitHub
parent 9f27827dcb
commit 3a07b2e968
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
3 changed files with 50 additions and 140 deletions

View file

@ -1,22 +0,0 @@
name: Deploy
# PLAT-77: push-to-main CDK deploy is frozen so HCP Terraform is the only
# path that can change this stack. workflow_dispatch stays for an explicit
# rollback of the management-account stack.
on:
workflow_dispatch:
permissions:
id-token: write
contents: read
concurrency:
group: deploy
cancel-in-progress: false
jobs:
deploy:
uses: Sea-Haven-Industries/.github/.github/workflows/cd-cdk.yaml@0a1010e63248c9ca9f042c870eb2c579ba6b9455 # v1.0.19
with:
node-version: "24"
secrets:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}

View file

@ -2,55 +2,38 @@
Instructions for coding agents working in this repository.
## Live path
The share runs in seahaven-prod under HCP Terraform workspace `file-share-prod`. Source is `terraform/`. Manual apply until the move is sealed. Do not enable auto-apply as part of a docs or cleanup change.
The management-account CDK stack was deleted on 2026-09-29. Do not run `cdk deploy`. `lib/` and `bin/` are the retired stack. The CDK deploy workflow has been removed.
The data volume is attached by the workspace variable `data_volume_id`. Terraform must not create or delete it. The previous management volume is retained until 2026-10-06 and is not the live disk.
## Infrastructure as Code principles
- **CDK + EC2**: this repo deploys an EC2 instance via CDK TypeScript. No Lambda, no SAM.
- **Lambda defaults (does not apply)** — this is a pure EC2 stack.
- **Exact-pin all CDK library versions** (`aws-cdk-lib`, `aws-cdk`, `constructs`). Never use `*` or `^` ranges.
- **Never commit** account IDs, role ARNs, VPC IDs, subnet IDs, or volume IDs in new code. The existing `cdk.context.json` already contains resolved values — do not add new deployment-specific identifiers without documented defaults.
- **Deploy with least-privilege IAM.** The instance role already has only `AmazonSSMManagedInstanceCore` + Secrets Manager read on `file-share/*`.
- **Verify**: `npm run build && npx cdk synth && npx cdk diff` before pushing. Do not commit `cdk.out/`.
- **Exact-pin CDK library versions** if you touch the retired CDK package. Never use `*` or `^` ranges.
- **Never commit** account IDs, role ARNs, VPC IDs, subnet IDs, or new volume IDs. The live volume id is an HCP variable.
- **Deploy with least-privilege IAM.** The instance role has SSM core plus Secrets Manager read on `file-share/*`.
- Do not commit `cdk.out/`.
## CDK directory layout
## Instance replacement
```
.
├── bin/
│ └── app.ts # CDK app entrypoint
├── lib/
│ └── file-share-stack.ts # All resource definitions
├── cdk.json # CDK context + config
├── cdk.context.json # Resolved context (committed)
├── package.json # Exact-pinned CDK dependencies
├── tsconfig.json
└── .github/
└── workflows/ # CI/CD
```
`user_data_replace_on_change` is false. Before any apply that would replace the instance (AMI, user data, instance type):
## `overrideLogicalId` — never remove
Removing `overrideLogicalId` on a deployed resource forces replacement. Never remove an existing call without documenting the replacement impact.
## Persistent EBS volumes
This repo's data volume (`vol-04d951cccacc435b5`) is **imported by ID**, not managed by CloudFormation. It survives instance replacement and stack deletion.
### Snapshot before an EC2-replacing deploy
Before any deploy that would replace the EC2 instance (AMI change, user-data change, instance type change):
1. Run `cdk diff` to confirm the instance will be replaced.
2. **Stop and ask for confirmation** — an instance replacement detaches the imported volume. A snapshot is the safety net.
3. If a DLM snapshot already exists from the same day, reference it rather than creating a duplicate.
1. Read the plan and confirm the instance is being replaced.
2. Stop and ask for confirmation. Replacement detaches the data volume.
3. Snapshot the volume first. If a DLM snapshot from the same day exists, use that instead of a second copy.
## Security group rules
- Never open 0.0.0.0/0. All rules use specific CIDR ranges (`10.10.0.0/16` VPN, `10.20.0.0/16` VPC).
- SSM Session Manager for instance access — no SSH key, no public port 22.
- No `0.0.0.0/0` ingress. Ingress is TCP 445, 8080, and 22 from `10.10.0.0/16` and `10.30.0.0/16`.
- Egress `0.0.0.0/0` stays so the instance can install packages and reach the pinned FileBrowser download.
- SSM Session Manager for instance access. No SSH key and no public port 22. SFTP for user `adam` is the office path, not an admin login.
## Secrets
Stored in AWS Secrets Manager (`file-share/smb-password`, `file-share/filebrowser-password`). The instance role reads them at boot. Do not embed secrets in UserData or source files.
Stored in AWS Secrets Manager (`file-share/smb-password`, `file-share/filebrowser-password`). The instance role reads them at boot. Do not embed secrets in user data or source files. Do not delete the management-account copies before 2026-10-06.
## Documentation

107
README.md
View file

@ -1,105 +1,54 @@
# file-share
# File Share
![CI](https://github.com/Sea-Haven-Industries/file-share/actions/workflows/ci.yaml/badge.svg)
![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
![AWS CDK](https://img.shields.io/badge/AWS-CDK-FF9900?logo=amazonaws&logoColor=white)
Samba, FileBrowser, and SFTP for the Sea Haven offices. The live host is an EC2 instance in seahaven-prod, managed by HCP Terraform workspace `file-share-prod`.
Personal file share server on AWS — Samba for macOS Finder integration and FileBrowser for web-based file management. Accessible exclusively over the site-to-site VPN.
Clients:
## Infrastructure
- `smb://10.40.20.185/files`
- `http://10.40.20.185:8080`
- SFTP as user `adam` on port 22
The live share is still the management-account CDK stack until cutover proof. The prod replacement is HCP Terraform in `terraform/`, workspace `file-share-prod`, trigger `terraform/**`. CDK deploy on push to main is frozen.
The instance has no public IP. Its route table sends `10.10.0.0/16` (Ronkonkoma) and `10.30.0.0/16` (Locust) through the VPN gateway in workspace variable `vpn_gateway_id`, and everything else through a NAT gateway in the syslog public subnet. Ingress is TCP 445, 8080, and 22 from those two office ranges only.
| Path | Role |
## Layout
| Path | What it is |
|---|---|
| `terraform/` | Prod EC2, subnet in the syslog VPC, DLM, and HCP roles |
| `lib/file-share-stack.ts` | Management-account CDK stack, still the live path until decommission |
| `terraform/` | Live infrastructure. Subnet `10.40.20.0/24` in the syslog VPC, security group, instance role, DLM, and the instance plus volume attachment. |
| `lib/`, `bin/` | Retired management-account CDK stack. Do not deploy it. The stack was deleted on 2026-09-29. |
```
bin/app.ts # CDK app entry point — instantiates the stack
lib/file-share-stack.ts # FileShareStack — all resource definitions
cdk.json # App command + CDK feature flags / context
cdk.context.json # Cached VPC/subnet lookups and the resolved AL2023 AMI
```
The data volume is not created by Terraform. Set `data_volume_id` on the workspace to the existing volume id (`vol-0f873de6adb59745f`). Terraform attaches it at `/dev/xvdf` and the boot script mounts the existing filesystem at `/data`. A `blkid` guard keeps a disk that already has a filesystem from being formatted.
`cdk.json` is the CDK app manifest. Its `app` command (`npx tsx bin/app.ts`) runs the TypeScript entry point directly via [tsx](https://github.com/privatenumber/tsx) — no separate compile step is needed for synth or deploy. `bin/app.ts` instantiates `FileShareStack` with an explicit `stackName: "file-share"` (kebab-case, per convention) and the target account/region.
## Apply
`aws-cdk-lib` is pinned to an exact version (`2.261.0`); the `aws-cdk` CLI is available as a dev dependency and via `npx cdk`.
Workspace `file-share-prod` is manual apply. Auto-apply stays off until the share has soaked. `user_data_replace_on_change` is false, so an AMI or user-data change does not replace the instance by itself. Snapshot the data volume and confirm before any apply that would replace the instance.
Common commands (also exposed as npm scripts):
| Command | Purpose |
|---|---|
| `npx cdk synth` (`npm run synth`) | Synthesize the CloudFormation template to `cdk.out/` |
| `npx cdk diff` (`npm run diff`) | Diff the synthesized stack against what is deployed |
| `npx cdk deploy` (`npm run deploy`) | Deploy the stack |
| `npm run build` | Type-check via `tsc` |
## Architecture
- **EC2** — `t4g.small` (ARM64, Amazon Linux 2023) in the private subnet. The AMI is cached in the committed `cdk.context.json` (`cachedInContext: true`), so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (`cdk context --reset <ami key> && cdk synth`).
- **Samba** — SMB file share at `/data/share`, optimized for macOS (`vfs_fruit`)
- **FileBrowser** — Web UI on port 8080, backed by the same `/data/share` directory
- **SFTP** — password auth for user `adam` (`ForceCommand internal-sftp`, same password as SMB); used by Hazel for automated uploads
- **EBS data volume** — `vol-04d951cccacc435b5`, 500 GiB gp3 encrypted, mounted at `/data`. **Unmanaged import**: the stack references it by ID (`Volume.fromVolumeAttributes` + `CfnVolumeAttachment`), so CloudFormation can attach it but can never create, replace, or delete it — the data survives instance replacement and even stack deletion. UserData waits for the attachment, then mounts the existing filesystem; a `blkid` guard ensures a disk that already has a filesystem is never formatted.
- **DLM** — Daily EBS snapshots at 06:00 UTC, 30-day retention (targets the `file-share-backup=true` tag, set directly on the volume)
- **SSM** — Session Manager for instance access (no SSH key)
> **History:** the data volume was originally an inline `blockDevice`, which destroyed data on instance replacement (2026-05-27 incident), then a stack-managed standalone volume, which was orphaned when an uncommitted deploy got reverted by CD (2026-06-05 incident). The unmanaged-import design ends that failure class.
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `file-share` stack is represented there as a Mermaid subgraph.
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
## Access
Requires VPN connection to the office network (10.10.0.0/16).
### Finder (SMB)
1. Finder > Go > Connect to Server
2. Enter `smb://<private-ip>/files`
3. Authenticate with `adam` and the password from `file-share/smb-password` in Secrets Manager
### FileBrowser (Web)
Open `http://<private-ip>:8080` in a browser.
The nightly DLM policy targets volumes tagged `file-share-backup=true` and keeps 30 snapshots.
## Secrets
Both stored in AWS Secrets Manager:
The instance role reads these secrets in seahaven-prod at boot. Do not put the values in Terraform:
| Secret | Purpose |
|---|---|
| `file-share/smb-password` | Samba user password |
| `file-share/filebrowser-password` | FileBrowser admin password |
Create these secrets before deploying the stack:
Copies of the same secret names still exist in the management account until 2026-10-06.
```bash
aws secretsmanager create-secret --name file-share/smb-password --secret-string '<password>'
aws secretsmanager create-secret --name file-share/filebrowser-password --secret-string '<password>'
```
## Rollback hold
## Deploy
The management CloudFormation stack is deleted. These stay until 2026-10-06, then they can be deleted:
Prod changes go through HCP Terraform workspace `file-share-prod` (manual apply until the move is sealed). The bootstrap apply creates the HCP roles and the instance boundary. The following apply, using `hcptf-file-share`, creates the subnet, security group, and DLM policy. `data_volume_id` stays empty until the snapshot copy exists, so those applies do not boot an instance.
- Data volume `vol-04d951cccacc435b5` (detached)
- Snapshot `snap-090186cf7e7b49b1c`
The management-account CDK workflow no longer runs on push. `workflow_dispatch` remains for an explicit rollback of that stack.
The management deploy role `githubdeploy-file-share` is left in place. The CDK deploy workflow is gone so a dispatch cannot recreate the stack.
The instance has no public IP. Its route table sends `10.10.0.0/16` and `10.30.0.0/16` through the VPN gateway in workspace variable `vpn_gateway_id` and everything else through a NAT gateway in the syslog public subnet. `10.10.0.0/16` is the Ronkonkoma office LAN and `10.30.0.0/16` is the Locust office LAN, the same pair the syslog VPN already routes. `10.20.0.0/16` is the management VPC and is not routed here.
## Expanding storage
Clients use the private IP. Office routing must include `10.40.20.0/24` on the existing syslog IPsec before SMB from the office will work. A check from `10.10.70.0/24` on 2026-09-28 reached the gateway for `10.40.10.254` and got no hop-1 reply for `10.40.20.1`. The nightly DLM policy targets volumes tagged `file-share-backup=true`. Tag the copied volume with that key at cutover.
Change the volume in seahaven-prod, then grow the filesystem. Terraform does not set the size.
## Expanding Storage
The data volume is **not managed by CloudFormation** (imported by ID), so changing a size in `lib/file-share-stack.ts` has no effect. Expand it directly — no downtime:
1. `aws ec2 modify-volume --volume-id vol-04d951cccacc435b5 --size <new-GiB>`
2. Wait for the modification to leave `modifying`: `aws ec2 describe-volumes-modifications --volume-ids vol-04d951cccacc435b5`
3. Resize the filesystem via an SSM session:
```bash
sudo resize2fs /dev/nvme1n1 # xvdf surfaces as nvme1n1 on Nitro instances
```
1. `aws ec2 modify-volume --volume-id vol-0f873de6adb59745f --size <new-GiB>`
2. Wait until the modification leaves `modifying`.
3. From an SSM session: `sudo resize2fs /dev/nvme1n1`