docs(infra): record the prod HCP path after management decommission (PLAT-77)

The management stack is deleted, so the docs now describe the live share and the CDK deploy workflow is removed.
This commit is contained in:
Adam Moussa 2026-09-29 19:28:26 -04:00
parent 9f27827dcb
commit 8768771de1
No known key found for this signature in database
3 changed files with 50 additions and 140 deletions

View file

@ -1,22 +0,0 @@
name: Deploy
# PLAT-77: push-to-main CDK deploy is frozen so HCP Terraform is the only
# path that can change this stack. workflow_dispatch stays for an explicit
# rollback of the management-account stack.
on:
workflow_dispatch:
permissions:
id-token: write
contents: read
concurrency:
group: deploy
cancel-in-progress: false
jobs:
deploy:
uses: Sea-Haven-Industries/.github/.github/workflows/cd-cdk.yaml@0a1010e63248c9ca9f042c870eb2c579ba6b9455 # v1.0.19
with:
node-version: "24"
secrets:
deploy-role-arn: ${{ secrets.AWS_DEPLOY_ROLE_ARN }}

View file

@ -2,56 +2,39 @@
Instructions for coding agents working in this repository.
## Live path
The share runs in seahaven-prod under HCP Terraform workspace `file-share-prod`. Source is `terraform/`. Manual apply until the move is sealed. Do not enable auto-apply as part of a docs or cleanup change.
The management-account CDK stack was deleted on 2026-09-29. Do not run `cdk deploy`. `lib/` and `bin/` are the retired stack. The CDK deploy workflow has been removed.
The data volume is attached by the workspace variable `data_volume_id`. Terraform must not create or delete it. The previous management volume is retained until 2026-10-06 and is not the live disk.
## Infrastructure as Code principles
- **CDK + EC2**: this repo deploys an EC2 instance via CDK TypeScript. No Lambda, no SAM.
- **Lambda defaults (does not apply)** — this is a pure EC2 stack.
- **Exact-pin all CDK library versions** (`aws-cdk-lib`, `aws-cdk`, `constructs`). Never use `*` or `^` ranges.
- **Never commit** account IDs, role ARNs, VPC IDs, subnet IDs, or volume IDs in new code. The existing `cdk.context.json` already contains resolved values — do not add new deployment-specific identifiers without documented defaults.
- **Deploy with least-privilege IAM.** The instance role already has only `AmazonSSMManagedInstanceCore` + Secrets Manager read on `file-share/*`.
- **Verify**: `npm run build && npx cdk synth && npx cdk diff` before pushing. Do not commit `cdk.out/`.
- **Exact-pin CDK library versions** if you touch the retired CDK package. Never use `*` or `^` ranges.
- **Never commit** account IDs, role ARNs, VPC IDs, subnet IDs, or new volume IDs. The live volume id is an HCP variable.
- **Deploy with least-privilege IAM.** The instance role has SSM core plus Secrets Manager read on `file-share/*`.
- Do not commit `cdk.out/`.
## CDK directory layout
## Instance replacement
```
.
├── bin/
│ └── app.ts # CDK app entrypoint
├── lib/
│ └── file-share-stack.ts # All resource definitions
├── cdk.json # CDK context + config
├── cdk.context.json # Resolved context (committed)
├── package.json # Exact-pinned CDK dependencies
├── tsconfig.json
└── .github/
└── workflows/ # CI/CD
```
`user_data_replace_on_change` is false. Before any apply that would replace the instance (AMI, user data, instance type):
## `overrideLogicalId` — never remove
Removing `overrideLogicalId` on a deployed resource forces replacement. Never remove an existing call without documenting the replacement impact.
## Persistent EBS volumes
This repo's data volume (`vol-04d951cccacc435b5`) is **imported by ID**, not managed by CloudFormation. It survives instance replacement and stack deletion.
### Snapshot before an EC2-replacing deploy
Before any deploy that would replace the EC2 instance (AMI change, user-data change, instance type change):
1. Run `cdk diff` to confirm the instance will be replaced.
2. **Stop and ask for confirmation** — an instance replacement detaches the imported volume. A snapshot is the safety net.
3. If a DLM snapshot already exists from the same day, reference it rather than creating a duplicate.
1. Read the plan and confirm the instance is being replaced.
2. Stop and ask for confirmation. Replacement detaches the data volume.
3. Snapshot the volume first. If a DLM snapshot from the same day exists, use that instead of a second copy.
## Security group rules
- Never open 0.0.0.0/0. All rules use specific CIDR ranges (`10.10.0.0/16` VPN, `10.20.0.0/16` VPC).
- SSM Session Manager for instance access — no SSH key, no public port 22.
- No `0.0.0.0/0` ingress. Ingress is TCP 445, 8080, and 22 from `10.10.0.0/16` and `10.30.0.0/16`.
- Egress `0.0.0.0/0` stays so the instance can install packages and reach the pinned FileBrowser download.
- SSM Session Manager for instance access. No SSH key and no public port 22. SFTP for user `adam` is the office path, not an admin login.
## Secrets
Stored in AWS Secrets Manager (`file-share/smb-password`, `file-share/filebrowser-password`). The instance role reads them at boot. Do not embed secrets in UserData or source files.
Stored in AWS Secrets Manager (`file-share/smb-password`, `file-share/filebrowser-password`). The instance role reads them at boot. Do not embed secrets in user data or source files. Do not delete the management-account copies before 2026-10-06.
## Documentation
The Confluence "AWS Architecture Map" (page 1540098) should be updated alongside any architecture change.
The Confluence "AWS Architecture Map" (page 1540098) should be updated alongside any architecture change.

107
README.md
View file

@ -1,105 +1,54 @@
# file-share
# File Share
![CI](https://github.com/Sea-Haven-Industries/file-share/actions/workflows/ci.yaml/badge.svg)
![TypeScript](https://img.shields.io/badge/TypeScript-3178C6?logo=typescript&logoColor=white)
![AWS CDK](https://img.shields.io/badge/AWS-CDK-FF9900?logo=amazonaws&logoColor=white)
Samba, FileBrowser, and SFTP for the Sea Haven offices. The live host is an EC2 instance in seahaven-prod, managed by HCP Terraform workspace `file-share-prod`.
Personal file share server on AWS — Samba for macOS Finder integration and FileBrowser for web-based file management. Accessible exclusively over the site-to-site VPN.
Clients:
## Infrastructure
- `smb://10.40.20.185/files`
- `http://10.40.20.185:8080`
- SFTP as user `adam` on port 22
The live share is still the management-account CDK stack until cutover proof. The prod replacement is HCP Terraform in `terraform/`, workspace `file-share-prod`, trigger `terraform/**`. CDK deploy on push to main is frozen.
The instance has no public IP. Its route table sends `10.10.0.0/16` (Ronkonkoma) and `10.30.0.0/16` (Locust) through the VPN gateway in workspace variable `vpn_gateway_id`, and everything else through a NAT gateway in the syslog public subnet. Ingress is TCP 445, 8080, and 22 from those two office ranges only.
| Path | Role |
## Layout
| Path | What it is |
|---|---|
| `terraform/` | Prod EC2, subnet in the syslog VPC, DLM, and HCP roles |
| `lib/file-share-stack.ts` | Management-account CDK stack, still the live path until decommission |
| `terraform/` | Live infrastructure. Subnet `10.40.20.0/24` in the syslog VPC, security group, instance role, DLM, and the instance plus volume attachment. |
| `lib/`, `bin/` | Retired management-account CDK stack. Do not deploy it. The stack was deleted on 2026-09-29. |
```
bin/app.ts # CDK app entry point — instantiates the stack
lib/file-share-stack.ts # FileShareStack — all resource definitions
cdk.json # App command + CDK feature flags / context
cdk.context.json # Cached VPC/subnet lookups and the resolved AL2023 AMI
```
The data volume is not created by Terraform. Set `data_volume_id` on the workspace to the existing volume id (`vol-0f873de6adb59745f`). Terraform attaches it at `/dev/xvdf` and the boot script mounts the existing filesystem at `/data`. A `blkid` guard keeps a disk that already has a filesystem from being formatted.
`cdk.json` is the CDK app manifest. Its `app` command (`npx tsx bin/app.ts`) runs the TypeScript entry point directly via [tsx](https://github.com/privatenumber/tsx) — no separate compile step is needed for synth or deploy. `bin/app.ts` instantiates `FileShareStack` with an explicit `stackName: "file-share"` (kebab-case, per convention) and the target account/region.
## Apply
`aws-cdk-lib` is pinned to an exact version (`2.261.0`); the `aws-cdk` CLI is available as a dev dependency and via `npx cdk`.
Workspace `file-share-prod` is manual apply. Auto-apply stays off until the share has soaked. `user_data_replace_on_change` is false, so an AMI or user-data change does not replace the instance by itself. Snapshot the data volume and confirm before any apply that would replace the instance.
Common commands (also exposed as npm scripts):
| Command | Purpose |
|---|---|
| `npx cdk synth` (`npm run synth`) | Synthesize the CloudFormation template to `cdk.out/` |
| `npx cdk diff` (`npm run diff`) | Diff the synthesized stack against what is deployed |
| `npx cdk deploy` (`npm run deploy`) | Deploy the stack |
| `npm run build` | Type-check via `tsc` |
## Architecture
- **EC2** — `t4g.small` (ARM64, Amazon Linux 2023) in the private subnet. The AMI is cached in the committed `cdk.context.json` (`cachedInContext: true`), so deploys never pick up a new AL2023 release implicitly — an AMI change forces instance replacement and must be deliberate (`cdk context --reset <ami key> && cdk synth`).
- **Samba** — SMB file share at `/data/share`, optimized for macOS (`vfs_fruit`)
- **FileBrowser** — Web UI on port 8080, backed by the same `/data/share` directory
- **SFTP** — password auth for user `adam` (`ForceCommand internal-sftp`, same password as SMB); used by Hazel for automated uploads
- **EBS data volume** — `vol-04d951cccacc435b5`, 500 GiB gp3 encrypted, mounted at `/data`. **Unmanaged import**: the stack references it by ID (`Volume.fromVolumeAttributes` + `CfnVolumeAttachment`), so CloudFormation can attach it but can never create, replace, or delete it — the data survives instance replacement and even stack deletion. UserData waits for the attachment, then mounts the existing filesystem; a `blkid` guard ensures a disk that already has a filesystem is never formatted.
- **DLM** — Daily EBS snapshots at 06:00 UTC, 30-day retention (targets the `file-share-backup=true` tag, set directly on the volume)
- **SSM** — Session Manager for instance access (no SSH key)
> **History:** the data volume was originally an inline `blockDevice`, which destroyed data on instance replacement (2026-05-27 incident), then a stack-managed standalone volume, which was orphaned when an uncommitted deploy got reverted by CD (2026-06-05 incident). The unmanaged-import design ends that failure class.
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence. This project's `file-share` stack is represented there as a Mermaid subgraph.
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
## Access
Requires VPN connection to the office network (10.10.0.0/16).
### Finder (SMB)
1. Finder > Go > Connect to Server
2. Enter `smb://<private-ip>/files`
3. Authenticate with `adam` and the password from `file-share/smb-password` in Secrets Manager
### FileBrowser (Web)
Open `http://<private-ip>:8080` in a browser.
The nightly DLM policy targets volumes tagged `file-share-backup=true` and keeps 30 snapshots.
## Secrets
Both stored in AWS Secrets Manager:
The instance role reads these secrets in seahaven-prod at boot. Do not put the values in Terraform:
| Secret | Purpose |
|---|---|
| `file-share/smb-password` | Samba user password |
| `file-share/filebrowser-password` | FileBrowser admin password |
Create these secrets before deploying the stack:
Copies of the same secret names still exist in the management account until 2026-10-06.
```bash
aws secretsmanager create-secret --name file-share/smb-password --secret-string '<password>'
aws secretsmanager create-secret --name file-share/filebrowser-password --secret-string '<password>'
```
## Rollback hold
## Deploy
The management CloudFormation stack is deleted. These stay until 2026-10-06, then they can be deleted:
Prod changes go through HCP Terraform workspace `file-share-prod` (manual apply until the move is sealed). The bootstrap apply creates the HCP roles and the instance boundary. The following apply, using `hcptf-file-share`, creates the subnet, security group, and DLM policy. `data_volume_id` stays empty until the snapshot copy exists, so those applies do not boot an instance.
- Data volume `vol-04d951cccacc435b5` (detached)
- Snapshot `snap-090186cf7e7b49b1c`
The management-account CDK workflow no longer runs on push. `workflow_dispatch` remains for an explicit rollback of that stack.
The management deploy role `githubdeploy-file-share` is left in place. The CDK deploy workflow is gone so a dispatch cannot recreate the stack.
The instance has no public IP. Its route table sends `10.10.0.0/16` and `10.30.0.0/16` through the VPN gateway in workspace variable `vpn_gateway_id` and everything else through a NAT gateway in the syslog public subnet. `10.10.0.0/16` is the Ronkonkoma office LAN and `10.30.0.0/16` is the Locust office LAN, the same pair the syslog VPN already routes. `10.20.0.0/16` is the management VPC and is not routed here.
## Expanding storage
Clients use the private IP. Office routing must include `10.40.20.0/24` on the existing syslog IPsec before SMB from the office will work. A check from `10.10.70.0/24` on 2026-09-28 reached the gateway for `10.40.10.254` and got no hop-1 reply for `10.40.20.1`. The nightly DLM policy targets volumes tagged `file-share-backup=true`. Tag the copied volume with that key at cutover.
Change the volume in seahaven-prod, then grow the filesystem. Terraform does not set the size.
## Expanding Storage
The data volume is **not managed by CloudFormation** (imported by ID), so changing a size in `lib/file-share-stack.ts` has no effect. Expand it directly — no downtime:
1. `aws ec2 modify-volume --volume-id vol-04d951cccacc435b5 --size <new-GiB>`
2. Wait for the modification to leave `modifying`: `aws ec2 describe-volumes-modifications --volume-ids vol-04d951cccacc435b5`
3. Resize the filesystem via an SSM session:
```bash
sudo resize2fs /dev/nvme1n1 # xvdf surfaces as nvme1n1 on Nitro instances
```
1. `aws ec2 modify-volume --volume-id vol-0f873de6adb59745f --size <new-GiB>`
2. Wait until the modification leaves `modifying`.
3. From an SSM session: `sudo resize2fs /dev/nvme1n1`