syslog-server/README.md

83 lines
4 KiB
Markdown
Raw Normal View History

# syslog-server
![Terraform](https://img.shields.io/badge/Terraform-844FBA?logo=terraform&logoColor=white)
![AWS](https://img.shields.io/badge/AWS-FF9900?logo=amazonaws&logoColor=white)
![CI](https://github.com/Sea-Haven-Industries/syslog-server/actions/workflows/ci.yaml/badge.svg)
EC2 collector that receives remote syslog (UDP/TCP 514) from the office UniFi
fleet over an Elastic IP and ships it to the `unifi-syslog` CloudWatch Logs
group via the CloudWatch agent.
Deploy path (PLAT-78): HCP Terraform in seahaven-prod (`011934824531`),
workspace `syslog-server-prod`. CDK CD in mgmt is retired.
## Architecture
```
office UniFi devices ──syslog/514──▶ EIP (prod) ──▶ EC2 (rsyslog)
│
/var/log/remote/<host>/*.log
│
CloudWatch agent ──▶ unifi-syslog (90d)
│
Syslog-NoIncomingLogs alarm ──▶ site-alerts
```
| Resource | Value |
|---|---|
| Account / region | seahaven-prod `011934824531` / us-east-1 |
| HCP workspace | `syslog-server-prod` (project `seahaven-prod`; VCS `main`; working dir `terraform`; trigger `terraform/**`) |
| HCP plan/apply roles | `hcptf-syslog-server-plan` / `hcptf-syslog-server` |
| Instance | `syslog-server`, t4g.nano, Amazon Linux 2023 (arm64), 30 GiB encrypted gp3 |
| VPC | dedicated `10.40.0.0/16`, public subnet `10.40.10.0/24` |
| Elastic IP | Terraform-managed (see HCP output `public_ip`) |
| Security group | `syslog-server` — 514 tcp/udp + 22 from office IPs + VPC/VPN CIDRs; 2055/2056 udp from office |
| Instance IAM | `/tf-managed/syslog-server-role` with `syslog-server-instance-boundary`; `AmazonSSMManagedInstanceCore` + `CloudWatchAgentServerPolicy` |
| Log group | `unifi-syslog` (90-day retention) |
| Alarms | `Syslog-NoIncomingLogs`, `EC2-StatusCheck-syslog-server`, `EC2-StatusCheckSystem-syslog-server-recover` → `site-alerts` |
## Access
SSM Session Manager (no key pair). SSH 22 is open from office/VPC for
break-glass only.
## HCP first apply
First apply uses the hcptf-bootstrap window (exact `StringEquals` trust, never
`StringLike`):
1. Create the HCP workspace. Auto-apply off. No project-level variable set.
Working directory `terraform`. File trigger prefix `terraform/**` only.
Speculative plans on. VCS on `main`.
2. From `seahaven-org-baseline`:
`scripts/create-hcptf-bootstrap-roles.sh --account prod --allow-workspace syslog-server-prod`
3. Point workspace `TFC_AWS_APPLY_ROLE_ARN` / `TFC_AWS_PLAN_ROLE_ARN` at
`hcptf-bootstrap` / `hcptf-bootstrap-plan`. Set `TFC_AWS_PROVIDER_AUTH=true`.
Never `TFC_AWS_RUN_ROLE_ARN`.
4. One manual apply. This creates the scoped `hcptf-*` roles, the instance
boundary, VPC, instance, EIP, log group, and alarms.
5. Retarget `TFC_AWS_*` to `hcptf-syslog-server` / `hcptf-syslog-server-plan`.
Re-run the create script with no `--allow-workspace`.
6. Second manual apply as the scoped role. After live-path proof, seal
auto-apply on.
`Syslog-NoIncomingLogs` defaults `treat_missing_data` to `notBreaching` so the
empty prod log group does not page `site-alerts` before UniFi is re-pointed.
After devices deliver to the new EIP, set `no_logs_treat_missing_data=breaching`.
HCP outputs to copy: `public_ip`, `instance_id`, `hcptf_apply_role_arn`,
`hcptf_plan_role_arn`.
AMI is pinned in `var.ami_id`. An AMI change forces instance replacement.
User-data changes also replace the instance (the box is stateless; logs live
in CloudWatch; the EIP re-associates).
## Documentation
The canonical map of Sea Haven's AWS infrastructure lives in Confluence.
- **[AWS Architecture Map](https://seahaven.atlassian.net/wiki/spaces/IT/pages/1540098)** (Confluence, IT space, page 1540098)
To widen device coverage of the forwarded syslog feed, see **INFRA-11**
(UniFi controller remote-logging config).