sh-openswe-traces/README.md
Adam Moussa 0dea585c1a
feat: sh-openswe-traces — LangSmith bulk-export trace archive
Storage-only SAM stack (S3 + KMS CMK + write-only IAM writer + Secrets Manager
holder) as the S3 destination for LangSmith Bulk Export of Open SWE traces, for
long-horizon auditing and prompt improvement (Athena over Parquet).

- template.yaml: versioned SSE-KMS bucket, access-log bucket, TLS-only policy,
  DEEP_ARCHIVE lifecycle; least-privilege LangSmith writer (bucket-wide PutObject,
  ViaService-scoped KMS, no read/delete).
- bootstrap.yaml: dedicated OIDC deploy role + least-privilege CFN exec role so CI
  never touches the shared execution role.
- CI/CD via reusable ci-python-sam / cd-sam workflows.

IAM passed GPT-4.1 cross-review + /sh-security-review (no blocking findings).
2026-07-10 15:45:23 -04:00

9.2 KiB

sh-openswe-traces

LangSmith Bulk Export destination for the Open SWE deployment. A storage-only SAM stack — no compute. LangSmith runs the export on its own schedule and writes Parquet run/trace data into this bucket; we retain it for periodic auditing and prompt/instruction improvement (query with Athena).

Why this exists

LangSmith retains traces for ~14 days. To audit agent behavior over longer horizons and mine it for prompt improvements, we own the data in S3 rather than paying for extended LangSmith retention. Capture is LangSmith-native Bulk Export; this repo is just the destination and its access controls.

Architecture

LangSmith (Bulk Export, scheduled LangSmith-side)
        │  s3:PutObject  (IAM user access key, least-privilege)
        ▼
  s3://sh-openswe-traces/langsmith/…    SSE-KMS (alias/sh-openswe-traces)
        │  lifecycle: → DEEP_ARCHIVE @ 90d
        ▼
  Athena / manual audit
Resource Name Notes
S3 bucket sh-openswe-traces BPA all-on, SSE-KMS (default), versioned, TLS-only, Retain on delete
Log bucket sh-openswe-traces-logs S3 server access logs (SSE-S3) for read attribution; logs expire 365d
KMS CMK alias/sh-openswe-traces Rotation on; bucket default + writer encrypt through it
IAM user auto-named (tag sh-openswe-langsmith-export) Write-only LangSmith writer; PutObject bucket-wide, no read/delete
Secret sh-openswe/langsmith-export-s3 Writer's access key (aws/secretsmanager key); populated post-deploy

The IAM user is not given an explicit name so the app stack deploys under CAPABILITY_IAM. It is referenced by ARN.

Deploy

Two stacks:

  • bootstrap.yaml → sh-openswe-traces-bootstrap — the CI IAM roles (OIDC deploy role + a least-privilege CFN exec role). Deployed once, manually, under admin (CAPABILITY_NAMED_IAM); rarely changes. Kept separate so CI never touches the shared github-cfn-execution-role, which is roles-only and can't create this stack's KMS key, secret, or IAM user.
  • template.yaml → sh-openswe-traces — the app (buckets, CMK, writer, secret). CI (cd-sam.yaml) deploys it on merge to main, assuming the OIDC deploy role and passing sh-openswe-traces-cfn-exec-role as --role-arn.
# One-time bootstrap (admin):
aws cloudformation deploy --template-file bootstrap.yaml \
  --stack-name sh-openswe-traces-bootstrap --capabilities CAPABILITY_NAMED_IAM \
  --region us-east-1

# App stack — CI does this on merge to main; for a local/admin deploy:
cp samconfig.toml.example samconfig.toml
sam validate --lint && sam build
sam deploy --capabilities CAPABILITY_IAM --resolve-s3

Wire CI (after bootstrap + repo exist)

  • Set repo secret AWS_DEPLOY_ROLE_ARN to the bootstrap DeployRoleArn output (arn:aws:iam::328440206208:role/githubdeploy-sh-openswe-traces).
  • deploy.yaml already passes cfn-role-arn = sh-openswe-traces-cfn-exec-role; the OIDC trust is pinned to this repo's main ref.

Post-deploy: mint and store the writer access key

The stack creates the IAM user and an empty secret; the access key is minted out-of-band so it never lands in CloudFormation state. Run this only on a trusted single-user workstation, never in CI — it handles a live credential.

USER=$(aws cloudformation describe-stacks --stack-name sh-openswe-traces \
  --query "Stacks[0].Outputs[?OutputKey=='ExportUserName'].OutputValue" --output text)

KEY_JSON=$(aws iam create-access-key --user-name "$USER" \
  | jq '{AccessKeyId: .AccessKey.AccessKeyId, SecretAccessKey: .AccessKey.SecretAccessKey}')

# Pass the secret via stdin, not argv, so it never appears in the process table.
aws secretsmanager put-secret-value \
  --secret-id sh-openswe/langsmith-export-s3 \
  --secret-string file:///dev/stdin <<<"$KEY_JSON"

unset KEY_JSON

Configure LangSmith Bulk Export

Driven by the LangSmith API (Plus/Enterprise only). Needs LS_API_KEY (LangSmith API key) and LS_TENANT (workspace id). Run after the writer key is minted/stored — the destination call validates by test-writing to the bucket.

1. Create the destination (creds pulled from Secrets Manager, never pasted):

CREDS=$(aws secretsmanager get-secret-value --secret-id sh-openswe/langsmith-export-s3 \
  --query SecretString --output text)
AKID=$(jq -r .AccessKeyId <<<"$CREDS"); SAK=$(jq -r .SecretAccessKey <<<"$CREDS")

curl -sS -X POST 'https://api.smith.langchain.com/api/v1/bulk-exports/destinations' \
  -H 'Content-Type: application/json' -H "X-API-Key: $LS_API_KEY" -H "X-Tenant-Id: $LS_TENANT" \
  --data @- <<JSON | jq .
{ "destination_type": "s3", "display_name": "sh-openswe-traces us-east-1",
  "config": { "bucket_name": "sh-openswe-traces", "prefix": "langsmith", "region": "us-east-1" },
  "credentials": { "access_key_id": "$AKID", "secret_access_key": "$SAK" } }
JSON

display_name must match ^[a-zA-Z0-9\-_ ']+$ (no parens); omit endpoint_url (S3-native, not GCS/MinIO). Save the returned destination id. LangSmith writes header-less, so bucket-default SSE-KMS encrypts every object under the CMK (confirm via head-object).

2. One scheduled export per tracing project (this workspace has 4). Get project UUIDs from GET /api/v1/sessions, then:

export LS_DEST_ID='<destination id>'
PROJECTS=( '<uuid-1>' '<uuid-2>' '<uuid-3>' '<uuid-4>' )   # all 4, or the subset you audit
for PID in "${PROJECTS[@]}"; do
  curl -sS -X POST 'https://api.smith.langchain.com/api/v1/bulk-exports' \
    -H 'Content-Type: application/json' -H "X-API-Key: $LS_API_KEY" -H "X-Tenant-Id: $LS_TENANT" \
    --data @- <<JSON | jq '{id, session_id, status}'
{ "bulk_export_destination_id": "$LS_DEST_ID", "session_id": "$PID",
  "start_time": "2026-06-26T00:00:00Z", "interval_hours": 24, "format_version": "v2_beta" }
JSON
done

start_time ~14d back backfills each project's retained window, then it continues daily. Keep inputs/outputs (omit export_fields) — they're the point of the audit. Data lands partitioned per project: langsmith/export_id=…/…/session_id=<id>/….

Monitor exports

# Every export in the workspace: id, project, schedule, status
curl -sS 'https://api.smith.langchain.com/api/v1/bulk-exports' \
  -H "X-API-Key: $LS_API_KEY" -H "X-Tenant-Id: $LS_TENANT" \
  | jq -r '(.bulk_exports // .exports // .)[]
           | [.id,
              (.session_id // "all_experiments"),
              (if .interval_hours then "every \(.interval_hours)h" else "one-off" end),
              .status] | @tsv' | column -t
  • A recurring schedule's own status is IntervalScheduled; the daily child exports it spawns have CREATED / RUNNING / COMPLETED / FAILED / CANCELLED / TIMEDOUT (child exports carry source_bulk_export_id).
  • One export's detail: GET /api/v1/bulk-exports/<id> — its per-run rows: GET /api/v1/bulk-exports/<id>/runs.
  • Stop a schedule: PATCH /api/v1/bulk-exports/<id> with {"status":"Cancelled"}. Already-spawned child exports must be cancelled separately, and a cancelled job can't be restarted — create a new one.

Operations

  • Rotate the writer access key quarterly: aws iam create-access-key, update the secret + LangSmith destination (PATCH/recreate), then delete the old key. Key age is monitored account-wide by the Security Hub ACCESS_KEYS_ROTATED Config rule (flags at 90d) — rotation keeps it compliant.
  • Audit: point Athena at s3://sh-openswe-traces/langsmith/ (Parquet). Objects older than 90 days are in Deep Archive — restore before querying.
  • Read attribution: object access is logged to s3://sh-openswe-traces-logs/s3-access/ (S3 server access logging).
  • Cost: Deep Archive ≈ $1/TB/mo; expect the archive to dominate storage cost.

Gotchas (from IAM cross-review)

  • Writer needs kms:Decrypt. SSE-KMS multipart uploads call kms:Decrypt at CompleteMultipartUpload; without it every multipart export fails AccessDenied. It's granted and safe — the writer has no s3:GetObject, so nothing to exfiltrate.
  • No ACL headers. The bucket is BucketOwnerEnforced; any PutObject carrying an ACL header is rejected with AccessControlListNotSupported (not fixable in policy). boto3/most SDKs send none by default — verify LangSmith's exporter likewise.
  • Downstream readers need their own CMK grant. An Athena/Glue role reading the Parquet needs explicit kms:Decrypt (+ kms:GenerateDataKey) on alias/sh-openswe-traces — see reference_kms_cmk_grant_migration.

Security

Traces can contain source code and secrets surfaced in tool I/O. Controls: SSE-KMS at rest (customer-managed CMK, bucket default), versioning (overwrite recovery), Block Public Access, TLS-only bucket policy, a write-only least-privilege writer (bucket-wide PutObject

  • kms:Decrypt gated to kms:ViaService=s3, no read/delete), the credential secret on a separate managed key, and S3 access logging. CI deploys via a dedicated least-privilege exec role (see bootstrap.yaml), not the shared execution role.

The IAM surface passed GPT-4.1 cross-review and a /sh-security-review fan-out + proof-or-kill verifier (no blocking findings). Deferred, non-blocking: CloudTrail S3 data-events (org trail carries none; access logging covers attribution for now).