mirror of
https://github.com/Sea-Haven-Industries/proposal-system.git
synced 2026-10-05 21:12:01 +00:00
feat(infra): migrate Bedrock KB vector store to Aurora pgvector (#125)
Some checks failed
Deploy / Deploy to AWS (push) Has been cancelled
Some checks failed
Deploy / Deploy to AWS (push) Has been cancelled
Replaces OpenSearch Serverless with Aurora PostgreSQL Serverless v2 + pgvector as the Bedrock Knowledge Base vector store (v1 PR3). Bedrock KB requires Aurora SSv2 (RDS Data API), not a plain RDS instance — so the DB engine moves to Aurora. - foundation: rds.DatabaseInstance (PG15) -> rds.DatabaseCluster Aurora SSv2 (0.5-4 ACU, enableDataApi). RDS alarms: free-storage -> freeable-memory. - compute: delete all AOSS (collection, policies, VPC endpoint, index-creator); add bedrock_user secret + KB role (scoped rds-data + secret read); repoint CfnKnowledgeBase to RDS storage (bedrock_integration.bedrock_kb, vector(1024)). - lambdas: oss-index-creator -> aurora-pgvector-init (bootstrap schema/table/ indexes/role via RDS Data API; transient-error retry; password guard). - ADR 0001 documents the decision. Eliminates the ~$175-350/mo AOSS OCU floor. NAT kept (egress still needed). GPT-4.1 cross-review: no BLOCK (FIX applied). tsc clean; foundation synth shows Aurora cluster with Data API enabled; 23 pytest pass.
This commit is contained in:
parent
bbd185b280
commit
aff38a11ae
8 changed files with 283 additions and 233 deletions
68
docs/adr/0001-bedrock-vector-store-aurora-pgvector.md
Normal file
68
docs/adr/0001-bedrock-vector-store-aurora-pgvector.md
Normal file
|
|
@ -0,0 +1,68 @@
|
||||||
|
# ADR 0001 — Bedrock Knowledge Base vector store: Aurora PostgreSQL + pgvector
|
||||||
|
|
||||||
|
- **Status:** Accepted (2026-06-12)
|
||||||
|
- **Decision owner:** Adam Moussa
|
||||||
|
- **Scope:** `proposal-system` infrastructure (foundation + compute stacks), v1 PR3
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The proposal system's RAG pipeline uses an Amazon Bedrock Knowledge Base (Titan Embed
|
||||||
|
v2) over historical proposals. The original vector store was **OpenSearch Serverless
|
||||||
|
(AOSS)**, which:
|
||||||
|
|
||||||
|
- carries a minimum ~2-OCU billing floor (~$175–350/mo) even at near-zero query volume —
|
||||||
|
the dominant line item of the monthly bill for an internal tool;
|
||||||
|
- is VPC-only, which forces the NAT gateway and a `oss-index-creator` bootstrap Lambda
|
||||||
|
that exists purely to pre-create the vector index while an IAM access policy propagates.
|
||||||
|
|
||||||
|
The database is a small RDS PostgreSQL instance. The v1 assessment flagged AOSS as
|
||||||
|
over-built for the corpus size and recommended a pgvector store on the existing database.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
Migrate the database from **RDS PostgreSQL 15 → Aurora PostgreSQL Serverless v2** and use
|
||||||
|
**pgvector** as the Bedrock KB vector store.
|
||||||
|
|
||||||
|
Aurora is required because **Bedrock Knowledge Bases support Aurora PostgreSQL
|
||||||
|
Serverless v2 (with the RDS Data API) as a pgvector store, but not a plain RDS
|
||||||
|
instance.** A standard RDS instance cannot back a Bedrock KB, so "pgvector on the
|
||||||
|
existing RDS" was infeasible without the Aurora move.
|
||||||
|
|
||||||
|
Implementation:
|
||||||
|
|
||||||
|
- Aurora Serverless v2 (min 0.5 / max 4 ACU), `enableDataApi: true`, `defaultDatabaseName:
|
||||||
|
'proposals'`. The cluster also serves the .NET API's application data (one database).
|
||||||
|
- A bootstrap custom-resource Lambda (`aurora-pgvector-init`, via the RDS Data API)
|
||||||
|
enables `vector`, creates the `bedrock_integration.bedrock_kb` table
|
||||||
|
(`vector(1024)` for Titan v2 + HNSW cosine + GIN indexes) and a dedicated
|
||||||
|
`bedrock_user` role. This **replaces** `oss-index-creator` — the bootstrap is swapped,
|
||||||
|
not eliminated.
|
||||||
|
- `CfnKnowledgeBase.storageConfiguration` → `type: 'RDS'` with `rdsConfiguration`
|
||||||
|
(cluster ARN, `bedrock_user` secret, `bedrock_integration.bedrock_kb`, field mapping
|
||||||
|
`id`/`embedding`/`chunks`/`metadata`). KB role IAM swaps `aoss:APIAccessAll` →
|
||||||
|
scoped `rds-data` + secret read.
|
||||||
|
- All AOSS constructs (collection, policies, VPC endpoint, index-creator) are deleted.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- **Cost:** eliminates the AOSS OCU floor (~$175–350/mo). Aurora Serverless v2 at
|
||||||
|
0.5 ACU min is ~$43/mo and scales toward zero idle — a net reduction.
|
||||||
|
- **Simplification:** one data engine (Aurora) instead of RDS + AOSS; fewer constructs.
|
||||||
|
A bootstrap Lambda remains (now for pgvector schema rather than the AOSS index).
|
||||||
|
- **NAT:** kept for now — Lambdas and the KB's Data API path still need AWS-service
|
||||||
|
egress. Dropping NAT would require VPC interface endpoints; tracked separately.
|
||||||
|
- **Migration:** the RDS→Aurora swap is a CloudFormation replacement. The deployed stacks
|
||||||
|
are **test-only with no production data**, so this is a clean redeploy.
|
||||||
|
|
||||||
|
## Alternatives considered
|
||||||
|
|
||||||
|
- **Keep AOSS:** rejected — the cost floor is the single biggest waste for the scale.
|
||||||
|
- **S3 Vectors:** viable Bedrock backend, but Aurora unifies app data + vectors and was
|
||||||
|
the owner's preference (also cheaper than AOSS).
|
||||||
|
- **pgvector on the existing RDS instance:** infeasible — Bedrock KB does not support a
|
||||||
|
plain RDS instance as a vector store.
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [Using Aurora PostgreSQL as a Bedrock Knowledge Base](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/AuroraPostgreSQL.VectorDB.html)
|
||||||
|
- [Bedrock KB vector-store prerequisites](https://docs.aws.amazon.com/bedrock/latest/userguide/knowledge-base-setup.html)
|
||||||
|
|
@ -26,6 +26,7 @@ const compute = new ComputeStack(app, `proposal-system-compute${stackSuffix}`, {
|
||||||
vpc: foundation.vpc,
|
vpc: foundation.vpc,
|
||||||
lambdaSecurityGroup: foundation.lambdaSecurityGroup,
|
lambdaSecurityGroup: foundation.lambdaSecurityGroup,
|
||||||
dbSecret: foundation.dbSecret,
|
dbSecret: foundation.dbSecret,
|
||||||
|
dbCluster: foundation.dbCluster,
|
||||||
uploadsBucket: foundation.uploadsBucket,
|
uploadsBucket: foundation.uploadsBucket,
|
||||||
generatedBucket: foundation.generatedBucket,
|
generatedBucket: foundation.generatedBucket,
|
||||||
libraryBucket: foundation.libraryBucket,
|
libraryBucket: foundation.libraryBucket,
|
||||||
|
|
|
||||||
|
|
@ -12,7 +12,7 @@ import * as cognito from 'aws-cdk-lib/aws-cognito';
|
||||||
import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager';
|
import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager';
|
||||||
import * as lambdaEventSources from 'aws-cdk-lib/aws-lambda-event-sources';
|
import * as lambdaEventSources from 'aws-cdk-lib/aws-lambda-event-sources';
|
||||||
import * as bedrock from 'aws-cdk-lib/aws-bedrock';
|
import * as bedrock from 'aws-cdk-lib/aws-bedrock';
|
||||||
import * as opensearchserverless from 'aws-cdk-lib/aws-opensearchserverless';
|
import * as rds from 'aws-cdk-lib/aws-rds';
|
||||||
import * as logs from 'aws-cdk-lib/aws-logs';
|
import * as logs from 'aws-cdk-lib/aws-logs';
|
||||||
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
|
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
|
||||||
import * as cloudwatchActions from 'aws-cdk-lib/aws-cloudwatch-actions';
|
import * as cloudwatchActions from 'aws-cdk-lib/aws-cloudwatch-actions';
|
||||||
|
|
@ -25,6 +25,7 @@ export interface ComputeStackProps extends cdk.StackProps {
|
||||||
vpc: ec2.IVpc;
|
vpc: ec2.IVpc;
|
||||||
lambdaSecurityGroup: ec2.ISecurityGroup;
|
lambdaSecurityGroup: ec2.ISecurityGroup;
|
||||||
dbSecret: secretsmanager.ISecret;
|
dbSecret: secretsmanager.ISecret;
|
||||||
|
dbCluster: rds.IDatabaseCluster;
|
||||||
uploadsBucket: s3.IBucket;
|
uploadsBucket: s3.IBucket;
|
||||||
generatedBucket: s3.IBucket;
|
generatedBucket: s3.IBucket;
|
||||||
libraryBucket: s3.IBucket;
|
libraryBucket: s3.IBucket;
|
||||||
|
|
@ -51,45 +52,23 @@ export class ComputeStack extends cdk.Stack {
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
// OpenSearch Serverless collection for Bedrock KB vector store
|
// Bedrock Knowledge Base vector store: Aurora PostgreSQL + pgvector (PR3 —
|
||||||
const ossEncryptionPolicy = new opensearchserverless.CfnSecurityPolicy(this, 'OssEncryptionPolicy', {
|
// replaced OpenSearch Serverless). A dedicated bedrock_user role queries the
|
||||||
name: `proposal-system-kb-enc${config.stackSuffix}`,
|
// pgvector table; the aurora-pgvector-init custom resource creates the schema.
|
||||||
type: 'encryption',
|
const dbCluster = props.dbCluster;
|
||||||
policy: JSON.stringify({
|
|
||||||
Rules: [{ ResourceType: 'collection', Resource: [`collection/proposal-system-kb${config.stackSuffix}`] }],
|
|
||||||
AWSOwnedKey: true,
|
|
||||||
}),
|
|
||||||
});
|
|
||||||
|
|
||||||
// Fix: INF-H3 — restrict OpenSearch Serverless to VPC (was AllowFromPublic: true).
|
// Credentials Bedrock uses to query the pgvector table as bedrock_user.
|
||||||
// Create a VPC endpoint so Lambdas in private subnets can reach the collection.
|
// SQL-unsafe characters are excluded so the bootstrap can inline the password.
|
||||||
const ossVpcEndpoint = new opensearchserverless.CfnVpcEndpoint(this, 'OssVpcEndpoint', {
|
const bedrockUserSecret = new secretsmanager.Secret(this, 'BedrockUserSecret', {
|
||||||
name: `proposal-system-kb-vpce${config.stackSuffix}`,
|
secretName: `proposal-system/bedrock-user${config.stackSuffix}`,
|
||||||
vpcId: props.vpc.vpcId,
|
generateSecretString: {
|
||||||
subnetIds: props.vpc.selectSubnets({ subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS }).subnetIds,
|
secretStringTemplate: JSON.stringify({ username: 'bedrock_user' }),
|
||||||
securityGroupIds: [props.lambdaSecurityGroup.securityGroupId],
|
generateStringKey: 'password',
|
||||||
|
excludePunctuation: true,
|
||||||
|
passwordLength: 32,
|
||||||
|
},
|
||||||
});
|
});
|
||||||
|
|
||||||
const ossNetworkPolicy = new opensearchserverless.CfnSecurityPolicy(this, 'OssNetworkPolicy', {
|
|
||||||
name: `proposal-system-kb-net${config.stackSuffix}`,
|
|
||||||
type: 'network',
|
|
||||||
policy: JSON.stringify([{
|
|
||||||
Rules: [
|
|
||||||
{ ResourceType: 'collection', Resource: [`collection/proposal-system-kb${config.stackSuffix}`] },
|
|
||||||
],
|
|
||||||
AllowFromPublic: false,
|
|
||||||
SourceVPCEs: [ossVpcEndpoint.attrId],
|
|
||||||
}]),
|
|
||||||
});
|
|
||||||
ossNetworkPolicy.addDependency(ossVpcEndpoint);
|
|
||||||
|
|
||||||
const ossCollection = new opensearchserverless.CfnCollection(this, 'OssCollection', {
|
|
||||||
name: `proposal-system-kb${config.stackSuffix}`,
|
|
||||||
type: 'VECTORSEARCH',
|
|
||||||
});
|
|
||||||
ossCollection.addDependency(ossEncryptionPolicy);
|
|
||||||
ossCollection.addDependency(ossNetworkPolicy);
|
|
||||||
|
|
||||||
// Bedrock KB execution role
|
// Bedrock KB execution role
|
||||||
const kbRole = new iam.Role(this, 'KnowledgeBaseRole', {
|
const kbRole = new iam.Role(this, 'KnowledgeBaseRole', {
|
||||||
roleName: `proposal-system-kb-role${config.stackSuffix}`,
|
roleName: `proposal-system-kb-role${config.stackSuffix}`,
|
||||||
|
|
@ -101,113 +80,56 @@ export class ComputeStack extends cdk.Stack {
|
||||||
resources: [props.libraryBucket.bucketArn, `${props.libraryBucket.bucketArn}/*`],
|
resources: [props.libraryBucket.bucketArn, `${props.libraryBucket.bucketArn}/*`],
|
||||||
}));
|
}));
|
||||||
|
|
||||||
kbRole.addToPolicy(new iam.PolicyStatement({
|
|
||||||
actions: ['aoss:APIAccessAll'],
|
|
||||||
resources: [ossCollection.attrArn],
|
|
||||||
}));
|
|
||||||
|
|
||||||
kbRole.addToPolicy(new iam.PolicyStatement({
|
kbRole.addToPolicy(new iam.PolicyStatement({
|
||||||
actions: ['bedrock:InvokeModel'],
|
actions: ['bedrock:InvokeModel'],
|
||||||
resources: [`arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0`],
|
resources: [`arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0`],
|
||||||
}));
|
}));
|
||||||
|
|
||||||
// Lambda to pre-create the vector index (retries until AOSS access policy propagates)
|
// KB queries the pgvector table via the RDS Data API as bedrock_user.
|
||||||
const indexCreatorFn = new lambda.Function(this, 'OssIndexCreator', {
|
kbRole.addToPolicy(new iam.PolicyStatement({
|
||||||
functionName: `proposal-system-oss-index-creator${config.stackSuffix}`,
|
actions: [
|
||||||
|
'rds-data:ExecuteStatement',
|
||||||
|
'rds-data:BatchExecuteStatement',
|
||||||
|
'rds-data:BeginTransaction',
|
||||||
|
'rds-data:CommitTransaction',
|
||||||
|
'rds-data:RollbackTransaction',
|
||||||
|
],
|
||||||
|
resources: [dbCluster.clusterArn],
|
||||||
|
}));
|
||||||
|
bedrockUserSecret.grantRead(kbRole);
|
||||||
|
|
||||||
|
// One-time bootstrap: enable pgvector + create the bedrock_integration schema,
|
||||||
|
// table, indexes, and bedrock_user role (via the RDS Data API as master).
|
||||||
|
const pgvectorInitFn = new lambda.Function(this, 'PgVectorInit', {
|
||||||
|
functionName: `proposal-system-pgvector-init${config.stackSuffix}`,
|
||||||
runtime: lambda.Runtime.PYTHON_3_12,
|
runtime: lambda.Runtime.PYTHON_3_12,
|
||||||
architecture: lambda.Architecture.ARM_64,
|
architecture: lambda.Architecture.ARM_64,
|
||||||
handler: 'app.handler',
|
handler: 'app.handler',
|
||||||
code: lambda.Code.fromAsset('../lambdas/oss-index-creator', {
|
code: lambda.Code.fromAsset('../lambdas/aurora-pgvector-init'),
|
||||||
bundling: {
|
timeout: cdk.Duration.minutes(5),
|
||||||
image: lambda.Runtime.PYTHON_3_12.bundlingImage,
|
environment: {
|
||||||
command: [
|
CLUSTER_ARN: dbCluster.clusterArn,
|
||||||
'bash', '-c',
|
MASTER_SECRET_ARN: props.dbSecret.secretArn,
|
||||||
'pip install -r requirements.txt -t /asset-output && cp -au . /asset-output',
|
BEDROCK_SECRET_ARN: bedrockUserSecret.secretArn,
|
||||||
],
|
DATABASE: 'proposals',
|
||||||
},
|
EMBED_DIM: '1024',
|
||||||
}),
|
},
|
||||||
timeout: cdk.Duration.minutes(6),
|
|
||||||
logRetention: logs.RetentionDays.TWO_MONTHS,
|
logRetention: logs.RetentionDays.TWO_MONTHS,
|
||||||
});
|
});
|
||||||
|
dbCluster.grantDataApiAccess(pgvectorInitFn);
|
||||||
|
bedrockUserSecret.grantRead(pgvectorInitFn);
|
||||||
|
|
||||||
indexCreatorFn.addToRolePolicy(new iam.PolicyStatement({
|
const pgvectorProvider = new cr.Provider(this, 'PgVectorInitProvider', {
|
||||||
actions: ['aoss:APIAccessAll'],
|
onEventHandler: pgvectorInitFn,
|
||||||
resources: [ossCollection.attrArn],
|
|
||||||
}));
|
|
||||||
|
|
||||||
// Fix: INF-M2 — scope AOSS data access policy permissions (was aoss:* on both
|
|
||||||
// collection and index). KB role needs read/write for embeddings. Index creator
|
|
||||||
// needs create/describe for bootstrapping the vector index.
|
|
||||||
const ossDataAccessPolicy = new opensearchserverless.CfnAccessPolicy(this, 'OssDataAccessPolicy', {
|
|
||||||
name: `proposal-system-kb-access${config.stackSuffix}`,
|
|
||||||
type: 'data',
|
|
||||||
policy: JSON.stringify([
|
|
||||||
{
|
|
||||||
Description: 'Bedrock KB role — read/write documents and describe collection',
|
|
||||||
Rules: [
|
|
||||||
{
|
|
||||||
ResourceType: 'collection',
|
|
||||||
Resource: [`collection/proposal-system-kb${config.stackSuffix}`],
|
|
||||||
Permission: [
|
|
||||||
'aoss:DescribeCollectionItems',
|
|
||||||
'aoss:CreateCollectionItems',
|
|
||||||
'aoss:UpdateCollectionItems',
|
|
||||||
],
|
|
||||||
},
|
|
||||||
{
|
|
||||||
ResourceType: 'index',
|
|
||||||
Resource: [`index/proposal-system-kb${config.stackSuffix}/*`],
|
|
||||||
Permission: [
|
|
||||||
'aoss:DescribeIndex',
|
|
||||||
'aoss:ReadDocument',
|
|
||||||
'aoss:WriteDocument',
|
|
||||||
],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
Principal: [kbRole.roleArn],
|
|
||||||
},
|
|
||||||
{
|
|
||||||
Description: 'Index creator Lambda — create and describe index during bootstrap',
|
|
||||||
Rules: [
|
|
||||||
{
|
|
||||||
ResourceType: 'collection',
|
|
||||||
Resource: [`collection/proposal-system-kb${config.stackSuffix}`],
|
|
||||||
Permission: [
|
|
||||||
'aoss:DescribeCollectionItems',
|
|
||||||
'aoss:CreateCollectionItems',
|
|
||||||
],
|
|
||||||
},
|
|
||||||
{
|
|
||||||
ResourceType: 'index',
|
|
||||||
Resource: [`index/proposal-system-kb${config.stackSuffix}/*`],
|
|
||||||
Permission: [
|
|
||||||
'aoss:CreateIndex',
|
|
||||||
'aoss:DescribeIndex',
|
|
||||||
'aoss:WriteDocument',
|
|
||||||
],
|
|
||||||
},
|
|
||||||
],
|
|
||||||
Principal: [indexCreatorFn.role!.roleArn],
|
|
||||||
},
|
|
||||||
]),
|
|
||||||
});
|
|
||||||
ossDataAccessPolicy.addDependency(ossCollection);
|
|
||||||
|
|
||||||
const indexProvider = new cr.Provider(this, 'OssIndexProvider', {
|
|
||||||
onEventHandler: indexCreatorFn,
|
|
||||||
});
|
});
|
||||||
|
|
||||||
const ossIndex = new cdk.CustomResource(this, 'OssIndex', {
|
const pgvectorInit = new cdk.CustomResource(this, 'PgVectorInitResource', {
|
||||||
serviceToken: indexProvider.serviceToken,
|
serviceToken: pgvectorProvider.serviceToken,
|
||||||
properties: {
|
properties: {
|
||||||
Endpoint: ossCollection.attrCollectionEndpoint,
|
// Bump to force the bootstrap to re-run when the schema/logic changes.
|
||||||
IndexName: 'proposal-system-index',
|
Version: '1',
|
||||||
VectorField: 'embedding',
|
|
||||||
TextField: 'text',
|
|
||||||
MetadataField: 'metadata',
|
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
ossIndex.node.addDependency(ossDataAccessPolicy);
|
|
||||||
|
|
||||||
const knowledgeBase = new bedrock.CfnKnowledgeBase(this, 'KnowledgeBase', {
|
const knowledgeBase = new bedrock.CfnKnowledgeBase(this, 'KnowledgeBase', {
|
||||||
name: `proposal-system-kb${config.stackSuffix}`,
|
name: `proposal-system-kb${config.stackSuffix}`,
|
||||||
|
|
@ -219,19 +141,22 @@ export class ComputeStack extends cdk.Stack {
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
storageConfiguration: {
|
storageConfiguration: {
|
||||||
type: 'OPENSEARCH_SERVERLESS',
|
type: 'RDS',
|
||||||
opensearchServerlessConfiguration: {
|
rdsConfiguration: {
|
||||||
collectionArn: ossCollection.attrArn,
|
resourceArn: dbCluster.clusterArn,
|
||||||
vectorIndexName: 'proposal-system-index',
|
credentialsSecretArn: bedrockUserSecret.secretArn,
|
||||||
|
databaseName: 'proposals',
|
||||||
|
tableName: 'bedrock_integration.bedrock_kb',
|
||||||
fieldMapping: {
|
fieldMapping: {
|
||||||
|
primaryKeyField: 'id',
|
||||||
vectorField: 'embedding',
|
vectorField: 'embedding',
|
||||||
textField: 'text',
|
textField: 'chunks',
|
||||||
metadataField: 'metadata',
|
metadataField: 'metadata',
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
},
|
},
|
||||||
});
|
});
|
||||||
knowledgeBase.node.addDependency(ossIndex);
|
knowledgeBase.node.addDependency(pgvectorInit);
|
||||||
|
|
||||||
// KB Data Source (S3 library bucket)
|
// KB Data Source (S3 library bucket)
|
||||||
const dataSource = new bedrock.CfnDataSource(this, 'KbDataSource', {
|
const dataSource = new bedrock.CfnDataSource(this, 'KbDataSource', {
|
||||||
|
|
|
||||||
|
|
@ -21,6 +21,7 @@ export class FoundationStack extends cdk.Stack {
|
||||||
public readonly vpc: ec2.IVpc;
|
public readonly vpc: ec2.IVpc;
|
||||||
public readonly lambdaSecurityGroup: ec2.ISecurityGroup;
|
public readonly lambdaSecurityGroup: ec2.ISecurityGroup;
|
||||||
public readonly dbSecret: secretsmanager.ISecret;
|
public readonly dbSecret: secretsmanager.ISecret;
|
||||||
|
public readonly dbCluster: rds.IDatabaseCluster;
|
||||||
public readonly uploadsBucket: s3.IBucket;
|
public readonly uploadsBucket: s3.IBucket;
|
||||||
public readonly generatedBucket: s3.IBucket;
|
public readonly generatedBucket: s3.IBucket;
|
||||||
public readonly libraryBucket: s3.IBucket;
|
public readonly libraryBucket: s3.IBucket;
|
||||||
|
|
@ -83,34 +84,34 @@ export class FoundationStack extends cdk.Stack {
|
||||||
'Allow PostgreSQL from Lambda SG'
|
'Allow PostgreSQL from Lambda SG'
|
||||||
);
|
);
|
||||||
|
|
||||||
// RDS PostgreSQL 15
|
// Aurora PostgreSQL Serverless v2 — pgvector store for the Bedrock Knowledge Base.
|
||||||
const dbInstance = new rds.DatabaseInstance(this, 'Database', {
|
// PR3: replaced the RDS instance + OpenSearch Serverless with Aurora + pgvector
|
||||||
instanceIdentifier: `proposal-system-db${config.stackSuffix}`,
|
// (kills the AOSS OCU floor; scales toward 0 ACU when idle). Data API is required
|
||||||
engine: rds.DatabaseInstanceEngine.postgres({
|
// by Bedrock Knowledge Bases to query the vector table.
|
||||||
version: rds.PostgresEngineVersion.VER_15,
|
const dbCluster = new rds.DatabaseCluster(this, 'Database', {
|
||||||
|
clusterIdentifier: `proposal-system-db${config.stackSuffix}`,
|
||||||
|
engine: rds.DatabaseClusterEngine.auroraPostgres({
|
||||||
|
version: rds.AuroraPostgresEngineVersion.VER_15_4,
|
||||||
}),
|
}),
|
||||||
instanceType: ec2.InstanceType.of(
|
|
||||||
ec2.InstanceClass.T4G,
|
|
||||||
ec2.InstanceSize.SMALL
|
|
||||||
),
|
|
||||||
vpc: this.vpc,
|
vpc: this.vpc,
|
||||||
vpcSubnets: { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS },
|
vpcSubnets: { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS },
|
||||||
securityGroups: [rdsSg],
|
securityGroups: [rdsSg],
|
||||||
multiAz: false,
|
writer: rds.ClusterInstance.serverlessV2('writer'),
|
||||||
allocatedStorage: 20,
|
serverlessV2MinCapacity: 0.5,
|
||||||
maxAllocatedStorage: 100,
|
serverlessV2MaxCapacity: 4,
|
||||||
|
enableDataApi: true,
|
||||||
storageEncrypted: true,
|
storageEncrypted: true,
|
||||||
backupRetention: cdk.Duration.days(7),
|
backup: { retention: cdk.Duration.days(7) },
|
||||||
deletionProtection: config.retainData,
|
deletionProtection: config.retainData,
|
||||||
removalPolicy: config.retainData ? cdk.RemovalPolicy.RETAIN : cdk.RemovalPolicy.DESTROY,
|
removalPolicy: config.retainData ? cdk.RemovalPolicy.RETAIN : cdk.RemovalPolicy.DESTROY,
|
||||||
databaseName: 'proposals',
|
defaultDatabaseName: 'proposals',
|
||||||
credentials: rds.Credentials.fromGeneratedSecret('proposalsadmin', {
|
credentials: rds.Credentials.fromGeneratedSecret('proposalsadmin', {
|
||||||
secretName: `proposal-system/db-credentials${config.stackSuffix}`,
|
secretName: `proposal-system/db-credentials${config.stackSuffix}`,
|
||||||
}),
|
}),
|
||||||
publiclyAccessible: false,
|
|
||||||
});
|
});
|
||||||
|
|
||||||
this.dbSecret = dbInstance.secret!;
|
this.dbSecret = dbCluster.secret!;
|
||||||
|
this.dbCluster = dbCluster;
|
||||||
|
|
||||||
// S3 Buckets
|
// S3 Buckets
|
||||||
// Fix: INF-M5 — enforce HTTPS-only access on all S3 buckets
|
// Fix: INF-M5 — enforce HTTPS-only access on all S3 buckets
|
||||||
|
|
@ -300,7 +301,7 @@ export class FoundationStack extends cdk.Stack {
|
||||||
new cloudwatch.Alarm(this, 'RdsCpuAlarm', {
|
new cloudwatch.Alarm(this, 'RdsCpuAlarm', {
|
||||||
alarmName: `proposal-system-rds-cpu${config.stackSuffix}`,
|
alarmName: `proposal-system-rds-cpu${config.stackSuffix}`,
|
||||||
alarmDescription: 'RDS CPU utilization above 80%',
|
alarmDescription: 'RDS CPU utilization above 80%',
|
||||||
metric: dbInstance.metricCPUUtilization({ period: cdk.Duration.minutes(5) }),
|
metric: dbCluster.metricCPUUtilization({ period: cdk.Duration.minutes(5) }),
|
||||||
threshold: 80,
|
threshold: 80,
|
||||||
evaluationPeriods: 3,
|
evaluationPeriods: 3,
|
||||||
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
||||||
|
|
@ -308,19 +309,21 @@ export class FoundationStack extends cdk.Stack {
|
||||||
new cloudwatch.Alarm(this, 'RdsConnectionsAlarm', {
|
new cloudwatch.Alarm(this, 'RdsConnectionsAlarm', {
|
||||||
alarmName: `proposal-system-rds-connections${config.stackSuffix}`,
|
alarmName: `proposal-system-rds-connections${config.stackSuffix}`,
|
||||||
alarmDescription: 'RDS database connections above 80',
|
alarmDescription: 'RDS database connections above 80',
|
||||||
metric: dbInstance.metricDatabaseConnections({ period: cdk.Duration.minutes(5) }),
|
metric: dbCluster.metricDatabaseConnections({ period: cdk.Duration.minutes(5) }),
|
||||||
threshold: 80,
|
threshold: 80,
|
||||||
evaluationPeriods: 2,
|
evaluationPeriods: 2,
|
||||||
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
|
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
|
||||||
}),
|
}),
|
||||||
new cloudwatch.Alarm(this, 'RdsFreeStorageAlarm', {
|
// Aurora storage auto-scales (no FreeStorageSpace); freeable memory is the
|
||||||
alarmName: `proposal-system-rds-free-storage${config.stackSuffix}`,
|
// meaningful health signal for a Serverless v2 cluster.
|
||||||
alarmDescription: 'RDS free storage below 2 GB',
|
new cloudwatch.Alarm(this, 'RdsLowMemoryAlarm', {
|
||||||
metric: dbInstance.metricFreeStorageSpace({ period: cdk.Duration.minutes(5) }),
|
alarmName: `proposal-system-rds-low-memory${config.stackSuffix}`,
|
||||||
threshold: 2_000_000_000,
|
alarmDescription: 'Aurora freeable memory below 256 MB',
|
||||||
|
metric: dbCluster.metricFreeableMemory({ period: cdk.Duration.minutes(5) }),
|
||||||
|
threshold: 256_000_000,
|
||||||
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
|
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
|
||||||
evaluationPeriods: 1,
|
evaluationPeriods: 3,
|
||||||
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
|
||||||
}),
|
}),
|
||||||
];
|
];
|
||||||
for (const alarm of rdsAlarms) {
|
for (const alarm of rdsAlarms) {
|
||||||
|
|
|
||||||
126
lambdas/aurora-pgvector-init/app.py
Normal file
126
lambdas/aurora-pgvector-init/app.py
Normal file
|
|
@ -0,0 +1,126 @@
|
||||||
|
"""CloudFormation custom-resource Lambda: initialize the Aurora pgvector store
|
||||||
|
for the Bedrock Knowledge Base (PR3 — replaces the OpenSearch index creator).
|
||||||
|
|
||||||
|
Runs the one-time bootstrap SQL via the RDS Data API as the master user:
|
||||||
|
enables pgvector, creates the `bedrock_integration` schema + `bedrock_kb` table
|
||||||
|
(with the column/index layout Bedrock requires), and creates/owns the dedicated
|
||||||
|
`bedrock_user` role whose credentials Bedrock uses to query the table.
|
||||||
|
|
||||||
|
Idempotent: safe to run on stack create/update (IF NOT EXISTS throughout). Only
|
||||||
|
boto3 (rds-data, secretsmanager) is used — both ship in the Lambda runtime.
|
||||||
|
"""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import time
|
||||||
|
|
||||||
|
import boto3
|
||||||
|
from botocore.exceptions import ClientError
|
||||||
|
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
logger.setLevel(os.environ.get("LOG_LEVEL", "INFO"))
|
||||||
|
|
||||||
|
rds_data = boto3.client("rds-data")
|
||||||
|
secrets = boto3.client("secretsmanager")
|
||||||
|
|
||||||
|
CLUSTER_ARN = os.environ["CLUSTER_ARN"]
|
||||||
|
MASTER_SECRET_ARN = os.environ["MASTER_SECRET_ARN"]
|
||||||
|
BEDROCK_SECRET_ARN = os.environ["BEDROCK_SECRET_ARN"]
|
||||||
|
DATABASE = os.environ.get("DATABASE", "proposals")
|
||||||
|
EMBED_DIM = int(os.environ.get("EMBED_DIM", "1024")) # Titan Embed v2
|
||||||
|
|
||||||
|
PHYSICAL_ID = "aurora-pgvector-init"
|
||||||
|
TABLE = "bedrock_integration.bedrock_kb"
|
||||||
|
|
||||||
|
|
||||||
|
def _exec(sql: str, attempts: int = 6) -> None:
|
||||||
|
"""Run one statement via the RDS Data API as the master user.
|
||||||
|
|
||||||
|
Retries transient errors while a Serverless v2 cluster is still becoming
|
||||||
|
reachable right after deploy (Data API can briefly report the cluster as
|
||||||
|
unavailable / resuming).
|
||||||
|
"""
|
||||||
|
for attempt in range(attempts):
|
||||||
|
try:
|
||||||
|
rds_data.execute_statement(
|
||||||
|
resourceArn=CLUSTER_ARN,
|
||||||
|
secretArn=MASTER_SECRET_ARN,
|
||||||
|
database=DATABASE,
|
||||||
|
sql=sql,
|
||||||
|
)
|
||||||
|
return
|
||||||
|
except ClientError as exc:
|
||||||
|
code = exc.response.get("Error", {}).get("Code", "")
|
||||||
|
message = str(exc)
|
||||||
|
transient = code == "DatabaseResumingException" or any(
|
||||||
|
s in message
|
||||||
|
for s in ("not currently available", "Communication link failure")
|
||||||
|
)
|
||||||
|
if transient and attempt < attempts - 1:
|
||||||
|
logger.warning(
|
||||||
|
"Transient Data API error (attempt %d/%d): %s",
|
||||||
|
attempt + 1,
|
||||||
|
attempts,
|
||||||
|
code or message,
|
||||||
|
)
|
||||||
|
time.sleep(10)
|
||||||
|
continue
|
||||||
|
raise
|
||||||
|
|
||||||
|
|
||||||
|
def handler(event, context):
|
||||||
|
request_type = event.get("RequestType")
|
||||||
|
logger.info("RequestType=%s", request_type)
|
||||||
|
|
||||||
|
# Leave the data in place on stack delete — nothing to undo.
|
||||||
|
if request_type == "Delete":
|
||||||
|
return {"PhysicalResourceId": PHYSICAL_ID}
|
||||||
|
|
||||||
|
# The password Bedrock will use to log in as bedrock_user. Generated by
|
||||||
|
# Secrets Manager with SQL-unsafe characters excluded (see CDK).
|
||||||
|
secret = json.loads(
|
||||||
|
secrets.get_secret_value(SecretId=BEDROCK_SECRET_ARN)["SecretString"]
|
||||||
|
)
|
||||||
|
password = secret["password"]
|
||||||
|
# Defense-in-depth: the password is inlined into CREATE/ALTER ROLE SQL (DDL can't
|
||||||
|
# bind parameters). The secret is generated with excludePunctuation=true, so it must
|
||||||
|
# be strictly alphanumeric — refuse anything else rather than risk SQL breakage.
|
||||||
|
if not password.isalnum():
|
||||||
|
raise ValueError(
|
||||||
|
"bedrock_user password is not alphanumeric; refusing to inline"
|
||||||
|
)
|
||||||
|
|
||||||
|
statements = [
|
||||||
|
"CREATE EXTENSION IF NOT EXISTS vector;",
|
||||||
|
"CREATE SCHEMA IF NOT EXISTS bedrock_integration;",
|
||||||
|
# Create the role if missing, then (re)set its password to match the secret.
|
||||||
|
"DO $$ BEGIN "
|
||||||
|
"IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'bedrock_user') "
|
||||||
|
f"THEN CREATE ROLE bedrock_user LOGIN PASSWORD '{password}'; END IF; END $$;",
|
||||||
|
f"ALTER ROLE bedrock_user WITH LOGIN PASSWORD '{password}';",
|
||||||
|
"GRANT ALL ON SCHEMA bedrock_integration TO bedrock_user;",
|
||||||
|
f"CREATE TABLE IF NOT EXISTS {TABLE} ("
|
||||||
|
"id uuid PRIMARY KEY, "
|
||||||
|
f"embedding vector({EMBED_DIM}), "
|
||||||
|
"chunks text, "
|
||||||
|
"metadata json, "
|
||||||
|
"custom_metadata jsonb);",
|
||||||
|
f"ALTER TABLE {TABLE} OWNER TO bedrock_user;",
|
||||||
|
f"CREATE INDEX IF NOT EXISTS bedrock_kb_embedding_idx ON {TABLE} "
|
||||||
|
"USING hnsw (embedding vector_cosine_ops) WITH (ef_construction = 256);",
|
||||||
|
f"CREATE INDEX IF NOT EXISTS bedrock_kb_chunks_idx ON {TABLE} "
|
||||||
|
"USING gin (to_tsvector('simple', chunks));",
|
||||||
|
f"CREATE INDEX IF NOT EXISTS bedrock_kb_custom_metadata_idx ON {TABLE} "
|
||||||
|
"USING gin (custom_metadata);",
|
||||||
|
"GRANT ALL ON ALL TABLES IN SCHEMA bedrock_integration TO bedrock_user;",
|
||||||
|
]
|
||||||
|
|
||||||
|
# Log only the position — the statements contain the bedrock_user password
|
||||||
|
# (CREATE/ALTER ROLE), so the SQL text itself must never reach the logs.
|
||||||
|
for i, sql in enumerate(statements, 1):
|
||||||
|
logger.info("executing bootstrap statement %d/%d", i, len(statements))
|
||||||
|
_exec(sql)
|
||||||
|
|
||||||
|
logger.info("pgvector store initialized: %s (dim=%d)", TABLE, EMBED_DIM)
|
||||||
|
return {"PhysicalResourceId": PHYSICAL_ID, "Data": {"TableName": TABLE}}
|
||||||
1
lambdas/aurora-pgvector-init/requirements.txt
Normal file
1
lambdas/aurora-pgvector-init/requirements.txt
Normal file
|
|
@ -0,0 +1 @@
|
||||||
|
boto3>=1.43.18,<2.0
|
||||||
|
|
@ -1,71 +0,0 @@
|
||||||
import time
|
|
||||||
import boto3
|
|
||||||
from opensearchpy import OpenSearch, RequestsHttpConnection
|
|
||||||
from requests_aws4auth import AWS4Auth
|
|
||||||
|
|
||||||
|
|
||||||
def handler(event, context):
|
|
||||||
if event["RequestType"] == "Delete":
|
|
||||||
return {"PhysicalResourceId": event.get("PhysicalResourceId", "none")}
|
|
||||||
|
|
||||||
props = event["ResourceProperties"]
|
|
||||||
endpoint = props["Endpoint"].replace("https://", "")
|
|
||||||
index_name = props["IndexName"]
|
|
||||||
vector_field = props["VectorField"]
|
|
||||||
text_field = props["TextField"]
|
|
||||||
metadata_field = props["MetadataField"]
|
|
||||||
|
|
||||||
session = boto3.Session()
|
|
||||||
credentials = session.get_credentials().get_frozen_credentials()
|
|
||||||
region = session.region_name
|
|
||||||
|
|
||||||
awsauth = AWS4Auth(
|
|
||||||
credentials.access_key,
|
|
||||||
credentials.secret_key,
|
|
||||||
region,
|
|
||||||
"aoss",
|
|
||||||
session_token=credentials.token,
|
|
||||||
)
|
|
||||||
|
|
||||||
client = OpenSearch(
|
|
||||||
hosts=[{"host": endpoint, "port": 443}],
|
|
||||||
http_auth=awsauth,
|
|
||||||
use_ssl=True,
|
|
||||||
verify_certs=True,
|
|
||||||
connection_class=RequestsHttpConnection,
|
|
||||||
timeout=30,
|
|
||||||
)
|
|
||||||
|
|
||||||
index_body = {
|
|
||||||
"settings": {"index": {"knn": True, "knn.algo_param.ef_search": 512}},
|
|
||||||
"mappings": {
|
|
||||||
"properties": {
|
|
||||||
vector_field: {
|
|
||||||
"type": "knn_vector",
|
|
||||||
"dimension": 1024,
|
|
||||||
"method": {
|
|
||||||
"engine": "faiss",
|
|
||||||
"name": "hnsw",
|
|
||||||
"space_type": "l2",
|
|
||||||
},
|
|
||||||
},
|
|
||||||
text_field: {"type": "text"},
|
|
||||||
metadata_field: {"type": "text"},
|
|
||||||
}
|
|
||||||
},
|
|
||||||
}
|
|
||||||
|
|
||||||
for attempt in range(30):
|
|
||||||
try:
|
|
||||||
client.indices.create(index=index_name, body=index_body)
|
|
||||||
return {"PhysicalResourceId": index_name}
|
|
||||||
except Exception as e:
|
|
||||||
error_str = str(e)
|
|
||||||
if "resource_already_exists_exception" in error_str:
|
|
||||||
return {"PhysicalResourceId": index_name}
|
|
||||||
if "403" in error_str and attempt < 29:
|
|
||||||
time.sleep(10)
|
|
||||||
continue
|
|
||||||
raise
|
|
||||||
|
|
||||||
raise Exception("Timeout waiting for AOSS access policy propagation")
|
|
||||||
|
|
@ -1,3 +0,0 @@
|
||||||
opensearch-py>=3.2.0
|
|
||||||
requests-aws4auth>=1.3.2
|
|
||||||
requests>=2.34.2
|
|
||||||
Loading…
Add table
Reference in a new issue