mirror of
https://github.com/Sea-Haven-Industries/proposal-system.git
synced 2026-09-30 03:03:13 +00:00
feat(infra): migrate Bedrock KB vector store to Aurora pgvector (#125)
Some checks failed
Deploy / Deploy to AWS (push) Has been cancelled
Some checks failed
Deploy / Deploy to AWS (push) Has been cancelled
Replaces OpenSearch Serverless with Aurora PostgreSQL Serverless v2 + pgvector as the Bedrock Knowledge Base vector store (v1 PR3). Bedrock KB requires Aurora SSv2 (RDS Data API), not a plain RDS instance — so the DB engine moves to Aurora. - foundation: rds.DatabaseInstance (PG15) -> rds.DatabaseCluster Aurora SSv2 (0.5-4 ACU, enableDataApi). RDS alarms: free-storage -> freeable-memory. - compute: delete all AOSS (collection, policies, VPC endpoint, index-creator); add bedrock_user secret + KB role (scoped rds-data + secret read); repoint CfnKnowledgeBase to RDS storage (bedrock_integration.bedrock_kb, vector(1024)). - lambdas: oss-index-creator -> aurora-pgvector-init (bootstrap schema/table/ indexes/role via RDS Data API; transient-error retry; password guard). - ADR 0001 documents the decision. Eliminates the ~$175-350/mo AOSS OCU floor. NAT kept (egress still needed). GPT-4.1 cross-review: no BLOCK (FIX applied). tsc clean; foundation synth shows Aurora cluster with Data API enabled; 23 pytest pass.
This commit is contained in:
parent
bbd185b280
commit
aff38a11ae
8 changed files with 283 additions and 233 deletions
68
docs/adr/0001-bedrock-vector-store-aurora-pgvector.md
Normal file
68
docs/adr/0001-bedrock-vector-store-aurora-pgvector.md
Normal file
|
|
@ -0,0 +1,68 @@
|
|||
# ADR 0001 — Bedrock Knowledge Base vector store: Aurora PostgreSQL + pgvector
|
||||
|
||||
- **Status:** Accepted (2026-06-12)
|
||||
- **Decision owner:** Adam Moussa
|
||||
- **Scope:** `proposal-system` infrastructure (foundation + compute stacks), v1 PR3
|
||||
|
||||
## Context
|
||||
|
||||
The proposal system's RAG pipeline uses an Amazon Bedrock Knowledge Base (Titan Embed
|
||||
v2) over historical proposals. The original vector store was **OpenSearch Serverless
|
||||
(AOSS)**, which:
|
||||
|
||||
- carries a minimum ~2-OCU billing floor (~$175–350/mo) even at near-zero query volume —
|
||||
the dominant line item of the monthly bill for an internal tool;
|
||||
- is VPC-only, which forces the NAT gateway and a `oss-index-creator` bootstrap Lambda
|
||||
that exists purely to pre-create the vector index while an IAM access policy propagates.
|
||||
|
||||
The database is a small RDS PostgreSQL instance. The v1 assessment flagged AOSS as
|
||||
over-built for the corpus size and recommended a pgvector store on the existing database.
|
||||
|
||||
## Decision
|
||||
|
||||
Migrate the database from **RDS PostgreSQL 15 → Aurora PostgreSQL Serverless v2** and use
|
||||
**pgvector** as the Bedrock KB vector store.
|
||||
|
||||
Aurora is required because **Bedrock Knowledge Bases support Aurora PostgreSQL
|
||||
Serverless v2 (with the RDS Data API) as a pgvector store, but not a plain RDS
|
||||
instance.** A standard RDS instance cannot back a Bedrock KB, so "pgvector on the
|
||||
existing RDS" was infeasible without the Aurora move.
|
||||
|
||||
Implementation:
|
||||
|
||||
- Aurora Serverless v2 (min 0.5 / max 4 ACU), `enableDataApi: true`, `defaultDatabaseName:
|
||||
'proposals'`. The cluster also serves the .NET API's application data (one database).
|
||||
- A bootstrap custom-resource Lambda (`aurora-pgvector-init`, via the RDS Data API)
|
||||
enables `vector`, creates the `bedrock_integration.bedrock_kb` table
|
||||
(`vector(1024)` for Titan v2 + HNSW cosine + GIN indexes) and a dedicated
|
||||
`bedrock_user` role. This **replaces** `oss-index-creator` — the bootstrap is swapped,
|
||||
not eliminated.
|
||||
- `CfnKnowledgeBase.storageConfiguration` → `type: 'RDS'` with `rdsConfiguration`
|
||||
(cluster ARN, `bedrock_user` secret, `bedrock_integration.bedrock_kb`, field mapping
|
||||
`id`/`embedding`/`chunks`/`metadata`). KB role IAM swaps `aoss:APIAccessAll` →
|
||||
scoped `rds-data` + secret read.
|
||||
- All AOSS constructs (collection, policies, VPC endpoint, index-creator) are deleted.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Cost:** eliminates the AOSS OCU floor (~$175–350/mo). Aurora Serverless v2 at
|
||||
0.5 ACU min is ~$43/mo and scales toward zero idle — a net reduction.
|
||||
- **Simplification:** one data engine (Aurora) instead of RDS + AOSS; fewer constructs.
|
||||
A bootstrap Lambda remains (now for pgvector schema rather than the AOSS index).
|
||||
- **NAT:** kept for now — Lambdas and the KB's Data API path still need AWS-service
|
||||
egress. Dropping NAT would require VPC interface endpoints; tracked separately.
|
||||
- **Migration:** the RDS→Aurora swap is a CloudFormation replacement. The deployed stacks
|
||||
are **test-only with no production data**, so this is a clean redeploy.
|
||||
|
||||
## Alternatives considered
|
||||
|
||||
- **Keep AOSS:** rejected — the cost floor is the single biggest waste for the scale.
|
||||
- **S3 Vectors:** viable Bedrock backend, but Aurora unifies app data + vectors and was
|
||||
the owner's preference (also cheaper than AOSS).
|
||||
- **pgvector on the existing RDS instance:** infeasible — Bedrock KB does not support a
|
||||
plain RDS instance as a vector store.
|
||||
|
||||
## References
|
||||
|
||||
- [Using Aurora PostgreSQL as a Bedrock Knowledge Base](https://docs.aws.amazon.com/AmazonRDS/latest/AuroraUserGuide/AuroraPostgreSQL.VectorDB.html)
|
||||
- [Bedrock KB vector-store prerequisites](https://docs.aws.amazon.com/bedrock/latest/userguide/knowledge-base-setup.html)
|
||||
|
|
@ -26,6 +26,7 @@ const compute = new ComputeStack(app, `proposal-system-compute${stackSuffix}`, {
|
|||
vpc: foundation.vpc,
|
||||
lambdaSecurityGroup: foundation.lambdaSecurityGroup,
|
||||
dbSecret: foundation.dbSecret,
|
||||
dbCluster: foundation.dbCluster,
|
||||
uploadsBucket: foundation.uploadsBucket,
|
||||
generatedBucket: foundation.generatedBucket,
|
||||
libraryBucket: foundation.libraryBucket,
|
||||
|
|
|
|||
|
|
@ -12,7 +12,7 @@ import * as cognito from 'aws-cdk-lib/aws-cognito';
|
|||
import * as secretsmanager from 'aws-cdk-lib/aws-secretsmanager';
|
||||
import * as lambdaEventSources from 'aws-cdk-lib/aws-lambda-event-sources';
|
||||
import * as bedrock from 'aws-cdk-lib/aws-bedrock';
|
||||
import * as opensearchserverless from 'aws-cdk-lib/aws-opensearchserverless';
|
||||
import * as rds from 'aws-cdk-lib/aws-rds';
|
||||
import * as logs from 'aws-cdk-lib/aws-logs';
|
||||
import * as cloudwatch from 'aws-cdk-lib/aws-cloudwatch';
|
||||
import * as cloudwatchActions from 'aws-cdk-lib/aws-cloudwatch-actions';
|
||||
|
|
@ -25,6 +25,7 @@ export interface ComputeStackProps extends cdk.StackProps {
|
|||
vpc: ec2.IVpc;
|
||||
lambdaSecurityGroup: ec2.ISecurityGroup;
|
||||
dbSecret: secretsmanager.ISecret;
|
||||
dbCluster: rds.IDatabaseCluster;
|
||||
uploadsBucket: s3.IBucket;
|
||||
generatedBucket: s3.IBucket;
|
||||
libraryBucket: s3.IBucket;
|
||||
|
|
@ -51,45 +52,23 @@ export class ComputeStack extends cdk.Stack {
|
|||
},
|
||||
});
|
||||
|
||||
// OpenSearch Serverless collection for Bedrock KB vector store
|
||||
const ossEncryptionPolicy = new opensearchserverless.CfnSecurityPolicy(this, 'OssEncryptionPolicy', {
|
||||
name: `proposal-system-kb-enc${config.stackSuffix}`,
|
||||
type: 'encryption',
|
||||
policy: JSON.stringify({
|
||||
Rules: [{ ResourceType: 'collection', Resource: [`collection/proposal-system-kb${config.stackSuffix}`] }],
|
||||
AWSOwnedKey: true,
|
||||
}),
|
||||
});
|
||||
// Bedrock Knowledge Base vector store: Aurora PostgreSQL + pgvector (PR3 —
|
||||
// replaced OpenSearch Serverless). A dedicated bedrock_user role queries the
|
||||
// pgvector table; the aurora-pgvector-init custom resource creates the schema.
|
||||
const dbCluster = props.dbCluster;
|
||||
|
||||
// Fix: INF-H3 — restrict OpenSearch Serverless to VPC (was AllowFromPublic: true).
|
||||
// Create a VPC endpoint so Lambdas in private subnets can reach the collection.
|
||||
const ossVpcEndpoint = new opensearchserverless.CfnVpcEndpoint(this, 'OssVpcEndpoint', {
|
||||
name: `proposal-system-kb-vpce${config.stackSuffix}`,
|
||||
vpcId: props.vpc.vpcId,
|
||||
subnetIds: props.vpc.selectSubnets({ subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS }).subnetIds,
|
||||
securityGroupIds: [props.lambdaSecurityGroup.securityGroupId],
|
||||
// Credentials Bedrock uses to query the pgvector table as bedrock_user.
|
||||
// SQL-unsafe characters are excluded so the bootstrap can inline the password.
|
||||
const bedrockUserSecret = new secretsmanager.Secret(this, 'BedrockUserSecret', {
|
||||
secretName: `proposal-system/bedrock-user${config.stackSuffix}`,
|
||||
generateSecretString: {
|
||||
secretStringTemplate: JSON.stringify({ username: 'bedrock_user' }),
|
||||
generateStringKey: 'password',
|
||||
excludePunctuation: true,
|
||||
passwordLength: 32,
|
||||
},
|
||||
});
|
||||
|
||||
const ossNetworkPolicy = new opensearchserverless.CfnSecurityPolicy(this, 'OssNetworkPolicy', {
|
||||
name: `proposal-system-kb-net${config.stackSuffix}`,
|
||||
type: 'network',
|
||||
policy: JSON.stringify([{
|
||||
Rules: [
|
||||
{ ResourceType: 'collection', Resource: [`collection/proposal-system-kb${config.stackSuffix}`] },
|
||||
],
|
||||
AllowFromPublic: false,
|
||||
SourceVPCEs: [ossVpcEndpoint.attrId],
|
||||
}]),
|
||||
});
|
||||
ossNetworkPolicy.addDependency(ossVpcEndpoint);
|
||||
|
||||
const ossCollection = new opensearchserverless.CfnCollection(this, 'OssCollection', {
|
||||
name: `proposal-system-kb${config.stackSuffix}`,
|
||||
type: 'VECTORSEARCH',
|
||||
});
|
||||
ossCollection.addDependency(ossEncryptionPolicy);
|
||||
ossCollection.addDependency(ossNetworkPolicy);
|
||||
|
||||
// Bedrock KB execution role
|
||||
const kbRole = new iam.Role(this, 'KnowledgeBaseRole', {
|
||||
roleName: `proposal-system-kb-role${config.stackSuffix}`,
|
||||
|
|
@ -101,113 +80,56 @@ export class ComputeStack extends cdk.Stack {
|
|||
resources: [props.libraryBucket.bucketArn, `${props.libraryBucket.bucketArn}/*`],
|
||||
}));
|
||||
|
||||
kbRole.addToPolicy(new iam.PolicyStatement({
|
||||
actions: ['aoss:APIAccessAll'],
|
||||
resources: [ossCollection.attrArn],
|
||||
}));
|
||||
|
||||
kbRole.addToPolicy(new iam.PolicyStatement({
|
||||
actions: ['bedrock:InvokeModel'],
|
||||
resources: [`arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0`],
|
||||
}));
|
||||
|
||||
// Lambda to pre-create the vector index (retries until AOSS access policy propagates)
|
||||
const indexCreatorFn = new lambda.Function(this, 'OssIndexCreator', {
|
||||
functionName: `proposal-system-oss-index-creator${config.stackSuffix}`,
|
||||
// KB queries the pgvector table via the RDS Data API as bedrock_user.
|
||||
kbRole.addToPolicy(new iam.PolicyStatement({
|
||||
actions: [
|
||||
'rds-data:ExecuteStatement',
|
||||
'rds-data:BatchExecuteStatement',
|
||||
'rds-data:BeginTransaction',
|
||||
'rds-data:CommitTransaction',
|
||||
'rds-data:RollbackTransaction',
|
||||
],
|
||||
resources: [dbCluster.clusterArn],
|
||||
}));
|
||||
bedrockUserSecret.grantRead(kbRole);
|
||||
|
||||
// One-time bootstrap: enable pgvector + create the bedrock_integration schema,
|
||||
// table, indexes, and bedrock_user role (via the RDS Data API as master).
|
||||
const pgvectorInitFn = new lambda.Function(this, 'PgVectorInit', {
|
||||
functionName: `proposal-system-pgvector-init${config.stackSuffix}`,
|
||||
runtime: lambda.Runtime.PYTHON_3_12,
|
||||
architecture: lambda.Architecture.ARM_64,
|
||||
handler: 'app.handler',
|
||||
code: lambda.Code.fromAsset('../lambdas/oss-index-creator', {
|
||||
bundling: {
|
||||
image: lambda.Runtime.PYTHON_3_12.bundlingImage,
|
||||
command: [
|
||||
'bash', '-c',
|
||||
'pip install -r requirements.txt -t /asset-output && cp -au . /asset-output',
|
||||
],
|
||||
},
|
||||
}),
|
||||
timeout: cdk.Duration.minutes(6),
|
||||
code: lambda.Code.fromAsset('../lambdas/aurora-pgvector-init'),
|
||||
timeout: cdk.Duration.minutes(5),
|
||||
environment: {
|
||||
CLUSTER_ARN: dbCluster.clusterArn,
|
||||
MASTER_SECRET_ARN: props.dbSecret.secretArn,
|
||||
BEDROCK_SECRET_ARN: bedrockUserSecret.secretArn,
|
||||
DATABASE: 'proposals',
|
||||
EMBED_DIM: '1024',
|
||||
},
|
||||
logRetention: logs.RetentionDays.TWO_MONTHS,
|
||||
});
|
||||
dbCluster.grantDataApiAccess(pgvectorInitFn);
|
||||
bedrockUserSecret.grantRead(pgvectorInitFn);
|
||||
|
||||
indexCreatorFn.addToRolePolicy(new iam.PolicyStatement({
|
||||
actions: ['aoss:APIAccessAll'],
|
||||
resources: [ossCollection.attrArn],
|
||||
}));
|
||||
|
||||
// Fix: INF-M2 — scope AOSS data access policy permissions (was aoss:* on both
|
||||
// collection and index). KB role needs read/write for embeddings. Index creator
|
||||
// needs create/describe for bootstrapping the vector index.
|
||||
const ossDataAccessPolicy = new opensearchserverless.CfnAccessPolicy(this, 'OssDataAccessPolicy', {
|
||||
name: `proposal-system-kb-access${config.stackSuffix}`,
|
||||
type: 'data',
|
||||
policy: JSON.stringify([
|
||||
{
|
||||
Description: 'Bedrock KB role — read/write documents and describe collection',
|
||||
Rules: [
|
||||
{
|
||||
ResourceType: 'collection',
|
||||
Resource: [`collection/proposal-system-kb${config.stackSuffix}`],
|
||||
Permission: [
|
||||
'aoss:DescribeCollectionItems',
|
||||
'aoss:CreateCollectionItems',
|
||||
'aoss:UpdateCollectionItems',
|
||||
],
|
||||
},
|
||||
{
|
||||
ResourceType: 'index',
|
||||
Resource: [`index/proposal-system-kb${config.stackSuffix}/*`],
|
||||
Permission: [
|
||||
'aoss:DescribeIndex',
|
||||
'aoss:ReadDocument',
|
||||
'aoss:WriteDocument',
|
||||
],
|
||||
},
|
||||
],
|
||||
Principal: [kbRole.roleArn],
|
||||
},
|
||||
{
|
||||
Description: 'Index creator Lambda — create and describe index during bootstrap',
|
||||
Rules: [
|
||||
{
|
||||
ResourceType: 'collection',
|
||||
Resource: [`collection/proposal-system-kb${config.stackSuffix}`],
|
||||
Permission: [
|
||||
'aoss:DescribeCollectionItems',
|
||||
'aoss:CreateCollectionItems',
|
||||
],
|
||||
},
|
||||
{
|
||||
ResourceType: 'index',
|
||||
Resource: [`index/proposal-system-kb${config.stackSuffix}/*`],
|
||||
Permission: [
|
||||
'aoss:CreateIndex',
|
||||
'aoss:DescribeIndex',
|
||||
'aoss:WriteDocument',
|
||||
],
|
||||
},
|
||||
],
|
||||
Principal: [indexCreatorFn.role!.roleArn],
|
||||
},
|
||||
]),
|
||||
});
|
||||
ossDataAccessPolicy.addDependency(ossCollection);
|
||||
|
||||
const indexProvider = new cr.Provider(this, 'OssIndexProvider', {
|
||||
onEventHandler: indexCreatorFn,
|
||||
const pgvectorProvider = new cr.Provider(this, 'PgVectorInitProvider', {
|
||||
onEventHandler: pgvectorInitFn,
|
||||
});
|
||||
|
||||
const ossIndex = new cdk.CustomResource(this, 'OssIndex', {
|
||||
serviceToken: indexProvider.serviceToken,
|
||||
const pgvectorInit = new cdk.CustomResource(this, 'PgVectorInitResource', {
|
||||
serviceToken: pgvectorProvider.serviceToken,
|
||||
properties: {
|
||||
Endpoint: ossCollection.attrCollectionEndpoint,
|
||||
IndexName: 'proposal-system-index',
|
||||
VectorField: 'embedding',
|
||||
TextField: 'text',
|
||||
MetadataField: 'metadata',
|
||||
// Bump to force the bootstrap to re-run when the schema/logic changes.
|
||||
Version: '1',
|
||||
},
|
||||
});
|
||||
ossIndex.node.addDependency(ossDataAccessPolicy);
|
||||
|
||||
const knowledgeBase = new bedrock.CfnKnowledgeBase(this, 'KnowledgeBase', {
|
||||
name: `proposal-system-kb${config.stackSuffix}`,
|
||||
|
|
@ -219,19 +141,22 @@ export class ComputeStack extends cdk.Stack {
|
|||
},
|
||||
},
|
||||
storageConfiguration: {
|
||||
type: 'OPENSEARCH_SERVERLESS',
|
||||
opensearchServerlessConfiguration: {
|
||||
collectionArn: ossCollection.attrArn,
|
||||
vectorIndexName: 'proposal-system-index',
|
||||
type: 'RDS',
|
||||
rdsConfiguration: {
|
||||
resourceArn: dbCluster.clusterArn,
|
||||
credentialsSecretArn: bedrockUserSecret.secretArn,
|
||||
databaseName: 'proposals',
|
||||
tableName: 'bedrock_integration.bedrock_kb',
|
||||
fieldMapping: {
|
||||
primaryKeyField: 'id',
|
||||
vectorField: 'embedding',
|
||||
textField: 'text',
|
||||
textField: 'chunks',
|
||||
metadataField: 'metadata',
|
||||
},
|
||||
},
|
||||
},
|
||||
});
|
||||
knowledgeBase.node.addDependency(ossIndex);
|
||||
knowledgeBase.node.addDependency(pgvectorInit);
|
||||
|
||||
// KB Data Source (S3 library bucket)
|
||||
const dataSource = new bedrock.CfnDataSource(this, 'KbDataSource', {
|
||||
|
|
|
|||
|
|
@ -21,6 +21,7 @@ export class FoundationStack extends cdk.Stack {
|
|||
public readonly vpc: ec2.IVpc;
|
||||
public readonly lambdaSecurityGroup: ec2.ISecurityGroup;
|
||||
public readonly dbSecret: secretsmanager.ISecret;
|
||||
public readonly dbCluster: rds.IDatabaseCluster;
|
||||
public readonly uploadsBucket: s3.IBucket;
|
||||
public readonly generatedBucket: s3.IBucket;
|
||||
public readonly libraryBucket: s3.IBucket;
|
||||
|
|
@ -83,34 +84,34 @@ export class FoundationStack extends cdk.Stack {
|
|||
'Allow PostgreSQL from Lambda SG'
|
||||
);
|
||||
|
||||
// RDS PostgreSQL 15
|
||||
const dbInstance = new rds.DatabaseInstance(this, 'Database', {
|
||||
instanceIdentifier: `proposal-system-db${config.stackSuffix}`,
|
||||
engine: rds.DatabaseInstanceEngine.postgres({
|
||||
version: rds.PostgresEngineVersion.VER_15,
|
||||
// Aurora PostgreSQL Serverless v2 — pgvector store for the Bedrock Knowledge Base.
|
||||
// PR3: replaced the RDS instance + OpenSearch Serverless with Aurora + pgvector
|
||||
// (kills the AOSS OCU floor; scales toward 0 ACU when idle). Data API is required
|
||||
// by Bedrock Knowledge Bases to query the vector table.
|
||||
const dbCluster = new rds.DatabaseCluster(this, 'Database', {
|
||||
clusterIdentifier: `proposal-system-db${config.stackSuffix}`,
|
||||
engine: rds.DatabaseClusterEngine.auroraPostgres({
|
||||
version: rds.AuroraPostgresEngineVersion.VER_15_4,
|
||||
}),
|
||||
instanceType: ec2.InstanceType.of(
|
||||
ec2.InstanceClass.T4G,
|
||||
ec2.InstanceSize.SMALL
|
||||
),
|
||||
vpc: this.vpc,
|
||||
vpcSubnets: { subnetType: ec2.SubnetType.PRIVATE_WITH_EGRESS },
|
||||
securityGroups: [rdsSg],
|
||||
multiAz: false,
|
||||
allocatedStorage: 20,
|
||||
maxAllocatedStorage: 100,
|
||||
writer: rds.ClusterInstance.serverlessV2('writer'),
|
||||
serverlessV2MinCapacity: 0.5,
|
||||
serverlessV2MaxCapacity: 4,
|
||||
enableDataApi: true,
|
||||
storageEncrypted: true,
|
||||
backupRetention: cdk.Duration.days(7),
|
||||
backup: { retention: cdk.Duration.days(7) },
|
||||
deletionProtection: config.retainData,
|
||||
removalPolicy: config.retainData ? cdk.RemovalPolicy.RETAIN : cdk.RemovalPolicy.DESTROY,
|
||||
databaseName: 'proposals',
|
||||
defaultDatabaseName: 'proposals',
|
||||
credentials: rds.Credentials.fromGeneratedSecret('proposalsadmin', {
|
||||
secretName: `proposal-system/db-credentials${config.stackSuffix}`,
|
||||
}),
|
||||
publiclyAccessible: false,
|
||||
});
|
||||
|
||||
this.dbSecret = dbInstance.secret!;
|
||||
this.dbSecret = dbCluster.secret!;
|
||||
this.dbCluster = dbCluster;
|
||||
|
||||
// S3 Buckets
|
||||
// Fix: INF-M5 — enforce HTTPS-only access on all S3 buckets
|
||||
|
|
@ -300,7 +301,7 @@ export class FoundationStack extends cdk.Stack {
|
|||
new cloudwatch.Alarm(this, 'RdsCpuAlarm', {
|
||||
alarmName: `proposal-system-rds-cpu${config.stackSuffix}`,
|
||||
alarmDescription: 'RDS CPU utilization above 80%',
|
||||
metric: dbInstance.metricCPUUtilization({ period: cdk.Duration.minutes(5) }),
|
||||
metric: dbCluster.metricCPUUtilization({ period: cdk.Duration.minutes(5) }),
|
||||
threshold: 80,
|
||||
evaluationPeriods: 3,
|
||||
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
||||
|
|
@ -308,19 +309,21 @@ export class FoundationStack extends cdk.Stack {
|
|||
new cloudwatch.Alarm(this, 'RdsConnectionsAlarm', {
|
||||
alarmName: `proposal-system-rds-connections${config.stackSuffix}`,
|
||||
alarmDescription: 'RDS database connections above 80',
|
||||
metric: dbInstance.metricDatabaseConnections({ period: cdk.Duration.minutes(5) }),
|
||||
metric: dbCluster.metricDatabaseConnections({ period: cdk.Duration.minutes(5) }),
|
||||
threshold: 80,
|
||||
evaluationPeriods: 2,
|
||||
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
|
||||
}),
|
||||
new cloudwatch.Alarm(this, 'RdsFreeStorageAlarm', {
|
||||
alarmName: `proposal-system-rds-free-storage${config.stackSuffix}`,
|
||||
alarmDescription: 'RDS free storage below 2 GB',
|
||||
metric: dbInstance.metricFreeStorageSpace({ period: cdk.Duration.minutes(5) }),
|
||||
threshold: 2_000_000_000,
|
||||
// Aurora storage auto-scales (no FreeStorageSpace); freeable memory is the
|
||||
// meaningful health signal for a Serverless v2 cluster.
|
||||
new cloudwatch.Alarm(this, 'RdsLowMemoryAlarm', {
|
||||
alarmName: `proposal-system-rds-low-memory${config.stackSuffix}`,
|
||||
alarmDescription: 'Aurora freeable memory below 256 MB',
|
||||
metric: dbCluster.metricFreeableMemory({ period: cdk.Duration.minutes(5) }),
|
||||
threshold: 256_000_000,
|
||||
comparisonOperator: cloudwatch.ComparisonOperator.LESS_THAN_THRESHOLD,
|
||||
evaluationPeriods: 1,
|
||||
treatMissingData: cloudwatch.TreatMissingData.BREACHING,
|
||||
evaluationPeriods: 3,
|
||||
treatMissingData: cloudwatch.TreatMissingData.NOT_BREACHING,
|
||||
}),
|
||||
];
|
||||
for (const alarm of rdsAlarms) {
|
||||
|
|
|
|||
126
lambdas/aurora-pgvector-init/app.py
Normal file
126
lambdas/aurora-pgvector-init/app.py
Normal file
|
|
@ -0,0 +1,126 @@
|
|||
"""CloudFormation custom-resource Lambda: initialize the Aurora pgvector store
|
||||
for the Bedrock Knowledge Base (PR3 — replaces the OpenSearch index creator).
|
||||
|
||||
Runs the one-time bootstrap SQL via the RDS Data API as the master user:
|
||||
enables pgvector, creates the `bedrock_integration` schema + `bedrock_kb` table
|
||||
(with the column/index layout Bedrock requires), and creates/owns the dedicated
|
||||
`bedrock_user` role whose credentials Bedrock uses to query the table.
|
||||
|
||||
Idempotent: safe to run on stack create/update (IF NOT EXISTS throughout). Only
|
||||
boto3 (rds-data, secretsmanager) is used — both ship in the Lambda runtime.
|
||||
"""
|
||||
|
||||
import json
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
|
||||
import boto3
|
||||
from botocore.exceptions import ClientError
|
||||
|
||||
logger = logging.getLogger(__name__)
|
||||
logger.setLevel(os.environ.get("LOG_LEVEL", "INFO"))
|
||||
|
||||
rds_data = boto3.client("rds-data")
|
||||
secrets = boto3.client("secretsmanager")
|
||||
|
||||
CLUSTER_ARN = os.environ["CLUSTER_ARN"]
|
||||
MASTER_SECRET_ARN = os.environ["MASTER_SECRET_ARN"]
|
||||
BEDROCK_SECRET_ARN = os.environ["BEDROCK_SECRET_ARN"]
|
||||
DATABASE = os.environ.get("DATABASE", "proposals")
|
||||
EMBED_DIM = int(os.environ.get("EMBED_DIM", "1024")) # Titan Embed v2
|
||||
|
||||
PHYSICAL_ID = "aurora-pgvector-init"
|
||||
TABLE = "bedrock_integration.bedrock_kb"
|
||||
|
||||
|
||||
def _exec(sql: str, attempts: int = 6) -> None:
|
||||
"""Run one statement via the RDS Data API as the master user.
|
||||
|
||||
Retries transient errors while a Serverless v2 cluster is still becoming
|
||||
reachable right after deploy (Data API can briefly report the cluster as
|
||||
unavailable / resuming).
|
||||
"""
|
||||
for attempt in range(attempts):
|
||||
try:
|
||||
rds_data.execute_statement(
|
||||
resourceArn=CLUSTER_ARN,
|
||||
secretArn=MASTER_SECRET_ARN,
|
||||
database=DATABASE,
|
||||
sql=sql,
|
||||
)
|
||||
return
|
||||
except ClientError as exc:
|
||||
code = exc.response.get("Error", {}).get("Code", "")
|
||||
message = str(exc)
|
||||
transient = code == "DatabaseResumingException" or any(
|
||||
s in message
|
||||
for s in ("not currently available", "Communication link failure")
|
||||
)
|
||||
if transient and attempt < attempts - 1:
|
||||
logger.warning(
|
||||
"Transient Data API error (attempt %d/%d): %s",
|
||||
attempt + 1,
|
||||
attempts,
|
||||
code or message,
|
||||
)
|
||||
time.sleep(10)
|
||||
continue
|
||||
raise
|
||||
|
||||
|
||||
def handler(event, context):
|
||||
request_type = event.get("RequestType")
|
||||
logger.info("RequestType=%s", request_type)
|
||||
|
||||
# Leave the data in place on stack delete — nothing to undo.
|
||||
if request_type == "Delete":
|
||||
return {"PhysicalResourceId": PHYSICAL_ID}
|
||||
|
||||
# The password Bedrock will use to log in as bedrock_user. Generated by
|
||||
# Secrets Manager with SQL-unsafe characters excluded (see CDK).
|
||||
secret = json.loads(
|
||||
secrets.get_secret_value(SecretId=BEDROCK_SECRET_ARN)["SecretString"]
|
||||
)
|
||||
password = secret["password"]
|
||||
# Defense-in-depth: the password is inlined into CREATE/ALTER ROLE SQL (DDL can't
|
||||
# bind parameters). The secret is generated with excludePunctuation=true, so it must
|
||||
# be strictly alphanumeric — refuse anything else rather than risk SQL breakage.
|
||||
if not password.isalnum():
|
||||
raise ValueError(
|
||||
"bedrock_user password is not alphanumeric; refusing to inline"
|
||||
)
|
||||
|
||||
statements = [
|
||||
"CREATE EXTENSION IF NOT EXISTS vector;",
|
||||
"CREATE SCHEMA IF NOT EXISTS bedrock_integration;",
|
||||
# Create the role if missing, then (re)set its password to match the secret.
|
||||
"DO $$ BEGIN "
|
||||
"IF NOT EXISTS (SELECT FROM pg_roles WHERE rolname = 'bedrock_user') "
|
||||
f"THEN CREATE ROLE bedrock_user LOGIN PASSWORD '{password}'; END IF; END $$;",
|
||||
f"ALTER ROLE bedrock_user WITH LOGIN PASSWORD '{password}';",
|
||||
"GRANT ALL ON SCHEMA bedrock_integration TO bedrock_user;",
|
||||
f"CREATE TABLE IF NOT EXISTS {TABLE} ("
|
||||
"id uuid PRIMARY KEY, "
|
||||
f"embedding vector({EMBED_DIM}), "
|
||||
"chunks text, "
|
||||
"metadata json, "
|
||||
"custom_metadata jsonb);",
|
||||
f"ALTER TABLE {TABLE} OWNER TO bedrock_user;",
|
||||
f"CREATE INDEX IF NOT EXISTS bedrock_kb_embedding_idx ON {TABLE} "
|
||||
"USING hnsw (embedding vector_cosine_ops) WITH (ef_construction = 256);",
|
||||
f"CREATE INDEX IF NOT EXISTS bedrock_kb_chunks_idx ON {TABLE} "
|
||||
"USING gin (to_tsvector('simple', chunks));",
|
||||
f"CREATE INDEX IF NOT EXISTS bedrock_kb_custom_metadata_idx ON {TABLE} "
|
||||
"USING gin (custom_metadata);",
|
||||
"GRANT ALL ON ALL TABLES IN SCHEMA bedrock_integration TO bedrock_user;",
|
||||
]
|
||||
|
||||
# Log only the position — the statements contain the bedrock_user password
|
||||
# (CREATE/ALTER ROLE), so the SQL text itself must never reach the logs.
|
||||
for i, sql in enumerate(statements, 1):
|
||||
logger.info("executing bootstrap statement %d/%d", i, len(statements))
|
||||
_exec(sql)
|
||||
|
||||
logger.info("pgvector store initialized: %s (dim=%d)", TABLE, EMBED_DIM)
|
||||
return {"PhysicalResourceId": PHYSICAL_ID, "Data": {"TableName": TABLE}}
|
||||
1
lambdas/aurora-pgvector-init/requirements.txt
Normal file
1
lambdas/aurora-pgvector-init/requirements.txt
Normal file
|
|
@ -0,0 +1 @@
|
|||
boto3>=1.43.18,<2.0
|
||||
|
|
@ -1,71 +0,0 @@
|
|||
import time
|
||||
import boto3
|
||||
from opensearchpy import OpenSearch, RequestsHttpConnection
|
||||
from requests_aws4auth import AWS4Auth
|
||||
|
||||
|
||||
def handler(event, context):
|
||||
if event["RequestType"] == "Delete":
|
||||
return {"PhysicalResourceId": event.get("PhysicalResourceId", "none")}
|
||||
|
||||
props = event["ResourceProperties"]
|
||||
endpoint = props["Endpoint"].replace("https://", "")
|
||||
index_name = props["IndexName"]
|
||||
vector_field = props["VectorField"]
|
||||
text_field = props["TextField"]
|
||||
metadata_field = props["MetadataField"]
|
||||
|
||||
session = boto3.Session()
|
||||
credentials = session.get_credentials().get_frozen_credentials()
|
||||
region = session.region_name
|
||||
|
||||
awsauth = AWS4Auth(
|
||||
credentials.access_key,
|
||||
credentials.secret_key,
|
||||
region,
|
||||
"aoss",
|
||||
session_token=credentials.token,
|
||||
)
|
||||
|
||||
client = OpenSearch(
|
||||
hosts=[{"host": endpoint, "port": 443}],
|
||||
http_auth=awsauth,
|
||||
use_ssl=True,
|
||||
verify_certs=True,
|
||||
connection_class=RequestsHttpConnection,
|
||||
timeout=30,
|
||||
)
|
||||
|
||||
index_body = {
|
||||
"settings": {"index": {"knn": True, "knn.algo_param.ef_search": 512}},
|
||||
"mappings": {
|
||||
"properties": {
|
||||
vector_field: {
|
||||
"type": "knn_vector",
|
||||
"dimension": 1024,
|
||||
"method": {
|
||||
"engine": "faiss",
|
||||
"name": "hnsw",
|
||||
"space_type": "l2",
|
||||
},
|
||||
},
|
||||
text_field: {"type": "text"},
|
||||
metadata_field: {"type": "text"},
|
||||
}
|
||||
},
|
||||
}
|
||||
|
||||
for attempt in range(30):
|
||||
try:
|
||||
client.indices.create(index=index_name, body=index_body)
|
||||
return {"PhysicalResourceId": index_name}
|
||||
except Exception as e:
|
||||
error_str = str(e)
|
||||
if "resource_already_exists_exception" in error_str:
|
||||
return {"PhysicalResourceId": index_name}
|
||||
if "403" in error_str and attempt < 29:
|
||||
time.sleep(10)
|
||||
continue
|
||||
raise
|
||||
|
||||
raise Exception("Timeout waiting for AOSS access policy propagation")
|
||||
|
|
@ -1,3 +0,0 @@
|
|||
opensearch-py>=3.2.0
|
||||
requests-aws4auth>=1.3.2
|
||||
requests>=2.34.2
|
||||
Loading…
Add table
Reference in a new issue