mirror of
https://github.com/Sea-Haven-Industries/sh-mcp.git
synced 2026-09-30 19:23:19 +00:00
Some checks are pending
deploy / deploy (push) Waiting to run
* Add shared transport: dispatch, MCP + OpenAPI adapters, local auth
Add the single authoritative tool-execution path (executeTool) plus the two
universal interfaces over it (design.md §2.5, §7.3):
- dispatch.ts: scope enforcement, ajv input validation, rate limiting, finance
egress redaction (redactDeep), and structured audit emission on one path.
- audit.ts / rate-limit.ts: injected AuditLogger + RateLimiter abstractions.
- mcp.ts: low-level MCP Server with scope-filtered tools/list (tool-hiding) and
tools/call routed through executeTool.
- http.ts: Express host mounting /mcp, /openapi.json, POST /tools/:name, /healthz.
- openapi.ts: buildOpenApiDocument wraps the existing path generator into a full
OpenAPI 3.1 document.
- local-auth.ts: LocalAuthProvider (dev bearer tokens) that refuses to construct
outside SH_MCP_ENV=local and enforces audience binding (design.md §3, §6).
* Add in-memory dev clients; make package tool exports lazy
Add an in-memory Client implementation per integration package (seeded fake
data, no network) selected when SH_MCP_ENV=local (build-plan §4). Gmail/calendar/
tasks dev clients partition by ctx.sub; payments/qbo seed sensitive-looking
fields so the redaction egress path has real targets to mask.
Make the eager default-tool exports in tasks/reminders/qbo LAZY (getDefaultTools)
so importing a package barrel no longer constructs an AWS client at module load
(build-plan §7 'no I/O at import time') — the previous eager construction broke
server startup. Fix payments tsconfig rootDir (src, was '.') so its declarations
resolve under dist/index.d.ts like the other 8 packages.
* Add runnable sh-mcp-ops and sh-mcp-finance servers
Two thin composition-root servers over the shared transport (design.md §3):
- ops: internal-data, knowledge-base, google-maps, gmail, calendar, tasks,
reminders. finance: qbo, payments (audited + redacted on egress).
- config from env only (no hardcoded ids/issuer/tables); SH_MCP_ENV selects
LocalAuthProvider + dev clients (local) vs CognitoAuthProvider + real stubs
(aws). Finance applies the 15-min finance-token TTL ceiling (design.md §2.5).
- index.ts is the only place .listen() is called; a Lambda handler placeholder
is exported but not depended on.
- synth-only CDK stubs (no real IAM/Cognito/WAF) so 'cdk synth' has a valid app
(build-plan §6); READMEs document local run, dev tokens, curl, MCP Inspector.
* Add security-weighted test suite + coverage gate; wire tooling
Add tests for the highest-risk surface (build-plan §5, design.md §7.3):
tool-hiding, server-side scope enforcement (incl. forced hidden calls),
audience binding, input-schema validation, finance redaction on egress, audit
emission with hashed args, prompt-injection regression (tool output is data),
rate limiting, MCP conformance (in-memory transport round-trip), OpenAPI 3.1
validity, and local-auth safety. Add HTTP integration tests (supertest) for both
servers and per-package dev-client tests. 405 tests pass.
Wire the coverage gate into vitest.config.ts: 80% overall, with per-file
thresholds on the auth + dispatch crown jewels; exclude deferred real client
stubs, entrypoints, cdk apps, and aws-only config from the gate (documented).
Extend eslint flat config + add .prettierignore to cover servers/. Commit the
updated package-lock.json.
* Suppress pre-existing dev-tooling + out-of-scope scanner findings
Add written-justification suppressions for the 4 confirmed crit/high pre-push
scanner findings, none of which are in this PR's Phase 1 production code:
- npmaudit vitest / @vitest/coverage-v8 / vite: dev/test-only deps that never
run in the deployed server/Lambda runtime (pins carried from Phase 0b;
Dependabot will bump).
- gitleaks docs/agentforce-plan.md secret: that file is not on this branch and
not in this changeset; flagged for the maintainer to scrub on its own branch.
The deep agentic /sh-security-review (required for this auth/authz-touching PR)
was NOT run by the agent and is flagged outstanding in the PR body.
* Address CodeQL findings: bound ajv error work + edge rate limiting
GHAS code-scanning alerts on this PR:
- dispatch.ts (js/resource-exhaustion): ajv ran with allErrors:true on
untrusted input, letting a crafted payload force unbounded error
enumeration. Switch to allErrors:false (default) so validation
short-circuits on the first failure; the 400 still names that path.
- http.ts (js/missing-rate-limiting): the authenticated routes (/mcp,
/tools/:name) had no edge throttle — auth/JWT verification ran on every
request before the per-sub dispatch limiter could apply. Add an IP-keyed
express-rate-limit in front of authenticate (120/60s default, configurable),
returning the standard 429 shape. Defense-in-depth over the per-sub +
per-tool limiter in executeTool; API GW/WAF remains the production edge.
Tests: +2 cases proving the edge limiter throttles before auth (429, not
401) on /tools and /mcp. 407 pass; tsc/eslint/prettier clean.
* Fix polynomial ReDoS in Bearer-token extraction (CodeQL js/polynomial-redos)
extractBearerToken matched /^Bearer\s+(.+)$/ — \s and . both match a space,
so the two quantifiers overlap and a crafted header can drive polynomial
backtracking. Require the capture to start with a non-whitespace char
(/^Bearer\s+(\S.*)$/), removing the ambiguity → linear match. Behavior is
unchanged for real tokens; +2 regression tests.
* Harden auth + finance redaction (sh-security-review confirmed mediums)
Two confirmed medium findings from the agentic security review:
- Fail-open SH_MCP_ENV: config defaulted to 'local' when the var was unset,
so a deploy that forgot SH_MCP_ENV=aws would silently run LocalAuthProvider
and accept static dev bearer tokens (dev-finance-admin -> finance:admin).
Now fail-closed: SH_MCP_ENV must be explicitly 'local' or 'aws' or the
server refuses to start. Plus an independent guard in LocalAuthProvider
that refuses to construct in an AWS runtime (AWS_LAMBDA_FUNCTION_NAME /
AWS_EXECUTION_ENV present), regardless of the env flag.
- Finance egress redaction gap: redactDeep only wholesale-masked a sensitive
key when its value was a scalar; an object/array under a sensitive key was
recursed into, letting a bare nested value (e.g. {account:{number:...}})
escape the keyword-gated pattern matcher. Now the entire subtree under a
sensitive key is masked. No current finance tool emitted such shapes (all
flat strings), so this closes a latent hole in the universal safety net.
+4 tests (subtree redaction, AWS-runtime guard). 411 pass; coverage gate green.
Review also produced lows (memo free-text digits, unsalted argsHash,
unauth /openapi.json by-design, session-cap no-reset by-design) tracked
separately; 0 confirmed critical/high — review verdict PASS.
291 lines
10 KiB
TypeScript
291 lines
10 KiB
TypeScript
/**
|
|
* Unit tests for @sh-mcp/knowledge-base.
|
|
*
|
|
* All tests use:
|
|
* - A mock AuthContext with the required ops:read scope.
|
|
* - A mock KnowledgeBaseClient — no AWS calls, no network.
|
|
*
|
|
* Coverage targets:
|
|
* - Happy path: results returned and shaped correctly.
|
|
* - Empty result: KB returns no matches.
|
|
* - Scope enforcement: missing scope throws ScopeError.
|
|
* - Client error: upstream error surfaces as a rejected promise.
|
|
* - Throttle / retry: upstream ThrottlingException propagates (retry
|
|
* logic, if added, would be tested here).
|
|
*/
|
|
|
|
import { describe, it, expect, vi, beforeEach } from 'vitest';
|
|
import { createKnowledgeBaseTools } from '../src/tools.js';
|
|
import type { KnowledgeBaseClient, KnowledgeBaseResult } from '../src/client.js';
|
|
import type { AuthContext } from '@sh-mcp/shared';
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Test fixtures
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** A valid AuthContext carrying the ops:read scope. */
|
|
const authorisedCtx: AuthContext = {
|
|
sub: 'lauren@seahavenind.com',
|
|
scopes: ['ops:read'],
|
|
aud: 'sh-mcp-ops',
|
|
};
|
|
|
|
/** An AuthContext with no scopes — used to test scope enforcement. */
|
|
const unauthorisedCtx: AuthContext = {
|
|
sub: 'guest@example.com',
|
|
scopes: [],
|
|
aud: 'sh-mcp-ops',
|
|
};
|
|
|
|
/** A finance-only AuthContext (finance:read but not ops:read). */
|
|
const financeOnlyCtx: AuthContext = {
|
|
sub: 'accounting@seahavenind.com',
|
|
scopes: ['finance:read'],
|
|
aud: 'sh-mcp-finance',
|
|
};
|
|
|
|
/** Sample KB results returned by the mock. */
|
|
const sampleResults: KnowledgeBaseResult[] = [
|
|
{
|
|
source: 's3://sh-kb-data/notion/procedures.md',
|
|
score: 0.92,
|
|
passage: 'All maintenance requests must be submitted via the work-order portal.',
|
|
},
|
|
{
|
|
source: 's3://sh-kb-data/work-orders/WO-1042.md',
|
|
score: 0.78,
|
|
passage: 'Work order 1042: HVAC inspection completed 2025-11-15.',
|
|
},
|
|
];
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Mock client factory
|
|
// ---------------------------------------------------------------------------
|
|
|
|
function makeMockClient(implementation?: Partial<KnowledgeBaseClient>): KnowledgeBaseClient {
|
|
return {
|
|
retrieve: vi.fn().mockResolvedValue(sampleResults),
|
|
...implementation,
|
|
};
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Tests
|
|
// ---------------------------------------------------------------------------
|
|
|
|
describe('search_knowledge_base', () => {
|
|
let mockClient: KnowledgeBaseClient;
|
|
|
|
beforeEach(() => {
|
|
mockClient = makeMockClient();
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Happy path
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('returns shaped results for a valid query', async () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
const output = await tool.handler(
|
|
{ query: 'maintenance procedures', maxResults: 5 },
|
|
authorisedCtx,
|
|
);
|
|
|
|
expect(output.count).toBe(2);
|
|
expect(output.results).toHaveLength(2);
|
|
|
|
const [first] = output.results;
|
|
expect(first.source).toBe('s3://sh-kb-data/notion/procedures.md');
|
|
expect(first.score).toBe(0.92);
|
|
expect(first.passage).toContain('maintenance requests');
|
|
});
|
|
|
|
it('passes query and maxResults through to the client', async () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
await tool.handler({ query: 'HVAC vendors', maxResults: 3 }, authorisedCtx);
|
|
|
|
expect(mockClient.retrieve).toHaveBeenCalledOnce();
|
|
expect(mockClient.retrieve).toHaveBeenCalledWith({
|
|
query: 'HVAC vendors',
|
|
maxResults: 3,
|
|
});
|
|
});
|
|
|
|
it('defaults maxResults to 5 when omitted', async () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
await tool.handler({ query: 'fire safety' }, authorisedCtx);
|
|
|
|
expect(mockClient.retrieve).toHaveBeenCalledWith({
|
|
query: 'fire safety',
|
|
maxResults: 5,
|
|
});
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Tool metadata assertions
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('has correct tool metadata', () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
expect(tool.name).toBe('search_knowledge_base');
|
|
expect(tool.tier).toBe('ops');
|
|
expect(tool.requiredScope).toBe('ops:read');
|
|
expect(tool.description).toMatch(/knowledge base/i);
|
|
});
|
|
|
|
it('exports exactly one tool', () => {
|
|
const tools = createKnowledgeBaseTools(mockClient);
|
|
expect(tools).toHaveLength(1);
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Empty result
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('returns an empty results array when the KB finds no matches', async () => {
|
|
const emptyClient = makeMockClient({
|
|
retrieve: vi.fn().mockResolvedValue([]),
|
|
});
|
|
const [tool] = createKnowledgeBaseTools(emptyClient);
|
|
|
|
const output = await tool.handler({ query: 'nonexistent topic xyz' }, authorisedCtx);
|
|
|
|
expect(output.count).toBe(0);
|
|
expect(output.results).toEqual([]);
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Scope enforcement
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('throws ScopeError when the caller has no scopes', async () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
await expect(tool.handler({ query: 'anything' }, unauthorisedCtx)).rejects.toThrow();
|
|
|
|
// The client must NOT be called when auth fails.
|
|
expect(mockClient.retrieve).not.toHaveBeenCalled();
|
|
});
|
|
|
|
it('throws ScopeError when the caller only has a finance scope (not ops:read)', async () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
await expect(tool.handler({ query: 'anything' }, financeOnlyCtx)).rejects.toThrow();
|
|
|
|
expect(mockClient.retrieve).not.toHaveBeenCalled();
|
|
});
|
|
|
|
it('succeeds when the caller has ops:read among multiple scopes', async () => {
|
|
const multiScopeCtx: AuthContext = {
|
|
sub: 'adam@seahavenind.com',
|
|
scopes: ['ops:read', 'ops:tasks', 'finance:read', 'finance:admin'],
|
|
aud: 'sh-mcp-ops',
|
|
};
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
|
|
const output = await tool.handler({ query: 'anything' }, multiScopeCtx);
|
|
expect(output.count).toBe(2);
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Client error
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('surfaces a client error as a rejected promise', async () => {
|
|
const errorClient = makeMockClient({
|
|
retrieve: vi.fn().mockRejectedValue(new Error('Bedrock Retrieve failed')),
|
|
});
|
|
const [tool] = createKnowledgeBaseTools(errorClient);
|
|
|
|
await expect(tool.handler({ query: 'HVAC' }, authorisedCtx)).rejects.toThrow(
|
|
'Bedrock Retrieve failed',
|
|
);
|
|
});
|
|
|
|
it('surfaces an unexpected error type without swallowing it', async () => {
|
|
const weirdClient = makeMockClient({
|
|
retrieve: vi.fn().mockRejectedValue('string error'),
|
|
});
|
|
const [tool] = createKnowledgeBaseTools(weirdClient);
|
|
|
|
await expect(tool.handler({ query: 'test' }, authorisedCtx)).rejects.toBe('string error');
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Throttle / retry
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('propagates a ThrottlingException from the client', async () => {
|
|
// Simulate the shape Bedrock SDK throws for throttling.
|
|
const throttleError = Object.assign(new Error('Too many requests'), {
|
|
name: 'ThrottlingException',
|
|
$fault: 'client',
|
|
$retryable: { throttling: true },
|
|
});
|
|
|
|
const throttledClient = makeMockClient({
|
|
retrieve: vi.fn().mockRejectedValue(throttleError),
|
|
});
|
|
const [tool] = createKnowledgeBaseTools(throttledClient);
|
|
|
|
const rejection = await tool
|
|
.handler({ query: 'anything' }, authorisedCtx)
|
|
.catch((e: unknown) => e);
|
|
|
|
expect((rejection as Error).name).toBe('ThrottlingException');
|
|
});
|
|
|
|
it('propagates throttle on first call (retry logic placeholder)', async () => {
|
|
// When retry logic is added (e.g. exponential back-off wrapper), update
|
|
// this test to assert the mock is called N times and eventually succeeds.
|
|
// For now assert the error propagates unchanged so the server layer can
|
|
// apply its own retry strategy.
|
|
const throttleError = Object.assign(new Error('Too many requests'), {
|
|
name: 'ThrottlingException',
|
|
});
|
|
|
|
const client = makeMockClient({
|
|
retrieve: vi.fn().mockRejectedValue(throttleError),
|
|
});
|
|
const [tool] = createKnowledgeBaseTools(client);
|
|
|
|
await expect(tool.handler({ query: 'test' }, authorisedCtx)).rejects.toMatchObject({
|
|
name: 'ThrottlingException',
|
|
});
|
|
|
|
// Exactly one attempt — no retry implemented yet.
|
|
expect(client.retrieve).toHaveBeenCalledOnce();
|
|
});
|
|
|
|
// -------------------------------------------------------------------------
|
|
// Input schema assertions (contract)
|
|
// -------------------------------------------------------------------------
|
|
|
|
it('declares query as a required string in the inputSchema', () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
const schema = tool.inputSchema as {
|
|
required: string[];
|
|
properties: Record<string, { type: string }>;
|
|
};
|
|
|
|
expect(schema.required).toContain('query');
|
|
expect(schema.properties['query'].type).toBe('string');
|
|
});
|
|
|
|
it('declares maxResults as an optional integer in the inputSchema', () => {
|
|
const [tool] = createKnowledgeBaseTools(mockClient);
|
|
const schema = tool.inputSchema as {
|
|
required: string[];
|
|
properties: Record<string, { type: string; minimum: number; maximum: number }>;
|
|
};
|
|
|
|
expect(schema.required).not.toContain('maxResults');
|
|
expect(schema.properties['maxResults'].type).toBe('integer');
|
|
expect(schema.properties['maxResults'].minimum).toBe(1);
|
|
expect(schema.properties['maxResults'].maximum).toBe(20);
|
|
});
|
|
});
|