mirror of
https://github.com/Sea-Haven-Industries/sh-mcp.git
synced 2026-09-30 22:53:17 +00:00
Some checks are pending
deploy / deploy (push) Waiting to run
* Add shared transport: dispatch, MCP + OpenAPI adapters, local auth
Add the single authoritative tool-execution path (executeTool) plus the two
universal interfaces over it (design.md §2.5, §7.3):
- dispatch.ts: scope enforcement, ajv input validation, rate limiting, finance
egress redaction (redactDeep), and structured audit emission on one path.
- audit.ts / rate-limit.ts: injected AuditLogger + RateLimiter abstractions.
- mcp.ts: low-level MCP Server with scope-filtered tools/list (tool-hiding) and
tools/call routed through executeTool.
- http.ts: Express host mounting /mcp, /openapi.json, POST /tools/:name, /healthz.
- openapi.ts: buildOpenApiDocument wraps the existing path generator into a full
OpenAPI 3.1 document.
- local-auth.ts: LocalAuthProvider (dev bearer tokens) that refuses to construct
outside SH_MCP_ENV=local and enforces audience binding (design.md §3, §6).
* Add in-memory dev clients; make package tool exports lazy
Add an in-memory Client implementation per integration package (seeded fake
data, no network) selected when SH_MCP_ENV=local (build-plan §4). Gmail/calendar/
tasks dev clients partition by ctx.sub; payments/qbo seed sensitive-looking
fields so the redaction egress path has real targets to mask.
Make the eager default-tool exports in tasks/reminders/qbo LAZY (getDefaultTools)
so importing a package barrel no longer constructs an AWS client at module load
(build-plan §7 'no I/O at import time') — the previous eager construction broke
server startup. Fix payments tsconfig rootDir (src, was '.') so its declarations
resolve under dist/index.d.ts like the other 8 packages.
* Add runnable sh-mcp-ops and sh-mcp-finance servers
Two thin composition-root servers over the shared transport (design.md §3):
- ops: internal-data, knowledge-base, google-maps, gmail, calendar, tasks,
reminders. finance: qbo, payments (audited + redacted on egress).
- config from env only (no hardcoded ids/issuer/tables); SH_MCP_ENV selects
LocalAuthProvider + dev clients (local) vs CognitoAuthProvider + real stubs
(aws). Finance applies the 15-min finance-token TTL ceiling (design.md §2.5).
- index.ts is the only place .listen() is called; a Lambda handler placeholder
is exported but not depended on.
- synth-only CDK stubs (no real IAM/Cognito/WAF) so 'cdk synth' has a valid app
(build-plan §6); READMEs document local run, dev tokens, curl, MCP Inspector.
* Add security-weighted test suite + coverage gate; wire tooling
Add tests for the highest-risk surface (build-plan §5, design.md §7.3):
tool-hiding, server-side scope enforcement (incl. forced hidden calls),
audience binding, input-schema validation, finance redaction on egress, audit
emission with hashed args, prompt-injection regression (tool output is data),
rate limiting, MCP conformance (in-memory transport round-trip), OpenAPI 3.1
validity, and local-auth safety. Add HTTP integration tests (supertest) for both
servers and per-package dev-client tests. 405 tests pass.
Wire the coverage gate into vitest.config.ts: 80% overall, with per-file
thresholds on the auth + dispatch crown jewels; exclude deferred real client
stubs, entrypoints, cdk apps, and aws-only config from the gate (documented).
Extend eslint flat config + add .prettierignore to cover servers/. Commit the
updated package-lock.json.
* Suppress pre-existing dev-tooling + out-of-scope scanner findings
Add written-justification suppressions for the 4 confirmed crit/high pre-push
scanner findings, none of which are in this PR's Phase 1 production code:
- npmaudit vitest / @vitest/coverage-v8 / vite: dev/test-only deps that never
run in the deployed server/Lambda runtime (pins carried from Phase 0b;
Dependabot will bump).
- gitleaks docs/agentforce-plan.md secret: that file is not on this branch and
not in this changeset; flagged for the maintainer to scrub on its own branch.
The deep agentic /sh-security-review (required for this auth/authz-touching PR)
was NOT run by the agent and is flagged outstanding in the PR body.
* Address CodeQL findings: bound ajv error work + edge rate limiting
GHAS code-scanning alerts on this PR:
- dispatch.ts (js/resource-exhaustion): ajv ran with allErrors:true on
untrusted input, letting a crafted payload force unbounded error
enumeration. Switch to allErrors:false (default) so validation
short-circuits on the first failure; the 400 still names that path.
- http.ts (js/missing-rate-limiting): the authenticated routes (/mcp,
/tools/:name) had no edge throttle — auth/JWT verification ran on every
request before the per-sub dispatch limiter could apply. Add an IP-keyed
express-rate-limit in front of authenticate (120/60s default, configurable),
returning the standard 429 shape. Defense-in-depth over the per-sub +
per-tool limiter in executeTool; API GW/WAF remains the production edge.
Tests: +2 cases proving the edge limiter throttles before auth (429, not
401) on /tools and /mcp. 407 pass; tsc/eslint/prettier clean.
* Fix polynomial ReDoS in Bearer-token extraction (CodeQL js/polynomial-redos)
extractBearerToken matched /^Bearer\s+(.+)$/ — \s and . both match a space,
so the two quantifiers overlap and a crafted header can drive polynomial
backtracking. Require the capture to start with a non-whitespace char
(/^Bearer\s+(\S.*)$/), removing the ambiguity → linear match. Behavior is
unchanged for real tokens; +2 regression tests.
* Harden auth + finance redaction (sh-security-review confirmed mediums)
Two confirmed medium findings from the agentic security review:
- Fail-open SH_MCP_ENV: config defaulted to 'local' when the var was unset,
so a deploy that forgot SH_MCP_ENV=aws would silently run LocalAuthProvider
and accept static dev bearer tokens (dev-finance-admin -> finance:admin).
Now fail-closed: SH_MCP_ENV must be explicitly 'local' or 'aws' or the
server refuses to start. Plus an independent guard in LocalAuthProvider
that refuses to construct in an AWS runtime (AWS_LAMBDA_FUNCTION_NAME /
AWS_EXECUTION_ENV present), regardless of the env flag.
- Finance egress redaction gap: redactDeep only wholesale-masked a sensitive
key when its value was a scalar; an object/array under a sensitive key was
recursed into, letting a bare nested value (e.g. {account:{number:...}})
escape the keyword-gated pattern matcher. Now the entire subtree under a
sensitive key is masked. No current finance tool emitted such shapes (all
flat strings), so this closes a latent hole in the universal safety net.
+4 tests (subtree redaction, AWS-runtime guard). 411 pass; coverage gate green.
Review also produced lows (memo free-text digits, unsalted argsHash,
unauth /openapi.json by-design, session-cap no-reset by-design) tracked
separately; 0 confirmed critical/high — review verdict PASS.
228 lines
9.4 KiB
TypeScript
228 lines
9.4 KiB
TypeScript
/**
|
|
* The single authoritative tool-execution path.
|
|
*
|
|
* Both transports — MCP (`mcp.ts`) and OpenAPI/HTTP (`http.ts`) — route every
|
|
* tool call through {@link executeTool}. Centralizing here is what makes the
|
|
* security guarantees of design.md §2.5 hold uniformly across interfaces:
|
|
*
|
|
* 1. tool lookup (unknown → `UnknownToolError` → 404)
|
|
* 2. server-side scope enforcement (`requireScope`; defense-in-depth — handlers
|
|
* also call it). Tool-hiding in the UI is NOT the boundary (design.md §2.5).
|
|
* 3. input-schema validation BEFORE the handler runs (ajv) — never pass
|
|
* unvalidated input to a handler (design.md §2.5 prompt-injection containment).
|
|
* 4. per-session cap + per-tool rate limit (design.md §7.3).
|
|
* 5. handler execution.
|
|
* 6. finance-tier redaction on egress — bank/routing/card/SSN masked before the
|
|
* value leaves the dispatcher (design.md §2.5, §7.3). Belt-and-braces over
|
|
* each finance tool's own internal redaction.
|
|
* 7. structured audit record for every finance (and future physical) call —
|
|
* args HASHED, never logged raw; no secrets (design.md §2.5, §7.3).
|
|
*
|
|
* Tool output is treated strictly as DATA, never as instructions: the dispatcher
|
|
* inspects/redacts it but never re-enters itself based on its content, so a
|
|
* prompt-injection payload in a tool response cannot trigger another tool call
|
|
* (design.md §2.5).
|
|
*/
|
|
|
|
import AjvModule from 'ajv';
|
|
import addFormatsModule from 'ajv-formats';
|
|
import type { Ajv as AjvInstance, Options as AjvOptions, ValidateFunction } from 'ajv';
|
|
|
|
// ajv / ajv-formats ship as CJS; under NodeNext ESM the callable lives on
|
|
// `.default`. Normalize so both module shapes work, then re-type the runtime
|
|
// values as the constructable class / callable plugin.
|
|
type AjvCtor = new (opts?: AjvOptions) => AjvInstance;
|
|
const Ajv = ((AjvModule as { default?: unknown }).default ?? AjvModule) as unknown as AjvCtor;
|
|
const addFormats = ((addFormatsModule as { default?: unknown }).default ?? addFormatsModule) as (
|
|
ajv: AjvInstance,
|
|
) => AjvInstance;
|
|
|
|
import { requireScope, ScopeError } from './auth.js';
|
|
import { hashArgs, type AuditLogger, type AuditRecord } from './audit.js';
|
|
import { redact } from './redact.js';
|
|
import { RateLimitError, type RateLimiter } from './rate-limit.js';
|
|
import type { ToolRegistry } from './registry.js';
|
|
import type { AuthContext, ToolDef } from './types.js';
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Errors (adapters translate these to status codes)
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** Unknown tool name → HTTP 404 / MCP "method not found"-equivalent. */
|
|
export class UnknownToolError extends Error {
|
|
readonly tool: string;
|
|
constructor(tool: string) {
|
|
super(`Unknown tool "${tool}".`);
|
|
this.name = 'UnknownToolError';
|
|
this.tool = tool;
|
|
Object.setPrototypeOf(this, new.target.prototype);
|
|
}
|
|
}
|
|
|
|
/** Input failed JSON-Schema validation → HTTP 400. */
|
|
export class InputValidationError extends Error {
|
|
readonly tool: string;
|
|
/** Human-readable validation messages; never echoes secrets. */
|
|
readonly issues: string[];
|
|
constructor(tool: string, issues: string[]) {
|
|
super(`Input validation failed for "${tool}": ${issues.join('; ')}`);
|
|
this.name = 'InputValidationError';
|
|
this.tool = tool;
|
|
this.issues = issues;
|
|
Object.setPrototypeOf(this, new.target.prototype);
|
|
}
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Dependencies injected into the dispatcher
|
|
// ---------------------------------------------------------------------------
|
|
|
|
export interface DispatchDeps {
|
|
/** Audit sink — finance/physical calls emit a record here. */
|
|
auditLogger: AuditLogger;
|
|
/** Per-session + per-tool limiter consulted before each handler runs. */
|
|
rateLimiter: RateLimiter;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// AJV — compiled validators cached per tool
|
|
// ---------------------------------------------------------------------------
|
|
|
|
// One Ajv instance for the process. Schemas are JSON Schema (draft-07 / the
|
|
// OpenAPI 3.1 subset our tools use). `strict: false` because tool authors use
|
|
// vocabulary (e.g. `description`) liberally; we only need structural validation.
|
|
//
|
|
// `allErrors: false` (the default) is deliberate and security-relevant: the
|
|
// input is UNTRUSTED, and `allErrors: true` makes ajv enumerate every schema
|
|
// violation, which an attacker can weaponize into CPU/memory exhaustion by
|
|
// sending a large/deeply-nested payload that fails many constraints at once
|
|
// (CodeQL js/resource-exhaustion). Short-circuiting on the first error caps the
|
|
// work per request; the 400 still names the first failing path, which is enough
|
|
// for a caller to fix their input.
|
|
const ajv = new Ajv({ allErrors: false, strict: false, coerceTypes: false });
|
|
addFormats(ajv);
|
|
|
|
const validatorCache = new WeakMap<object, ValidateFunction>();
|
|
|
|
function getValidator(tool: ToolDef<unknown, unknown>): ValidateFunction {
|
|
const schema = tool.inputSchema as object;
|
|
const cached = validatorCache.get(schema);
|
|
if (cached) return cached;
|
|
const validate = ajv.compile(schema);
|
|
validatorCache.set(schema, validate);
|
|
return validate;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// Finance egress redaction
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/** Field names whose values are masked wholesale on finance egress. */
|
|
const SENSITIVE_FIELD_RE = /(account|routing|card|ssn|tax[_-]?id|iban|swift)/i;
|
|
|
|
/**
|
|
* Deep-redact a finance tool's output before it leaves the dispatcher.
|
|
*
|
|
* Two complementary passes (design.md §2.5):
|
|
* - String values run through `redact()` (pattern-based: routing/account/card/SSN).
|
|
* - Any field whose KEY looks sensitive is fully replaced with `[REDACTED]`,
|
|
* catching isolated values (e.g. a bare `bankAccountNumber: "123456789"`)
|
|
* that the inline pattern matcher would miss without keyword context.
|
|
*
|
|
* Returns a NEW structure; the handler's value is not mutated.
|
|
*/
|
|
export function redactDeep(value: unknown): unknown {
|
|
if (typeof value === 'string') return redact(value);
|
|
if (Array.isArray(value)) return value.map(redactDeep);
|
|
if (value !== null && typeof value === 'object') {
|
|
const out: Record<string, unknown> = {};
|
|
for (const [key, v] of Object.entries(value as Record<string, unknown>)) {
|
|
if (SENSITIVE_FIELD_RE.test(key) && v !== null && v !== undefined) {
|
|
// Key looks sensitive → mask the ENTIRE value wholesale, whether it is a
|
|
// scalar, an array, or a nested object. Recursing into a non-scalar here
|
|
// would lose the keyword context redact() needs, letting a bare nested
|
|
// value (e.g. { account: { number: "021000021" } }) escape unmasked.
|
|
// Over-masking on the finance tier is the correct trade: a false negative
|
|
// is a PII leak (design.md §2.5).
|
|
out[key] = '[REDACTED]';
|
|
} else {
|
|
out[key] = redactDeep(v);
|
|
}
|
|
}
|
|
return out;
|
|
}
|
|
return value;
|
|
}
|
|
|
|
// ---------------------------------------------------------------------------
|
|
// executeTool — the one path
|
|
// ---------------------------------------------------------------------------
|
|
|
|
/**
|
|
* Look up, authorize, validate, rate-limit, run, redact, and audit a single
|
|
* tool call. Used identically by both transports.
|
|
*
|
|
* @throws {UnknownToolError} unknown tool name (→ 404)
|
|
* @throws {ScopeError} caller lacks the required scope (→ 403)
|
|
* @throws {InputValidationError} input failed schema validation (→ 400)
|
|
* @throws {RateLimitError} limiter rejected the call (→ 429)
|
|
* @throws {Error} handler threw (→ 500; message not leaked verbatim)
|
|
*/
|
|
export async function executeTool(
|
|
registry: ToolRegistry,
|
|
ctx: AuthContext,
|
|
toolName: string,
|
|
rawInput: unknown,
|
|
deps: DispatchDeps,
|
|
): Promise<unknown> {
|
|
const tool = registry.get(toolName) as ToolDef<unknown, unknown> | undefined;
|
|
if (!tool) {
|
|
throw new UnknownToolError(toolName);
|
|
}
|
|
|
|
const audited = tool.tier === 'finance';
|
|
let decision: AuditRecord['decision'] = 'deny';
|
|
let result: AuditRecord['result'] = 'error';
|
|
|
|
try {
|
|
// 2. Server-side scope enforcement (authoritative; not UI tool-hiding).
|
|
requireScope(ctx, tool.requiredScope);
|
|
decision = 'allow';
|
|
|
|
// 3. Input-schema validation BEFORE the handler sees the input.
|
|
const validate = getValidator(tool);
|
|
if (!validate(rawInput)) {
|
|
const issues = (validate.errors ?? []).map(
|
|
(e) => `${e.instancePath || '(root)'} ${e.message ?? 'is invalid'}`,
|
|
);
|
|
throw new InputValidationError(toolName, issues.length ? issues : ['invalid input']);
|
|
}
|
|
|
|
// 4. Rate limit / session cap.
|
|
deps.rateLimiter.check(ctx.sub, toolName);
|
|
|
|
// 5. Run the handler. Its output is data only.
|
|
const output = await tool.handler(rawInput, ctx);
|
|
|
|
// 6. Finance egress redaction (belt-and-braces over internal redaction).
|
|
const safeOutput = tool.tier === 'finance' ? redactDeep(output) : output;
|
|
|
|
result = 'ok';
|
|
return safeOutput;
|
|
} finally {
|
|
// 7. Audit every finance/physical call regardless of outcome.
|
|
if (audited) {
|
|
deps.auditLogger.log({
|
|
sub: ctx.sub,
|
|
tool: toolName,
|
|
argsHash: hashArgs(rawInput),
|
|
decision,
|
|
result,
|
|
ts: new Date().toISOString(),
|
|
});
|
|
}
|
|
}
|
|
}
|
|
|
|
// Re-export the error types adapters need to translate outcomes.
|
|
export { ScopeError, RateLimitError };
|