G/00 // Guide — Compliance

How do you build an in-house AI KYC system?

You build an in-house AI KYC system in 2026 by treating identity verification vendors (Onfido, Sumsub, Jumio, Veriff) as interchangeable data sources and owning the risk-decision layer above them. The full stack is a firm-owned data spine, a firm-specific risk taxonomy, an AI triage and narrative model trained on your own outcomes, a human sign-off workspace and a full audit trail — typically shipped in a 6–12 week initial build.

Short answer

You build an in-house AI KYC system in 2026 by treating identity verification vendors (Onfido, Sumsub, Jumio, Veriff) as interchangeable data sources and owning the risk-decision layer above them. The full stack is a firm-owned data spine, a firm-specific risk taxonomy, an AI triage and narrative model trained on your own outcomes, a human sign-off workspace and a full audit trail — typically shipped in a 6–12 week initial build.

Why "in-house AI KYC" is not "build our own IDV"

The confusion most firms fall into is thinking "in-house AI KYC" means building document scanning and biometric checks from scratch. It doesn't, and it shouldn't. IDV vendors (Onfido/Entrust IDV, Sumsub, Jumio, Veriff) are technically credible and improve constantly. Replicating them buys nothing.

The right frame: those vendors return a pass, fail or manual-review flag. What lives above the check — the risk narrative, the case triage, the correlation with fraud signals, the SAR narrative drafting, the audit trail — is where firm-owned intelligence produces durable value. That is the in-house AI KYC system.

The architecture, layer by layer

  1. Source layer — IDV vendor (or two: primary + secondary), device intelligence (Sift, Sardine), PEP/sanctions screening (ComplyAdvantage, WorldCheck, Dow Jones Risk), core banking / product data. All ingested as data sources.
  2. Firm data spine — Warehouse (Snowflake / BigQuery / Databricks) plus vector store for documents. Every verification, every fraud signal, every historical decision and every downstream outcome (chargeback, closure, fraud loss) streams in.
  3. Risk taxonomy — Firm-specific categories mapped from vendor taxonomies. Product-specific concentration limits, region-specific rules, customer-segment-specific policies. Owned in code, versioned like code.
  4. Model layer — LLM-based triage of ambiguous cases, narrative agent that drafts the customer-risk story, cross-source fraud correlation, outcome-trained risk scoring. Trained on the firm's own historical decisions and outcomes.
  5. Human sign-off workspace — Analyst cockpit with drafted narrative, evidence linked to source documents, recommended action. Every decision human-signed with explicit rationale capture.
  6. Audit trail — Full provenance from every AI proposal back to source data. Every human override captured. Immutable log, regulator-ready.

The build sequence in 6–12 weeks

The pattern in production: weeks 1–2, integrate the primary IDV vendor's callbacks into the firm spine and stand up the warehouse and vector store. Weeks 3–5, encode the firm risk taxonomy and ship the first AI triage model on a defined customer segment. Weeks 6–8, wire in the analyst cockpit with narrative drafting and human sign-off. Weeks 9–12, add cross-source fraud correlation, outcome-trained scoring and the regulator-ready audit trail.

The critical mistake is trying to do everything for every customer segment on day one. Ship the intelligence layer for the highest-volume, highest-cost segment first, prove the ROI, then extend segment by segment.

Regulator posture: how in-house AI KYC strengthens it

The instinct in some compliance functions is that AI weakens regulator posture. In practice, a properly built in-house AI KYC system materially strengthens it, for three reasons.

First, provenance. Every AI proposal links back to source documents, every human decision is captured with rationale, every model version is logged. The audit trail is stronger than any manual process it replaces.

Second, consistency. A firm-owned model applies the firm's own risk taxonomy to every case in the same way. Human-only KYC teams are inherently inconsistent across shifts, offices and time; a firm-owned model reduces that variance and makes remaining exceptions explicit.

Third, defensibility. When a regulator asks how a decision was reached, "we applied our own documented risk taxonomy, using firm-owned code, with full provenance and human sign-off" is a stronger answer than "we followed the vendor's recommendation".

In-house AI KYC build: what actually gets built

LayerWhat it isOwned by
IDV checkDocument + biometric + databaseVendor (Onfido / Sumsub / Jumio / Veriff)
ScreeningSanctions / PEP / adverse mediaVendor (ComplyAdvantage / WorldCheck / Dow Jones)
Data spineWarehouse + vector storeFirm
Risk taxonomyFirm risk categories + rulesFirm (versioned in code)
Model layerAI triage + narrative + correlationFirm (trained on own outcomes)
Analyst cockpitSign-off workspaceFirm
Audit trailImmutable provenance logFirm

Vendors stay for the commodity check; the firm owns the intelligence and the decision.

How to build this in production

  1. 01

    Integrate the IDV vendor(s) into the firm data spine

    Every verification result, callback and manual-review event streams into a governed firm-owned warehouse alongside device intelligence, screening hits and product data.

  2. 02

    Encode the firm risk taxonomy in code

    Firm-specific risk categories, product-specific limits, region-specific rules — owned in code, versioned like code, mapped from vendor taxonomies.

  3. 03

    Ship the AI triage + narrative model on the priority segment

    Trained on the firm's own historical decisions and outcomes. Writes drafted narrative and recommended action into the analyst workspace.

  4. 04

    Build the analyst cockpit and human sign-off workflow

    Every decision captured with rationale, every AI proposal linked to source documents, every human override becomes training data.

  5. 05

    Wire in cross-source correlation, outcome-training and audit trail

    Correlate IDV, device intel and screening per customer; retrain risk model on downstream outcomes; ship the immutable regulator-ready audit log.

Related questions

Should we build our own IDV instead of using Onfido / Sumsub?

No. IDV vendors are technically credible and improve constantly. Replicating them buys nothing. Where firms produce durable value is in the intelligence layer above the check — risk narrative, triage, correlation, audit trail — which is firm-owned.

How is this regulator-friendly?

Full provenance from every AI proposal back to source data, every human decision captured with rationale, every model version logged. The audit trail is stronger than any manual process it replaces, and "we applied our own documented risk taxonomy" is a defensible position with any serious regulator.

How much does an in-house AI KYC system cost to build?

A production first-segment deployment typically ships in a 6–12 week Build engagement at £150k–£400k, with ongoing infra costs of £15k–£40k/month depending on volume. The break-even against manual review reduction is usually inside 12–18 months.

Can this replace our L1/L2 analyst team?

It shouldn't and doesn't. It reduces the number of cases each analyst has to work by 30–60%, and materially improves the quality of the ones they do work by pre-drafting narrative and evidence. Analysts are freed for higher-value investigations and complex cases.

Which regulators is this pattern compatible with?

The pattern is compatible with FCA (UK), FinCEN (US), MAS (Singapore), FINMA (Switzerland), MiCA/EBA (EU) and every major AML regime we've reviewed. The audit trail and firm-owned taxonomy typically exceed what regulators see from most vendor-only deployments.

Firms that build an in-house AI KYC system in 2026 own their onboarding conversion, their fraud economics and their regulator posture. Firms that don't are betting their compliance function on a vendor's roadmap.

Start with the £15,000 Financial AI Blueprint — two weeks with our team and you have a board-ready plan, sized to your customer base, data quality and starting posture.