What alternative data actually improves SME underwriting?
The alternative data sources that materially move SME underwriting in 2026 are: Companies House filings and director history, VAT return data via HMRC, marketplace revenue (Amazon/eBay/Shopify), payment processor data (Stripe/Square), and device/behavioural signals. Everything else is noise or unavailable at scale.
The alternative data sources that materially move SME underwriting in 2026 are: Companies House filings and director history, VAT return data via HMRC, marketplace revenue (Amazon/eBay/Shopify), payment processor data (Stripe/Square), and device/behavioural signals. Everything else is noise or unavailable at scale.
The signal-to-noise problem
Every SME lender has been offered 'alt-data' feeds — social scraping, satellite imagery, news sentiment, sector proxies. Almost none of it moves the underwriting model. The sources that do are unglamorous, structured and boringly reliable.
The five alt-data sources that actually move the score
- Companies House — director history, filing regularity, group structure, charges.
- VAT return data — cash-basis and standard-scheme returns, per-quarter progression.
- Marketplace revenue — Amazon Seller, Shopify, eBay APIs where borrower consents.
- Payment processor data — Stripe, Square, SumUp where borrower consents.
- Device / behavioural — session behaviour on application, device fingerprint, referrer patterns.
How to blend without over-fitting
The trap is loading every source into a monolithic model and celebrating in-sample Gini. Modern lenders use a two-stage blend: a per-source sub-score with strong monotonic constraints, then a shallow blender. This is explainable, robust to source drops and survives regulator scrutiny.
Related questions
Which source has highest single-source lift?
Open banking, by a distance. Of the alt-data sources listed here, Companies House and VAT returns share second place.
Consent to marketplace / processor?
First-class UX flow at application; SMEs consent in 60-80% of cases when the reason is clear.
Data quality issues?
Every source has known failure modes documented per-feature; monotonic constraints protect the model when a source drops.
Cost?
Companies House and VAT are near-free at scale. Marketplace/processor via customer consent. Device intelligence is per-application priced.
Timeline?
Eight to twelve weeks to bring 3-5 sources into the ensemble after the Diagnostic.
The AI-first lenders are pulling away on unit economics, not on rate. Every quarter you defer the architecture is a quarter of compounding disadvantage.
The £15k AI Diagnostic maps your lender stack, prioritises the systems that pay back fastest and produces a costed sequenced build plan.