The vendor scorecard: 40 AI-for-finance vendors, ranked by what actually ships.
Over the last two years KJ Capital has sat inside evaluations of more than forty AI vendors selling into financial services. Most of them are the same wrapper with different branding. A handful are serious. Here is how we separate the two in fifteen minutes.

One of the most common questions I get from finance-firm CTOs and COOs is some version of: how do we tell which of these AI vendors is real? The decks look identical. The demos show the same three use cases. The pricing pages have been through the same consultant. And the AI-labelled procurement pipeline has quadrupled in eighteen months.
We have now sat through evaluations of more than forty vendors — some as an advisor, some as an architect brought in to make the buy/build call, occasionally as a build partner brought in after a vendor call went sideways. The honest scorecard is uncomfortable. Most of them are the same OpenAI or Anthropic wrapper with a different logo, a nice front-end, and a two-page white paper about financial-services-grade AI that means nothing.
This note is the fifteen-minute filter we now run, the four questions that separate serious vendors from repackaged wrappers, and the shape of the eight or nine vendors that pass. Names are withheld — we still work with several of them — but the shape of the scorecard is public.
Question 1 — show me the evals harness
Not the demo. The actual harness. The test set, the scoring rubric, the version log of runs against it. A vendor without one is not a serious AI vendor. They are a design agency with an API key. The test set should be domain-specific — if they cannot show one, the domain claim is marketing.
Serious vendors will hand you a document, or better a live dashboard, showing every model version they have shipped, scored against a fixed rubric, with the deltas explained. The unserious ones will describe a QA process. Those are different things. QA measures whether the button works. Evals measure whether the answer is right.
Question 2 — what happens when the model gets it wrong
The right answer is a specific mechanism. Human-in-the-loop routing above a confidence threshold. Fallback behaviour to a deterministic system. Incident logging and post-mortem process. Named severity levels. A shrug and a claim about model accuracy is the wrong answer — that vendor has never operated the system at real volume.
This question filters brutally. Most vendors give an aspirational answer because they have not yet operated a customer whose failure is expensive. If you are a regulated financial firm, you cannot be their first.
Question 3 — show me the compliance envelope
Which data goes where. Which model provider sees what. How the vendor handles right-to-erasure, data residency, model-version transparency. How they evidence outputs to a regulator, including yours. If they hand you a security questionnaire and a SOC 2 report instead of an architecture diagram, you have your answer.
Serious vendors have a per-customer compliance envelope. The good ones will let you inspect it. The best ones will let your CCO reshape it before signing.
Question 4 — tell me about the last customer you turned down
This one is the killer. Serious AI vendors turn firms away — because the use case is not ready, because the data is not there, because the accuracy bar cannot be cleared, because the regulator posture is wrong. Vendors that say yes to everyone are selling revenue, not systems.
The answer to this question is a proxy for whether the vendor has an engineering culture or a sales culture. Both can be legitimate businesses. Only one of them should be embedded inside your customer-facing stack.
The scorecard, aggregated
Across the forty-vendor cohort, the aggregate results were:
- Passed all four questions: 9 vendors. Genuinely worth talking to. Six of the nine are UK or EU based; three are US.
- Passed three of four: 7 vendors. Usually strong in one dimension (evals, or compliance) and weak in another. Fine for narrow, non-customer-facing use cases.
- Passed two of four: 11 vendors. Mostly wrappers with a nice front-end and one genuine capability. Bad fit for regulated deployments.
- Passed one of four: 9 vendors. Selling AI-flavoured software. Skip unless the use case is trivial.
- Passed zero: 4 vendors. In several cases, actively dangerous — no compliance envelope, no evals, no failure mechanism. One is on our do-not-recommend list.
When to buy and when to build
Buy when the use case is horizontal, non-customer-facing, and someone else is going to invest more in the roadmap than you can. Meeting-note summarisers, internal knowledge-base search, developer copilots — all buy. The vendor’s advantage compounds and you are unlikely to build a better version.
Build when the use case is customer-facing, sits inside your compliance envelope, or is where you intend to compete. Research substrates, KYC copilots, trader onboarding, personalised communications — all build. A shared vendor cannot be your differentiator. If it is shared, every competitor has it.
Hybrid — buy the substrate, build the surface — is the most common shape we ship for firms in the £5m–£100m revenue band. Buy the model, buy the vector database, buy the observability. Build the policy layer, the retrieval curation, the customer surface, the evals rubric. That combination captures 80% of the benefit at 30% of the build cost.
The three vendor categories most firms get wrong
First, the compliance-labelled RAG vendors. Two-thirds of the compliance-grade vendors in the cohort do not actually enforce policy above the model — they just log. Logging is not a control. If the model can produce an out-of-policy output and the vendor’s only response is to record it, that is not compliance-grade.
Second, the horizontal AI platforms with a financial-services skin. The skin is real; the domain depth is usually not. Three months in, the firm is doing all the domain adaptation itself, at platform prices, and the vendor is not helping.
Third, the AI-driven trading tools that are actually decision-support tools. This is a definitional problem more than a technology problem — but many firms buy them thinking they are getting one thing and get another. Read the contract, not the marketing.
How to run the fifteen-minute test
You do not need an AI architect in the room to run this. You need the four questions, a notepad, and someone willing to stay in the awkward silence when the vendor pivots to a demo. The demo is not the answer to any of the four questions. Do not accept it as one.
If the vendor refuses to answer any of the four questions in writing after the meeting, that is your answer. If they answer three and struggle with the fourth, ask for a technical follow-up with the person who owns that part of the system. If they cannot produce that person, that is also your answer.
Firms that adopt this test consistently save around 70% of the money they were going to spend on AI vendors in the first year, and reallocate it to the ones that actually deliver. That is the real return of the exercise.
FAQ
Will you share the full 40-vendor list?
Not publicly. We work with several of them under NDAs and it would compromise the working relationship. The four-question test is public, and any vendor who wants to demonstrate their answers can do so in a call with your team.
Do you take referral fees from vendors?
No. KJ Capital does not accept referral fees, revenue share, or affiliate arrangements from any vendor. Our economics come only from client engagements, which keeps recommendations honest.
Which categories of vendor should firms most avoid?
Anything sold as compliance-grade AI without an inspectable policy layer above the model. Anything sold as domain-tuned without a domain evals set. Anything sold as an AI trading assistant without a clear line between decision-support and recommendation.
How does this test apply to build partners rather than SaaS vendors?
It works even better. A build partner should be able to answer all four questions from prior projects. If they cannot, they have not shipped enough AI systems into regulated firms to be your first.
Can we run the fifteen-minute test internally?
Yes. That is the point of publishing it. If the firm has a technical enough team to interpret the answers, running it in-house is the right first step. If not, the Financial AI Diagnostic includes a vendor filter as one of its outputs.
Want this rigour applied inside your firm?
Start with the free 5-minute AI Readiness Score, or go straight to the £15k Financial AI Diagnostic — a two-week engagement that produces a costed build plan mapped to your regulator, your stack and your P&L.