Steel Vision Ltd · Edinburgh · AI Assurance
Independent evaluation for AI in environments where failure is expensive.
Before a regulator asks, a risk committee signs, or a system touches production — someone has to prove the AI actually works. That is what we do: independent, evidence-first evaluation of deployed and edge AI systems, delivered as fixed-price engagements.
Why our evaluations hold up
We build what we judge.
Our founder ships production AI — including a safety-critical, offline-first Edge AI system for industrial field crews. Evaluation grounded in engineering, not slideware.
Validated where money is real.
Our methodology runs daily on a live, real-money machine-learning trading system: combinatorially purged cross-validation, sealed holdouts, drift monitoring, and fail-closed deployment gates.
Published on evaluation integrity.
Four papers on machine-learning evaluation — including how models game the metrics used to judge them. We test for the failure modes we documented.
Evaluation at the edge.
We evaluate AI where it actually runs — cloud, on-premise, or fully on-device. Air-gapped and offline systems are our native territory, not an exclusion clause.
The Engagement
The Two-Week Evaluation Sprint
Fixed price. Fixed scope. Two weeks from kickoff to an evidence pack your risk committee can act on.
What we test
- fact_checkReliability & hallucination
Does the system state falsehoods with confidence — and under what conditions?
- securityRobustness & gaming
Adversarial probing: prompt manipulation, metric gaming, and behaviour at the edges of the training distribution.
- trending_downRegression & drift
Behaviour versus baseline across versions — what silently changed, and whether monitoring would catch it.
- gavelGovernance evidence
The documentation trail regulators expect: what was tested, how, by whom, and with what result.
What you receive
A complete evidence pack: findings ranked by severity, the full test methodology, reproducible harness outputs, and governance documentation written for the committee that has to sign — not for the engineers who already believe.
- Findings report with severity ranking
- Test methodology & coverage map
- Reproducible evaluation harness outputs
- Risk-committee-ready governance pack
The Product
The Steel.vision Evaluation Platform
In development — hardened by every engagement we deliver.
Every sprint runs on our evaluation harness — the same methodology validated daily on a live, real-money machine-learning system and documented in our founder's published research. We are productizing that harness into a platform: automated pre-deployment testing, continuous drift monitoring, and evidence-pack generation for AI deployed in regulated sectors. Early engagements are how the platform meets its first users — each one ships product, not slides.
Early access — request an evaluation below.
Where we work
Financial Services
Independent validation for AI and ML models under model-risk expectations — the evidence line your second line of defence and regulators expect to see.
Energy & Industrial
Evaluation for AI operating in safety-critical, connectivity-constrained environments — including fully offline and on-device systems.
Public Sector
Procurement-ready, transparent evaluation for AI in public services — documented to the standard public accountability demands.
Proof we build what we evaluate
Steel.vision began by building a multimodal Edge AI system for industrial field technicians — computer vision, retrieval, and language models running entirely on-device in environments with no connectivity and no second chances. That system is our standing case study in evaluation-grade engineering: every claim on this page is a discipline we first imposed on ourselves. Meet the founder →
Start a conversation
Request an evaluation
Tell us what your AI system does and where it runs. We reply within one working day — usually with three questions and a proposed scope.