Steel Vision Ltd · Edinburgh · AI Assurance

Independent evaluation for AI in environments where failure is expensive.

Before a regulator asks, a risk committee signs, or a system touches production — someone has to prove the AI actually works. That is what we do: independent, evidence-first evaluation of deployed and edge AI systems, delivered as fixed-price engagements.

Why our evaluations hold up

We build what we judge.

Our founder ships production AI — including a safety-critical, offline-first Edge AI system for industrial field crews. Evaluation grounded in engineering, not slideware.

Validated where money is real.

Our methodology runs daily on a live, real-money machine-learning trading system: combinatorially purged cross-validation, sealed holdouts, drift monitoring, and fail-closed deployment gates.

Published on evaluation integrity.

Four papers on machine-learning evaluation — including how models game the metrics used to judge them. We test for the failure modes we documented.

Evaluation at the edge.

We evaluate AI where it actually runs — cloud, on-premise, or fully on-device. Air-gapped and offline systems are our native territory, not an exclusion clause.

The Engagement

The Two-Week Evaluation Sprint

Fixed price. Fixed scope. Two weeks from kickoff to an evidence pack your risk committee can act on.

What we test

  • fact_check
    Reliability & hallucination

    Does the system state falsehoods with confidence — and under what conditions?

  • security
    Robustness & gaming

    Adversarial probing: prompt manipulation, metric gaming, and behaviour at the edges of the training distribution.

  • trending_down
    Regression & drift

    Behaviour versus baseline across versions — what silently changed, and whether monitoring would catch it.

  • gavel
    Governance evidence

    The documentation trail regulators expect: what was tested, how, by whom, and with what result.

What you receive

A complete evidence pack: findings ranked by severity, the full test methodology, reproducible harness outputs, and governance documentation written for the committee that has to sign — not for the engineers who already believe.

  • Findings report with severity ranking
  • Test methodology & coverage map
  • Reproducible evaluation harness outputs
  • Risk-committee-ready governance pack

The Product

The Steel.vision Evaluation Platform

In development — hardened by every engagement we deliver.

Every sprint runs on our evaluation harness — the same methodology validated daily on a live, real-money machine-learning system and documented in our founder's published research. We are productizing that harness into a platform: automated pre-deployment testing, continuous drift monitoring, and evidence-pack generation for AI deployed in regulated sectors. Early engagements are how the platform meets its first users — each one ships product, not slides.

Early access — request an evaluation below.

Where we work

account_balance

Financial Services

Independent validation for AI and ML models under model-risk expectations — the evidence line your second line of defence and regulators expect to see.

bolt

Energy & Industrial

Evaluation for AI operating in safety-critical, connectivity-constrained environments — including fully offline and on-device systems.

domain

Public Sector

Procurement-ready, transparent evaluation for AI in public services — documented to the standard public accountability demands.

precision_manufacturing

Proof we build what we evaluate

Steel.vision began by building a multimodal Edge AI system for industrial field technicians — computer vision, retrieval, and language models running entirely on-device in environments with no connectivity and no second chances. That system is our standing case study in evaluation-grade engineering: every claim on this page is a discipline we first imposed on ourselves. Meet the founder →

Start a conversation

Request an evaluation

Tell us what your AI system does and where it runs. We reply within one working day — usually with three questions and a proposed scope.