BG Assurance
Enterprise AI Assurance

Your people are using AI. Can you prove they're using it well?

BG Assurance sits between your teams and the models they reach for — measuring every prompt for quality, cost, and risk before the work goes out, and turning it into evidence you can govern.

63+ models compared Task-level routing Accuracy before price Human oversight where it counts
Portfolio assurance score LIVE
0 · UNGOVERNED100 · ASSURED
Prompts / wk
14,208
Cost saved
£61k /mo
Flagged
3.1%

The gap

Enterprises bought the models.
Almost nobody instrumented the use.

95%

of enterprise generative-AI pilots showed no measurable return. The failure was integration and oversight — not the models themselves.

MIT NANDA, via Fortune · 2025
$30–40bn

spent on enterprise generative AI to date — much of it with no reliable way to verify how it's actually being used day to day.

MIT NANDA, via Fortune · 2025
11.5×

the cost of putting a frontier model on every task — for roughly 5% more quality the task usually didn't need. Routing routine work the right way removes up to 85% of that spend.

Illustrative · published 2026 model-cost analyses (CloudZero)

Assurance is the missing layer: not another model — the measurement, guidance, and evidence that make the models you already have accountable.


What we do

An assurance layer, not another chatbot.

Your teams reach the leading models through us. Every prompt passes a real-time assurance layer — deterministic rules for the obvious, a purpose-built model for judgment — before an answer is ever generated. The result is better work, lower spend, and a defensible record of how AI is being used across the organisation.

The practice

Independent AI audit & assurance

An advisory practice grounded in documented AI failure cases at major firms — reviewing where models are trusted, how they're governed, and what would stand up to scrutiny.

The platform

Real-time, per-prompt assurance

A gateway to 63+ models with a live scoring layer on top. Individual analytics for every user; organisation-wide analytics and per-person drill-down for leaders.


The assurance layer

What happens the moment a prompt is sent.

Seven checks, in real time, before and after the model answers — so quality, cost, and risk are handled on the way out, not discovered after the fact.

01

Prompt written

A user writes a prompt in the BG Assurance workspace.

02

Rules & safety gate

Deterministic checks read structure and scan for sensitive or restricted content. Unsafe prompts held; risky ones flagged to anonymise.

03

Model routing check

Does this task need the model chosen — or would a smaller, cheaper one do it just as well?

04

Skills & tools check

Should this have used retrieval, code execution, or a connected tool? Mismatches that quietly degrade output are surfaced.

05

Quality review

A purpose-built model judges the softer questions: is the task clear, the context enough, the framing right?

06

Risk scoring

Hallucination, security, and confidentiality risk are scored. High-risk responses route to human review.

07

Score & evidence

One clear score and two or three practical fixes — logged to the user's analytics and the organisation's.


What we measure

Seven lenses on every interaction.

Each interaction is scored the same way, every time — so a habit that's quietly costing money or creating risk shows up as a number, not a hunch.

01

Prompting & context

Vague instructions, missing context, and the wrong framing are where costly, hard-to-spot errors begin. We score each prompt and coach the fix.

Prompt score
Clear taskBusiness contextSpecific instructionsOutput formatReliability guidanceSafety & compliance
02

Model selection

Not just which provider — which specific model, and whether it's the most accurate and cost-effective choice for this task. Grounded in our comparison of 63+ models.

Fit / task
Cost to run a comparable task · smaller ≈ 1× · mid-tier ≈ 3× · frontier ≈ 11.5×
03

Skills, tools & plugins

Whether people pair the right capabilities — retrieval, code execution, connected tools — with the right models, and where a mismatch is degrading output or adding risk.

Tooling fit
04

Reliance balance

Where AI is over-trusted for work that needs a human — and where it's underused, leaving productivity on the table. We map both across the business.

Over-reliance
05

Hallucination risk

We can't detect hallucinations directly — so we estimate their likelihood from five measurable indicators and route the riskiest responses to review.

Risk 0–100
Grounded in evidenceAnswer consistencyDomain importanceFactual-claim densityConfidence calibration
06

Security risk

Prompts and responses are checked for injection attempts, data-exfiltration patterns, and unsafe requests before they can reach — or leave — a model.

Exposure
07

Confidentiality risk

Sensitive and client-confidential information — names, financials, credentials, protected data — is detected before it reaches a model, and flagged for removal or anonymisation, with a record of what was caught.

Leak risk

Try it

Run a prompt through the assurance layer.

Type a prompt — or load an example — and see the assessment your people would get back, in real time.

Prompt · draft
Illustrative local assessment0 words

Runs entirely in your browser on an illustrative version of the scoring engine. The production layer also samples model responses for consistency and grounding, and calibrates against audited outcomes for each client.


The evidence

Two views. One record.

The same scoring produces a coaching view for every individual and a governance view for leaders — down to any single person.

Every person sees how they're doing — their assurance score over time, where their prompts are strong, and the two habits that would improve them most.

Assurance score · 12 weeksA. Okafor · Legal
Currentvs team avg 71
84 ▲ 12
Top quartile in your function. Strongest on output format and reliability guidance.
Your two highest-leverage habitscoaching · this month
Add contextYou name the deliverable clearly, but 4 in 10 of your prompts don't say who the output is for. Adding the audience lifts your average prompt score by an estimated +1.3.
Right-sizeYou default to a frontier model for routine drafting. Routing those to a mid-tier model would cut your inference cost by an estimated 54% with no measured quality loss.

Leaders see the whole estate — an organisation-wide assurance index, the cost routing has saved, where risk is concentrated by team, and the ability to drill into any individual.

Org assurance indexrolling 30d
78 ▲ 9
Up from 69 at rollout. 1,240 active users across 9 functions.
Inference cost savedrouting · /mo
£61k
37% of spend removed by routing routine work to right-sized models. Illustrative.
Responses to reviewrisk ≥ 65
3.1%
Auto-routed to human review. Concentrated in legal and finance.
Risk by team & lensdarker = higher risk
Adoption vs relianceby function
Drill into individualslegal function · sample
PersonAssurancePrompt qual.Halluc. riskFlag rateModel fit

Deployment

Platform first. Coverage next.

We launch where value is proven fastest — the workspace your teams use directly — then extend enforcement outward, on one governance model.

Phase 1 · now

Platform

The workspace: prompt checking, clear feedback, editing, and appeals. Prove the core value with real usage.

Live focus
Phase 2

Team controls

Admin controls, policies, reporting, and audit logs for managers and risk functions.

Phase 3

Hybrid

Extend coverage to AI used outside the platform, keeping one appeals and governance process.

Phase 4

Enterprise network

Network-level protection across tools for the largest enterprises, with the same record of use.

Every prompt, measured.
Every model, accountable.

BG Assurance turns AI from something you hope is going well into something you can see, govern, and prove.