The AI decision your people make today, made here first.
Real work with an AI system inside it. Every decision is scored. Every consequence plays out. Nobody is told what to look for.
Work Queue
● LIVETK-4082
Complaint — Account closure
TK-4083
Payment authorisation — Voice
TK-4084
Vendor pack — Renewal review
AI Decision Score
84 / 100
Tracks
Pick the sandbox your people work in.
Where AI touches money, customers and regulators
Complaints, client notes, payment instructions, code changes and vendor renewals — with an AI tool already inside the workflow and a governance call nobody flags.
Where the agents themselves get built
Model routing, context and caching, tool permissions, untrusted input and data access. The engineers building your agents set your AI bill and your blast radius in code, long before anyone reviews it.
Where AI touches patients and clinical records
AI scribes, patient messages, case summaries and regulatory submissions — where an unverified AI output reaches a patient record or a regulator.
Where AI speaks to your customers
Service agents, product content, pricing and loyalty data — where an AI promise, claim or data export reaches customers at scale, and the brand is held to it.
How a simulation runs
Three screens. One decision each.
Recognition is the first screen. Handling is the other two.
01
The task
Ordinary work with a deadline. Nothing is flagged. The participant opens what they choose to open, takes a real action, and records why.
02
The consequence
The decision plays out as a system event, not as feedback text. Then they have to deal with it.
03
Fix and prevent
What gets contained, who gets told, and what standing control stops it happening again.
A wrong first call caps the score. It doesn't end the simulation — handling your own mistake correctly is most of the value.
Why it feels different
Not a course. A working environment.
Real work, not exercises
The queue and the repo are full of legitimate items built to the same quality as the one that matters. Finding it is the work.
Evidence, or the score says so
Every artefact is a real artefact. Whether the critical one was opened before deciding is recorded.
Consequences, then containment
A wrong call becomes the incident the participant has to contain and report.
Standards, not reminders
Each simulation ends in one control statement the organisation adopts, tested more than once across the library.
For leadership
What you hold at the end.
Not an attendance sheet.
A control standard set in force — written so your audit or platform function can test against it
Readiness by function and category — where the team is strong and where it is thin
Every decision with its reasoning — captured where it arose, exportable
A named gap list — the specific categories and cohorts to fix next
A reconciled, tamper-evident record — scores cross-checked, discrepancies flagged
This measures human decision capability. It is one indicator. It is not a measure of your organisation's overall AI risk posture.
Run one team through it.
Pick the queue or the repo where AI already touches a decision, and see what your record looks like.