AI compliance · advisory
We measure how well your model follows your rules — on which cases, with which numbers — and deliver the documentation that shows it. We start with one bounded measurement.
Bounded · fixed price · no lock-in
The problem
In an inspection it is the numbers that count: how well the AI follows your rules, and on which cases. A measurement makes that visible.
This matters for two reasons. First, you want the rules you set for the system to actually be followed — that was the point of setting them. Second, you need to be able to document it: the EU AI Act requires you to declare your model's accuracy and robustness (Art. 15) and to show continuous risk management (Art. 9).
And systems often follow their own rules worse than people assume. Including the ones built to follow them. You do not see it until it is measured.
Our job is to make it measurable: how well your AI follows your rules, and the paperwork that proves it.
Do not do this
The more rules compete for the model's attention at once, the fewer are actually followed. "Covering yourself" by pasting in the entire handbook is the very thing that makes compliance fall. We fix it with two moves: show only the rules relevant to the question, and choose the right model size.
We measured it on our own service søgfonde.dk: with 56 rules in play at once, average adherence fell to 33 %, and 32 of the 52 relevant rules failed — several all the way down to 0 %. When the system instead looks up only the rule relevant to the question, adherence rises towards 100 %.
The EU AI Act is in force and the requirements for high-risk uses are phasing in. The question is not whether you will need to show the numbers, but when an authority asks for them. That is what we deliver.
From the article “large or small model?” (in Danish)
How we work
Every recommendation comes with a visual explanation you can use and verify. One example: a model that is too small cannot solve the task safely, and a larger one quickly becomes unnecessary — it cannot run on your own servers, and it costs more than the value it adds. The right choice is the smallest model that handles your task with good margin.
What you get
Concrete deliverables. You know what you are paying for the whole way, and you can prove it afterwards.
We find out how advanced a model your task actually requires, so you neither overpay nor pick too little.
We tune the model to your risk profile, so it is neither too loose nor refuses too much.
When the model is unsure it refers on instead of inventing an answer. That is built into the system.
Documentation showing your auditor that you meet the AI Act's requirements on accuracy (Art. 15) and risk management (Art. 9).
The smallest model that solves the task safely. Cheaper, can run on your own servers, and uses less energy.
The method works on the model you prefer, and we install the risk controls in your own setup. The safety net is not universal: change the base model and we recalibrate — a new model does not inherit the old one's protection.
See also Friktionskompasset — find the friction in your own everyday work →
| Declared metric | Measured | Requirement |
|---|---|---|
| Correct refusal on unsettled cases | 98 % | ≥ 95 % |
| Wrongful refusal of legitimate questions | 4 % | ≤ 8 % |
| Calibration error (ECE) | 0.06 | ≤ 0.10 |
Illustrative figures — not a guarantee. The actual numbers are measured per model, language and corpus, and they vary (on our own EU demo, over-refusal sat higher than this before calibration). That is precisely why we measure yours — every metric on your own cases, dated, ready for the auditor.
How it runs
No large contract to get started. First we measure. Then you know what needs doing, and whether you want us to do it.
A bounded review: how complex the task is, where the floor and the ceiling are, where the risk sits. Fixed scope, fixed price.
You get a concrete recommendation: which model, which safety nets, which settings. Do it yourselves, or let us do it. No lock-in.
Calibration, gate architecture and audit pack installed in your model, auditable from day one.
The whole method is open — you can verify it: the full measurement protocol → · AI Act & GDPR, the short version → (both in Danish)
About
You can verify everything we measure.
Behind the advisory is Tomas Lund. We research how language models make decisions — when they commit to an answer, and when they should have abstained — and turn that research into measurable, documented AI compliance for companies. The research is publicly published as Behavioral Friction Theory — the method behind everything on this site.
The method is public, so you can verify everything we recommend: the measurement protocol and the sources come with it, and the measurement can be run on your own rule-set. You know exactly what you are buying before you buy it.
A bounded analysis with no lock-in. Afterwards you know exactly where your AI compliance stands, and what it takes.
Tell us briefly what you use AI for — we will come back with a bounded measurement offer.