How the measurement works
Why the test is short, and why it stops when it does.
It measures a decision, not a score
For every competency, the diagnostic is answering one question: do you need teaching here or not. That is a much smaller question than “how good are you”, and it can be settled far faster.
So the test stops asking about a competency as soon as you are clearly on one side of that line. Someone comfortably above it is done in about three questions. Someone sitting almost exactly on it might get six, because that is the case where the extra evidence actually changes what happens next.
The first version of this used a fixed precision target instead, and it could never have worked: the target was mathematically unreachable inside the question limit, so every competency would have run to its cap and the early stop would have been decoration. A test caught it. The deeper mistake was measuring the wrong quantity.
Each question is chosen to be maximally informative
Ability and item difficulty live on the same scale, so the question that tells us most is the one you have roughly an even chance of getting right. That is where an answer carries the most information, and it is also why a well-built adaptive test feels challenging rather than either humiliating or pointless.
The model behind it is deliberately simple — one ability parameter per person per competency, one difficulty parameter per item. Richer models need thousands of responses per item before their extra parameters are anything but noise, and a new bank does not have that. It gets replaced when the evidence justifies it, not before.
Uncertainty is reported, never assumed away
If the diagnostic never got to a competency, it is reported as unassessed rather than quietly counted as passed. Silence is not evidence. The same rule governs the credential: it states what was assessed and what was not.
Where there is genuine doubt, you get taught. Being shown a lesson on something you already knew costs a few minutes. Being waved past a real gap costs you the exam.
Items are written here, and reviewed by a person
Every question is original. Exam boards own the copyright in their papers, and reusing them commercially is not fair use, so nothing here is taken or adapted from a real test.
Items are drafted by a language model, then put through an automatic screen and then in front of a human, who approves or rejects each one by name. An item nobody approved is never shown to anyone. Once items are live, their statistics are watched: a question the strongest candidates get wrong is almost always a question with the wrong answer keyed, and it is pulled automatically.
What is not covered is stated
Speaking is not assessed. An item bank cannot rate live speech, and a credential that quietly omitted a quarter of an exam while implying full coverage would make every other credential worth less. It either becomes a paid add-on with a human rater or it stays out. It does not become a machine-scored approximation.