How it works · The test

A test you can stand behind

Tests are coming soon to GoYou. How we do it and where we draw the line is already here, so you can explain it to your department, your exam board and your students.

This is how a test is built.

A test is not a list of questions a language model coughs up. It is a chain with a fixed order, in which every step checks something the previous one could not, and in which the last step is always you.

  1. Step 1

    Only what your class has covered

    A test comes from your lessons, a series, or text and learning goals you supply yourself. Every question points to the piece of material it rests on. A question without that reference does not exist.

  2. Step 2

    The structure first, then the questions

    You see, per learning goal, how many questions there will be, at which thinking level and for how many marks. That is the test blueprint. For a test that counts, you approve it before a single question is written.

  3. Step 3

    The numbers are worked out by code, not by the AI

    How many questions and marks a learning goal gets follows fixed rules: every goal is covered and no single goal takes up more than forty percent of the test.

  4. Step 4

    A rule that code can see is enforced in code

    No “all of the above”, no double negatives, no correct answer that happens to be the longest, a key for every closed question, sentences that fit the reading level of your class. What fails goes back to the writer once, with the reason attached.

  5. Step 5

    A second reader who is told nothing

    A different AI model reads every question without seeing which thinking level was intended, and re-checks the key. If it says “this is recall”, it must quote the sentence in the material where the answer sits, and the code verifies that quote. A question below its level is rewritten once.

  6. Step 6

    You approve, and nothing reaches students before that

    You see the questions as your students will see them, with the key, the learning goal, the material and every signal we could not settle ourselves. You edit, discard, add. A colleague can check the whole test on one printable review sheet.

Two kinds

Formative or summative is a different question from scored or not.

A formative test shows where your class stands and may well have a score. A summative test counts and is stricter: the structure is a separate stop that you approve, questions that appeared word for word in the lesson are not allowed (unless you switch that on, for vocabulary or formulas), and the form is a form. Alongside these comes the targeted test: only what is still shaky in this class, with the reason next to each question.

What we do not claim.

A test is only good if it measures what it is meant to measure. We are honest about that, also where it does not suit us.

That a question is good because AI approved it

Our AI readers are a sieve, not a seal of quality. They catch giveaways, overlap and questions outside the material. A subtle ambiguity can slip through. That is why every signal sits next to the question, and why the review sheet for a colleague exists.

That we know how hard a question is

That only shows once a class has taken the test. The screen says so in as many words. After a sitting we work out how each question behaved, and only then is there anything to say about difficulty.

That a test from one lesson can carry a grade

Below 25 marks GoYou warns you: one guess or one slip then shifts the outcome. A graded test built from a single lesson we call a quiz, and we say so.

That the thinking level is right because it says so

A question labelled “apply” whose answer sits word for word in the lesson is a recall question. This was the biggest weakness we measured in ourselves, and the reason for the second reader with a duty of proof.

Our testing policy

What we do, and what we never do.

No surveillance, ever

No AI detector, no proctoring, no webcam, no emotion recognition. What we do record (that an answer was pasted, that a screen was left) is a fact, deterministic, and the student knows it before starting.

No student data goes to the model

To build a test, only the lesson material and the education level go to the model. No name, no class, no school, no town.

A test that counts is a form

Every student gets the same questions, the same order of parts, the same time. An AI that keeps probing during a graded test makes results incomparable, so it does not. In formative work the tutor may join in.

The pass mark is a line of reasoning, not a habit

GoYou shows what guessing alone would yield and what a pass mark therefore really means. Per learning goal you can mark what every student who just passes must be able to do; the pass mark follows from that and can be explained to a parent or an exam board.

Marks are earned per part

Four out of five placed correctly is not zero. Every question form has a scoring rule in plain language that sits next to the question, and that a student can check for themselves. An ordering question is corrected for chance.

A grade is never changed by a machine

GoYou proposes, you approve. After release a student can ask the tutor why something was wrong; the score never changes as a result. Anyone who disagrees ends up with you. We keep no grade book.

Built as if the strictest rules already apply

The European AI Act counts assessment in education among the high-risk uses. We build as if those rules already apply: a human who can stop every step, a log of every model call, instructions for use and a test file with every choice and every measurement.

We do not do examinations

School exams and national exams fall outside GoYou. We do already record what an exam board would later ask for: a version number, freezing per test, complete logging.

How we check ourselves.

A rule in an instruction to the AI only counts with us once a measurement shows that it works. So we run a fixed measurement set: ten fixed sources, from upper primary to pre-university and vocational, the same every time, with thresholds fixed in advance. If the set is red, it stays red until things are genuinely better; we do not talk it green.

What that set taught us in September 2026: too many questions asked for a lower thinking level than their label said. A question was called “apply” while the answer sat word for word in the lesson. So an independent model now reads every question, must prove “recall” with a quote from the material, and a question below its level is rewritten. Whatever still reads too low carries a visible signal for the teacher. This is not solved, which is why it is here.

What we cannot measure ourselves, we have others do. A fixed set of generated questions is ready for review by an independent assessment expert, blind: without our intended level and without our signals. How many rejected questions had slipped past us without a signal is the figure that counts. After a real sitting the classical measures follow: how hard a question turned out to be, how well it discriminates, and how reliable the test is as a whole.

The full account, with every choice, what was rejected and every measurement, is in our test file. We share it with schools that want to read it; ask viacontact.

For those who want to go deeper

What this rests on.

None of this is new; it is what assessment theory has said for decades, put into software. A test that follows from the learning goals and the lessons in between (constructive alignment). A test blueprint before the questions. Thinking levels after the revised Bloom taxonomy, measured against what was covered and not against the verb in the question. Correction for guessing on guessable forms. A pass mark that follows from the content rather than from habit. Partial credit because all-or-nothing throws information away. And after a sitting: p-values, discrimination and reliability, the measures a Dutch assessment expert asks for.

Another question about tests?

Ask GoYou’s assistant: what a test does and does not do, how the pass mark works, and what happens to students’ answers.