Formative assessment examples, and which ones hold up
Thirty-two techniques, grouped by when they run, and what the evidence actually covers
Black and Wiliam put typical effect sizes for the formative assessment experiments at between 0.4 and 0.7 in a 1998 article for Phi Delta Kappan. Fifteen years on, the peer-reviewed literature still called it the range often cited. Kingston and Nash went looking for the studies underneath it. Of more than 300 K-12 studies, 13 carried enough information to compute an effect size at all, and their weighted mean was 0.20.
That does not make the practice worthless. It makes the headline unusable, and it changes which of the thirty-two techniques below deserve your minutes. They are grouped by when in the lesson they run: before you teach, while you teach, at the door, and across a unit.
What makes an assessment formative
Not the instrument. Black and Wiliam's definition puts the weight on what happens after the answers arrive. Practice is formative to the extent that evidence about student achievement is elicited, interpreted, and used to decide the next steps in instruction.
The reasoning behind that wording is the useful part. Locating it on intention rather than use would mean that evidence collected, but not used still counted as formative, which they call unfortunate. Requiring the adjustment to actually improve learning would be too strict, because that is a claim about a lesson you did not teach.

So the test of every technique below is the same. Did the next thing you did differ? A question you ask and then teach past is a warm-up, and calling it assessment does not change what it did.
The framework Wiliam and Thompson set out, which Black and Wiliam reprint, names five key strategies:
- clarifying and sharing learning intentions and criteria for success
- engineering effective classroom discussions and other learning tasks that elicit evidence of student understanding
- providing feedback that moves learners forward
- activating students as instructional resources for one another
- activating students as the owners of their own learning
Only the second of those is a question on a screen. The first sets up what the question is for, and the last three are what the answers get used for.
One more distinction is worth carrying into the list. Black and Wiliam separate synchronous moments of contingency from asynchronous ones. A synchronous moment is the real-time adjustment inside a discussion. Asynchronous ones include students' own summaries made at the end of a lesson, which they call exit passes. Both are formative, and they pay out on completely different timescales. A technique whose payout is the next lesson cannot rescue this one, which is worth knowing before you pick one.
The number everybody quotes, and what it will support
The 0.4 to 0.7 range is quoted accurately and understood wrongly. Black and Wiliam's sentence describes typical effect sizes across a set of experiments they reviewed. It summarises a spread, not the measured effect of one programme you could adopt.
When somebody did pool the studies into a single number, it came out much smaller. Kingston and Nash's weighted mean across the 13 usable studies was 0.20, with an observed median of 0.25.

Then that number was contested too. McMillan, Venable and Varier went back through the same 13 studies. They argue that weaknesses in how studies were selected, their methodology, and the many different things called formative assessment inside them all moderate the conclusions. Their sharpest point is about definition: a study got in as long as its authors stated that their intervention was formative, and by that rule the pile is too mixed to average.
So the honest position is this. The field's headline number is a range across a mixed set of experiments, and its best counter-number rests on thirteen studies a later team judged unfit to pool. Neither side of that argument tells you whether to run a hinge question on Tuesday.
What helps instead is knowing what size of effect to expect. Kraft assembled 1,942 effect sizes from 747 randomised trials of education interventions and found a median of 0.10 standard deviations. He proposes counting 0.20 and above as large. Effects of 0.15 or even 0.10 should count as large and impressive, he writes, when they come from big pre-registered field experiments measuring broad achievement.
Against that scale, the trial that tested formative assessment as a whole practice reads as a success rather than a disappointment. The Education Endowment Foundation randomised 140 secondary schools and 25,393 pupils into Embedding Formative Assessment. Attainment 8 scores moved by an effect size of 0.10, the equivalent of two additional months' progress, at a very high security rating. That result is significant at the 10% level rather than the 5% one, so it is a modest finding held to a modest bar. The programme cost about £1.20 per pupil per year averaged over three years.
Read the same report's second conclusion before you get excited. It found no evidence that the programme improved English or Maths GCSE attainment specifically. A whole-school programme moved the aggregate and did not show up in the two subjects everybody watches.
Treat any large round number in this field with suspicion, including the flattering ones. Hattie and Timperley put the average effect of feedback across 12 meta-analyses at 0.79, in the top five to ten influences on achievement. Bergeron, a statistician, went through the wider Visible Learning synthesis and reports common language effect sizes running to negative probabilities, an error Norwegian researchers flagged in 2012. He calls failing to notice them an enormous blunder.
And feedback averages hide the direction. Kluger and DeNisi pooled 607 effect sizes across 23,663 observations of feedback interventions. The weighted mean was 0.41, and over 38% of the effects were negative. A negative effect there means the group given feedback finished behind the group given none.
Formative assessment examples for the start of a lesson
Eight techniques that run before you have taught anything, or in the settling minutes at the start. None of them costs more than three minutes.

1. Prequestions. Three questions on material you have not taught yet, answered cold. Ninety seconds. It shows you what the room already believes, including where it is confidently wrong. Teach the answers in the order the wrong answers point at. Pan and Carpenter's review of pre-instruction testing reports Soderstrom and Bjork's undergraduates scoring 8 to 9% better on final exam questions covering content that had been pretested. The same review calls that literature relatively small next to retrieval practice, so keep prequestions in proportion.
2. The entrance ticket. One question from last lesson, on the way in. Two minutes. It measures what survived the gap rather than what survived the hour. If a third of the room has lost it, the first ten minutes of today are already different.
3. The recall dump. Write everything you can remember about last week's topic, no notes. Three minutes. This separates what students can retrieve from what they recognise when they see it, and it is the only one of these eight that puts no limit on the answer.
4. The misconception poll. One question, four options, three of which are the specific wrong answers you expect. Sixty seconds. A plain right-or-wrong count tells you the room is confused. This tells you which confusion, which is the only version you can teach against.
5. One word for the key term. Name today's term and ask for one word they associate with it. Sixty seconds as a word cloud. Repeated words grow, so you see the working vocabulary the room brought in with it. When the biggest word is the everyday near synonym rather than the technical one, teaching the difference is the lesson's first job.
6. Predict the result. Before the demonstration, the experiment or the worked example, everyone writes down an outcome. Thirty seconds. Collect them before you start, because a prediction you can still revise while watching is not one.
7. Show me the first line only. Give a problem, ask for the opening step and nothing else. One minute. The first line is where the choice of method is visible, and a wrong choice wastes everything written after it.
8. The confidence sort. Rate today's success criteria from one to five before the lesson starts. One minute. Treat this as one of the weakest instruments on the page. Hacker and colleagues tracked 96 undergraduates predicting their own test performance across a semester and found the lowest performing students grossly overconfident in both predictions and postdictions.
On the big screen
SLD.FUN/OFFSITE
Where should we host the Q3 offsite?
Formative assessment examples for the middle of a lesson
Ten techniques for the stretch where a lesson goes wrong quietly. The failure Rosenshine describes is the default: less effective teachers simply asked "Are there any questions?" and, hearing none, assumed the material had landed.

9. The hinge question. One question at the pivot of the lesson, with the decision rule written before the lesson starts: this many right and you move on, fewer and you reteach. Sixty seconds. The rule written in advance is the whole technique, because a rule invented after you see the answers is not a rule.
10. All-student response. Mini whiteboards or lettered cards, everyone answering the same question at the same moment. Twenty seconds a question, once the boards and pens are already on the desks. Randolph's meta-analysis of 18 response card studies used hand raising as the control condition and reported large, statistically significant effects on test achievement, quiz achievement, participation and reduced off-task behaviour. Read the boards for the split rather than the total, and reteach against the wrong answer most of them share.
11. The pivot poll. Phones instead of boards, for the two things boards cannot do. Thirty seconds. A board answer is visible to whoever sits either side of it, and an unsure student can read it before writing. And the count is kept, which is what technique 29 needs five weeks later. Boards are quicker off the desk, so save this for the questions whose answers have to outlive the lesson.
12. The process question. Ask how they got the answer. Thirty seconds a student, and you will do three of them in a lesson rather than thirty. Rosenshine's summary is that the most effective teachers also ask students to explain the process they used, and that less successful teachers ask fewer questions and almost no process questions. Correct the method rather than the answer, because a right answer by an unreliable route fails next week.
13. Tell your neighbour. Everyone says their answer to one other person before anybody says it to the room. Forty seconds. The rehearsal is the point, because an answer said once in private is cheaper to say in public. Walk while they talk: a pair with nothing to say is the signal the volunteers will never give you. The full version is think-pair-share, including the phase that breaks and what to do about it.
14. Wait time. Ask, then say nothing for five seconds. Free, and the hardest of these to actually do. Rowe's analysis of over 300 tape recordings put the mean wait before a teacher moved on at one second, which is the habit the five seconds is fighting. The think-pair-share page has what changes when you hold the pause.
15. The self-explanation prompt. Ask students to explain to themselves why a step follows from the one before, in writing, before you explain it. Two minutes. Bisra and colleagues pooled 69 effect sizes from 64 research reports on prompted self-explanation and found a weighted mean of g = 0.55. Collect them and reteach the step where the explanations run out.
16. Find the mistake. A worked solution with exactly one wrong line, and the question is which line and what the student was thinking. Two minutes. The task is to diagnose rather than to produce, and the answers name misconceptions in the students' own words.
17. Circulate with a tally. Walk the room during practice and count how many people made the same error, rather than fixing each one. Free, and it changes what you do at the board next. Ruiz-Primo and Furtak modelled this as a cycle: the teacher elicits, the student responds, and the teacher recognises and uses what came back. The teacher who more frequently used complete cycles had students with higher performance on the embedded assessment.
18. A target success rate. A number you watch rather than a question you ask. Rosenshine reports a fourth-grade mathematics study where the most successful teachers' classes got 82% of answers correct against 73% for the least successful, and puts the optimal rate at about 80%. His reason for the number is that 80% shows both that students are learning the material and that they are still being challenged. A room well above it is telling you to make the next question harder rather than to slow down.
Formative assessment examples for the end of a lesson
Seven techniques for the last three minutes. All of them are asynchronous in Black and Wiliam's sense: nothing changes today, and everything they are worth is spent in the next lesson.
19. The exit ticket. One question with a right answer, tied to the one thing the lesson existed to teach. Three minutes. Its own page carries seventy questions by subject, five formats, and the argument for never grading them. There is a printable version at Exit Ticket Template.
20. One sentence to somebody who missed it. Explain today's idea in a line, for a named absent classmate. Two minutes. The audience constraint is what stops the answer being a list of the words you used.
21. The muddiest point. Which part of today was least clear. One minute. It surfaces the confusion students can name, which is not the same as the confusion they have, so read it as a list of candidates rather than a diagnosis.
22. The question you would still ask. One thing you want answered at the start of next lesson. One minute. You are collecting the opening of the next lesson rather than a measurement of this one, so read them for the one question worth answering for everybody.
23. Three words for today. Three words describing the lesson, collected as a cloud. Ninety seconds. Useful for tone and vocabulary and useless for correctness, so do not run it on a day where you needed to know whether they got it.
24. Self-assessment against the criteria. Students rate their own work against the success criteria you shared at the start. Two minutes. It only works if the criteria were specific enough to disagree about, and it is worth collecting for the mismatches rather than the ratings.
25. Predict tomorrow's first question. Write the question you think will open the next lesson. Ninety seconds. A room that predicts it has a model of where the topic is going. A room that writes down the recap question you open with every week has told you about your routine instead.
Formative assessment examples that run across a unit
Seven techniques that need something one lesson cannot give you: old material, a class set of work, a five-week gap, a colleague.
26. Low-stakes quizzing on a schedule. A short unmarked quiz on old material, weekly. Five minutes. Yang and colleagues pooled 222 classroom studies covering 48,478 students and put the effect of classroom testing on achievement at g = 0.499. Run it on content from three weeks ago, not from this week, or you are measuring short-term memory.
27. Peer assessment against a rubric. Students mark each other's work using criteria you wrote. Fifteen minutes, on a page of work and a rubric the class has used before. Double, McGrane and Hopfenbeck pooled 54 studies and 141 effect sizes. Peer assessment improved academic performance at g = 0.31, better than no assessment and better than teacher assessment. It did not differ significantly from self-assessment. Mark the disagreements between the markers, because a criterion two students read differently is one you have not taught yet.
28. The whole-class feedback sheet. One page of the errors the whole set made, read out and worked through, instead of thirty individual comments. Twenty minutes for a class set, if you are reading for the pattern rather than writing on the work. The Education Endowment Foundation's review of practice in English schools found whole-class feedback and live marking among the most commonly cited feedback practices in the schools it visited. Put the three commonest errors on the board as questions and make the class fix them.
29. The same question twice. Ask one hard question in week one and again in week six, unchanged. Two minutes each time. Do not tell them it is coming back, or the second answer measures who revised. It shows what survived the unit rather than what survived the lesson, and the answer is sometimes uncomfortable.
30. A misconception log. A running list of the wrong models this cohort actually held, written down when you meet them. Free. Next year it becomes the distractor set for the misconception polls in that unit, which is the point.
31. A gallery of anonymised work. Four unnamed answers on the screen, the room ranking them against the criteria. Ten minutes. Use work from a previous year or write the four yourself, because a class of thirty recognises its own handwriting. A standard is easier to recognise than to define, and the argument about the ranking is where it gets defined.
32. A teacher learning community. A monthly meeting where a group of teachers commit to one formative technique and report back. Two hours a month is what the Education Endowment Foundation asked of teaching staff in its trial, covering the meeting and the paired lesson observation between meetings. Its process evaluation found teachers valued the dialogue between teachers and experimented more because of it. Bring the count your technique produced rather than the impression it left.
Which of the thirty-two actually have evidence
Sorted by what the evidence was actually measured on, the thirty-two fall into three boxes.

Direct evidence, measured on the technique itself. Retrieval practice and low-stakes quizzing, at g = 0.499 across 222 classroom studies. Prompted self-explanation, at g = 0.55. Peer assessment, at g = 0.31. All-student response cards, against hand raising as the control condition. These four you can run on the strength of the studies rather than on the strength of the argument.
Evidence for the mechanism, not for the named routine. Most of the list. Hinge questions, misconception polls, exit tickets, muddiest points, entrance tickets and the rest inherit their case from the general finding that eliciting and acting on evidence beats not doing so. That is a real inheritance and it is weaker than a trial. It also means the quality of your question matters more than the choice of format, which is not what a page of thirty-two techniques usually wants to tell you. Prequestions are here on different grounds: they have a literature of their own, and the review technique 1 cites calls it small next to retrieval practice.
Little or nothing, at least directly. The confidence sort, and the cousins it belongs with: thumbs up and down, traffic-light cards. They measure how sure a student is rather than what a student can do, and Hacker and colleagues found the lowest performing students grossly overconfident about their own results. A three-word cloud is the same: it reports mood and vocabulary, which are worth knowing and are not understanding.
One finding turns up in both the 1998 review and the 2018 trial, and it is about who gains. Black and Wiliam report that many of the studies they reviewed reached the same conclusion: improved formative assessment helps low achievers more than other students. That narrows the range of achievement while raising it overall. The Education Endowment Foundation trial saw the same shape. Additional progress for children in the lowest third for prior attainment was greater than for the highest third. The report rates that subgroup result as less secure than its headline.
Reading thirty answers without taking them home
Sorting beats marking. Count how many got it, how many half got it, and how many did not, then act on the biggest pile that did not get there. If both of those piles are small, the decision is to move on, and that is a result too. You are not producing a record. You are producing a decision about what the next lesson opens with.
The reason to keep it that cheap is that expensive versions get abandoned. The Education Endowment Foundation's guidance report puts Key Stage 3 teachers at 6.3 hours per week on written feedback in both 2013 and 2018. In the same report, 65% of secondary and 58% of primary teachers called their marking workload too much. A routine that adds to that number is asking for time those teachers have already said they do not have.
Cheap also happens to be safer. The same guidance report is blunt that the impact of feedback varies and, in some cases, can even hamper pupil progress, which is the classroom version of Kluger and DeNisi's negative third. Its first recommendation is to lay the foundations before giving any: set the learning intentions the feedback will aim at, and assess the learning gaps it will address.
Two rules are worth keeping whatever you run. Say the pattern out loud to the whole class, because a room that never hears what the check found learns that the check does not matter. And keep marks off anything you collected in order to teach better, for reasons the exit tickets page sets out with the experiment behind them.
Running these from the front of the room
Paper works and costs you the carrying. A live version puts the same question on the screen and the answers on your laptop before the room has packed up. Participants scan a QR code or open a short link on their own phones. Nothing installs and nobody makes an account.
Live Polls covers the largest share of this page. A poll can be single choice, multiple choice or open text, so one activity handles the four-option misconception check and the open request for the question they would still ask. A choice answer carries no name at all, and an open-text answer carries one only when its author ticks Share my name. Either way the default is the honest one for a check on the material.
Word Clouds suit the one-word entries: the key term, the three words for today, the vocabulary check. Repeated words grow, which turns thirty answers into one picture of what the room means by a word. The Word Cloud Generator does a different job: you paste text you already have, and it shows you what is in it.
Live Quizzes fit the scheduled retrieval quiz, with a timer and a leaderboard. One limit is worth knowing before you plan around it: quiz questions are presenter-authored, so a class cannot write the quiz for each other.
Reveal holds every answer face down until you turn the cards over. That matters for a prediction round, where the third answer is worthless if its author has already read the first two. It is also the one activity where you decide up front whether cards come back with names on them.
Afterwards, Activity Reports keeps each session's results in one place rather than in screenshots of the big screen. Two things about that are worth saying plainly. A choice poll reports counts and a cloud reports words, so neither will tell you which child is stuck. And a report covers one session, so technique 29 means opening two of them five weeks apart and doing the comparison yourself.
Be honest about the cost of running this daily. The free plan meters your sessions two ways a month: how many people you reach in total, and how many of those sessions had a real class in the room. Small sessions count against neither. A whole-class check every lesson starts a qualifying session every school day, and the session meter, not the headcount, is what runs out first. The current numbers are on the pricing page.
When a formative check is the wrong move
A check you collect and read, every lesson, forever, is the failure mode, and it has a specific shape: the checks keep happening and the teaching stops changing. Black and Wiliam are explicit that for assessment to function formatively the results have to be used to adjust teaching and learning. A weekly technique you act on beats a daily one you file.

Skip the ones you collect and read later when there is no next lesson to spend them on. The last lesson before the exam. A topic whose next instalment is a month away. A lesson you are covering and will not teach again. In each case you are collecting evidence you cannot spend, and a hinge question you act on inside the hour is a different decision.
Skip it when you already know, and be strict about what knowing means. Marking last night's work and finding the same error in six books is knowing. A feeling you formed at the front of the room while nobody said anything is the read this page exists to replace.
Be careful with anonymity when the answer is about a person rather than a topic. Anonymity is there so that being wrong in front of everybody costs nothing, which is what you want for a check on the material. It is exactly wrong for a wellbeing check, because a note that somebody is struggling, with no name on it, is one you cannot act on.
And be patient with the timescale. The Education Endowment Foundation's process evaluation concluded that it may take more time for improvements in teaching practice and pupil learning strategies to feed fully into attainment. That trial ran for two years to find two months of progress. A technique you drop after a fortnight because nothing visibly moved was never going to show you anything.
FAQ
Common questions
What is formative assessment?
Formative assessment is any check on learning whose results change what happens next. Black and Wiliam define practice as formative to the extent that evidence about student achievement is elicited, interpreted and used to decide the next steps in instruction. The definition deliberately turns on the use rather than the instrument, because evidence collected and never used would otherwise count.
What are examples of formative assessment?
Hinge questions, exit tickets, entrance tickets, mini whiteboards, misconception polls, think-pair-share, the muddiest point, low-stakes quizzes, self-explanation prompts, peer assessment against a rubric and whole-class feedback sheets. This page lists thirty-two of them, grouped by whether they run before you teach, during the lesson, at the door or across a unit.
What is the difference between formative and summative assessment?
The same test can be either, and the difference is what you do with the result. A summative assessment reports where a student got to. A formative one changes the teaching that follows it. A mock exam used to write next term's plan is formative; the same mock filed as a grade is summative.
Does formative assessment actually improve results?
Modestly, on the evidence of the large trial that tested it as a whole practice. The Education Endowment Foundation randomised 140 secondary schools into a two-year formative assessment programme. It found the equivalent of two additional months of progress in Attainment 8, at a very high security rating. That result was significant at the 10% level rather than the 5% one, and the trial found no evidence of an effect on English or Maths GCSE specifically.
Why is the effect size range quoted for formative assessment disputed?
Because that range was a summary of typical effect sizes across a diverse set of experiments, not a measured effect of a defined programme. When Kingston and Nash tried to compute one, only a small fraction of the studies carried enough information to do it, and their weighted mean came out well below the quoted range. McMillan and colleagues then argued that the studies which did qualify were too varied to pool at all.
What is a hinge question?
A hinge question is a single question at the pivot point of a lesson. Each wrong option is written to correspond to a specific misconception, and the decision rule is set before the lesson. You decide in advance what proportion of correct answers means you move on and what proportion means you reteach. Writing the rule afterwards is what turns it back into a warm-up.
How often should you use formative assessment?
Often enough that the answers still change something, which is a test rather than a number. Two techniques on this page cost nothing and belong in every lesson: waiting five seconds after a question, or counting the same error as you walk the room. Ration the ones that produce a pile to read. A check nobody acts on teaches a class that the check does not matter. The useful test is whether you can name a lesson that went differently because of one.
Run this with your own team
Start free, no credit card. Your audience joins from their phones with a code — nothing to install.