Here is the most uncomfortable finding in the whole feedback literature. In 1996, Avraham Kluger and Angelo DeNisi pulled together more than 600 effect sizes from feedback interventions and found that more than a third of them made performance worse. Not neutral. Worse. People received feedback, in good faith, from people trying to help, and they got worse at the thing they were being helped with.
Every leader who has ever sat in a post-observation conversation should find that number clarifying. Feedback is not automatically good. It is one of the most powerful interventions we have and one of the most fragile, and the difference between the two lies almost entirely in its design. After twenty years in classrooms and school leadership, I have come to believe that four properties do most of the work: feedback needs to be specific, timely, actionable and regular. Each of those words carries a genuine evidence base behind it, and each is more interesting than it first appears.
Specific: information is the active ingredient
The largest modern meta-analysis of feedback, Wisniewski, Zierer and Hattie (2020), pooled 435 studies covering more than 61,000 learners. The average effect was solid, but the average hides the real finding: the variation was enormous, and the strongest predictor of whether feedback worked was its information content. Feedback rich in precise, task-level information substantially outperformed praise, grades and general commentary. “Good questioning” carries almost no usable information. “In the first ten minutes you asked eleven recall questions and two that required reasoning, and the reasoning questions produced the longest pupil responses of the lesson” gives a teacher something they can inspect, verify and act on.
Kluger and DeNisi explained the backfire effect the same way. Feedback that keeps attention on the task tends to help. Feedback that draws attention to the self, through judgement, comparison or evaluation of the person, tends to harm. This is why a lesson grade is not merely blunt but actively counterproductive: it moves the conversation from what happened in the room to what the number says about me.
There is a caveat worth knowing. Research from organisational psychology by Goodman, Wood and Hendrickx (2004) found that extremely specific feedback can boost immediate performance while quietly suppressing the exploration that builds independent capability. If the feedback does all the thinking, the recipient does none. The resolution is a principle I would offer any leadership team: be maximally specific about the evidence, and deliberately open about the interpretation. Show precisely what happened; invite the teacher to reason about why and what next.
Timely: it is not about speed, it is about the next lesson
Timeliness has the most contested evidence base of the four, and it is worth being honest about that. The seminal review by Kulik and Kulik (1988) found that applied classroom studies generally favoured immediate feedback, while laboratory studies often favoured a delay. Valerie Shute’s influential 2008 review suggested immediate feedback serves procedural learning best, while a delay can sometimes support deeper transfer. Faster is not always better.
For teachers reflecting on lessons, though, three things are clearly true. First, memory decays fast: feedback built on a colleague’s recollection of a lesson loses detail within hours, and reconstruction errors creep in. Second, feedback lands better while the experience is still alive; a conversation about a lesson I taught this morning engages me differently from one about a lesson I can barely remember from a fortnight ago. Third, and most important, the deadline that actually matters is not “as soon as possible” but “before I next teach this class”. Feedback is valuable at the moment it can be used. A development point received on Tuesday and applied on Wednesday has better timing, in the sense that counts, than one received in minutes and never revisited.
Actionable: the best-evidenced word of the four
If I had to rank the four properties by weight of evidence, actionable comes first. Hattie and Timperley’s landmark model of feedback identifies “where to next” as the component with the greatest potential to accelerate learning, and the one most often missing in practice. And the strongest causal evidence we have for improving teaching points the same way: Kraft, Blazar and Hogan’s meta-analysis of teacher coaching found effects on instructional practice of around half a standard deviation, which is remarkable by the standards of education research. The defining ingredient of those coaching programmes was not the observation. It was the small, concrete, rehearsable action step that followed it.
The research suggests a good action step has four properties. It is small enough to attempt in the next lesson. It is behavioural and observable, so the teacher can tell whether they did it. It is high-leverage, chosen because it unlocks other improvements. And it comes alone, or nearly so, because working memory punishes long lists. “Improve your questioning” fails every test. “After asking a question, count three seconds before taking an answer” passes all four.
Regular: rhythm beats volume
One conversation, however brilliant, rarely changes practice. The learning sciences are unambiguous that spaced encounters with an idea beat a single intense exposure, and the EEF’s review of effective professional development (Sims and colleagues, 2021) identifies revisiting material, goal setting, prompts and follow-up among the mechanisms that make development stick. Habits form through repetition in a stable context, and teaching is nothing if not a stable context: the same room, the same classes, the same Tuesday afternoon.
There is a nuance here that busy leaders should find liberating. The coaching evidence shows no simple relationship between total hours of support and impact; beyond a modest level, more volume does not reliably mean more improvement. What matters is cadence: a sustainable rhythm of short cycles, each with one focus, each followed up. A school that gives every teacher one small development conversation a fortnight, every fortnight, will outgrow a school that delivers one exhaustive observation cycle a term.
The four work as a loop, not a list
Ranked by evidence, I would order them: actionable, specific, regular, timely. But the ranking matters less than the relationship. Regular feedback that is not actionable is noise on a schedule. Specific feedback that never returns is diagnosis without treatment. The four properties behave multiplicatively, and a serious weakness in any one devalues the rest. The schools that get this right are not running four separate initiatives; they are running one loop. Evidence from the lesson, arriving while it can shape the next one, carrying one small rehearsable step, revisited cycle after cycle.
That loop is worth auditing honestly. Most feedback in most schools, mine included at various points, fails at least two of the four tests. It is general rather than evidenced, delayed until the lesson is a memory, descriptive rather than actionable, and episodic rather than regular. None of this reflects a lack of goodwill. It reflects the arithmetic of leadership time. The evidence on what works has been clear for years; the constraint has always been capacity.
That constraint is the problem I have spent the last two years working on with Starlight, which turns a whole-lesson audio recording into a private, transcript-grounded coaching report for the teacher within minutes: evidenced strengths, a small number of high-leverage next steps, and reflection prompts, built deliberately around the four properties above. I have written elsewhere on the Starlight blog about how the research translates into practice, if you want to go deeper.
Book a demo. If you are thinking about how feedback works in your school and want to see this in action, you can schedule a personalised demo here.
Spark Insight with Starlight, and let the evidence do the talking in every feedback conversation.
The Insight Engine is written by Adam Sturdee, co-founder of Starlight, the UK’s teacher-first AI-powered coaching platform, and a senior leader with responsibility for teaching, learning and coaching. This blog is part of a wider mission to support educators through meaningful reflection, not performance metrics. It documents the journey of building Starlight from the ground up, and explores how AI, when shaped with care, can reduce workload, surface insight, and help teachers think more deeply about their practice. Rooted in the belief that growth should be private, professional, and purposeful, The Insight Engine offers ideas and stories that put insight, not judgment, at the centre of development.