diff --git a/README.md b/README.md new file mode 100644 index 0000000..5b95bfb --- /dev/null +++ b/README.md @@ -0,0 +1,124 @@ +# Nurture — formative assessment skills for AI agents + +An [Agent Plugin](https://agent-plugins.org/) that gives an AI agent the pedagogy +to be useful to a teacher: how to give feedback that actually changes what a +learner does, how to build assessments that produce a decision rather than a +score, and how to read a class set and know what to teach next. + +Grounded throughout in **Dylan Wiliam's** work on formative assessment — +*Inside the Black Box* (Black & Wiliam, 1998), *Embedded Formative Assessment* +(2011/2018) and *Embedding Formative Assessment* (Wiliam & Leahy, 2015). + +## Why + +Ask a general-purpose model to "give feedback on this essay" and you get four +hundred fluent words that praise the student, correct every error, and leave them +with nothing to do. That is not a small quality problem — Kluger & DeNisi's +meta-analysis found feedback interventions *lowered* performance in 38% of cases, +mostly when attention moved from the task to the self. The default behaviour is +the failure mode. + +These skills encode what is known about doing it well, and — as importantly — the +guardrails against the specific ways AI-generated teaching material goes wrong: +volume, generic praise, invented misconceptions, over-correction, and confident +claims about curricula that vary by jurisdiction. + +## Skills + +| Skill | What it does | +|---|---| +| [`formative-assessment`](skills/formative-assessment/) | Orients a task around Wiliam's five key strategies and routes to the right skill below | +| [`giving-feedback`](skills/giving-feedback/) | Writes and critiques feedback that is task-focused, actionable, and more work for the recipient than the donor | +| [`designing-assessments`](skills/designing-assessments/) | Blueprints quizzes, tests and tasks backwards from the decision they inform; item-writing and fairness checks | +| [`hinge-questions`](skills/hinge-questions/) | Single diagnostic items at a lesson's decision point, where every wrong answer names a misconception | +| [`success-criteria`](skills/success-criteria/) | Learning intentions, success criteria, rubrics, and the exemplar work that makes them mean anything | +| [`eliciting-evidence`](skills/eliciting-evidence/) | Questioning and all-student response techniques that surface what everyone thinks, not just the volunteers | +| [`peer-and-self-assessment`](skills/peer-and-self-assessment/) | Protocols that make learners resources for one another and owners of their own learning | +| [`analysing-student-work`](skills/analysing-student-work/) | Reads a class set, finds the shared misconceptions, and produces the next teaching move | + +Each skill carries its own `references/` with the deeper material — technique +banks, worked examples, the evidence base, and the honest limits of each claim. + +## The framework + +Wiliam's five key strategies, which the skills map onto directly: + +| | Where the learner is going | Where the learner is right now | How to get there | +|--------------|---------------------------|--------------------------------|------------------| +| **Teacher** | Clarifying and sharing learning intentions and success criteria | Engineering discussions and tasks that elicit evidence of learning | Providing feedback that moves learners forward | +| **Peer** | ← Understanding learning intentions and criteria → | ← Activating learners as instructional resources for one another → | | +| **Learner** | | ← Activating learners as owners of their own learning → | | + +And the one big idea underneath: **use evidence of learning to adapt teaching to +meet learner needs.** + +## Structure + +``` +nurture/ +├── plugin.json +├── mcp.json +└── skills/ + ├── formative-assessment/ + ├── giving-feedback/ + ├── designing-assessments/ + ├── hinge-questions/ + ├── success-criteria/ + ├── eliciting-evidence/ + ├── peer-and-self-assessment/ + └── analysing-student-work/ +``` + +Conforms to Agent Plugins 1.0.0. Skills are plain `SKILL.md` files with +`references/` and `scripts/` alongside — portable to any client implementing the +specification. + +## MCP + +`mcp.json` declares the **Nurture Signal** MCP server +(`https://signal.gonurture.com/mcp`, OAuth-protected), which gives skills access +to live session data: engagement reactions, comments, questions and analytics. + +Every skill also works with no server connected — the pedagogy is the substance, +and the data is an accelerant. Where a skill can use live class data it says so, +and it says what that data can and cannot license you to conclude. Engagement +reactions are not evidence of understanding, and the skills say so every time. + +## Tools + +`skills/analysing-student-work/scripts/item_analysis.py` — item analysis from a +CSV of responses: facility, distractor frequencies, discrimination, and flags for +items that are broken rather than hard. Standard library only. + +``` +python3 item_analysis.py responses.csv +python3 item_analysis.py responses.csv --key B,A,C,D --json +``` + +## Using this responsibly + +The skills are written on these assumptions, and state them in their own output: + +- **A teacher reviews anything that reaches a learner.** Nothing here is designed + to run unsupervised between a model and a child. +- **Nothing generated here should carry a high-stakes grade** without human + marking. +- **Learners are named only in teacher-facing output.** Anything shown to a class + is anonymised. +- **Claims are made at the strength the evidence supports.** Where the research is + thin — peer assessment quality, self-assessment accuracy, engagement metrics — + the skills say so rather than overselling. + +## References + +- Wiliam, D. (2018) *Embedded Formative Assessment*, 2nd ed. Solution Tree. +- Wiliam, D. & Leahy, S. (2015) *Embedding Formative Assessment*. Learning Sciences International. +- Black, P. & Wiliam, D. (1998) *Inside the Black Box*. King's College London. +- Black, P. & Wiliam, D. (2009) "Developing the theory of formative assessment", *Educational Assessment, Evaluation and Accountability* 21(1). +- Sadler, D. R. (1989) "Formative assessment and the design of instructional systems", *Instructional Science* 18. +- Kluger, A. N. & DeNisi, A. (1996) "The effects of feedback interventions on performance", *Psychological Bulletin* 119(2). +- Butler, R. (1988) "Enhancing and undermining intrinsic motivation", *British Journal of Educational Psychology* 58. +- Hattie, J. & Timperley, H. (2007) "The power of feedback", *Review of Educational Research* 77(1). + +See [`skills/formative-assessment/references/evidence-base.md`](skills/formative-assessment/references/evidence-base.md) +for what each study actually found, and where the common overclaims are. diff --git a/mcp.json b/mcp.json new file mode 100644 index 0000000..23af7d6 --- /dev/null +++ b/mcp.json @@ -0,0 +1,9 @@ +{ + "$schema": "https://agent-plugins.org/schemas/1.0.0/mcp.schema.json", + "mcpServers": { + "nurture-signal": { + "type": "streamable-http", + "url": "https://signal.gonurture.com/mcp" + } + } +} diff --git a/plugin.json b/plugin.json new file mode 100644 index 0000000..c4127d8 --- /dev/null +++ b/plugin.json @@ -0,0 +1,25 @@ +{ + "$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json", + "name": "nurture", + "version": "0.1.0", + "description": "Formative assessment skills grounded in Dylan Wiliam's pedagogy: feedback that moves learners forward, assessments and hinge questions that elicit evidence of learning, success criteria, peer and self-assessment, and responsive analysis of student work.", + "author": { + "name": "Nurture", + "url": "https://gonurture.com" + }, + "homepage": "https://gonurture.com", + "repository": "https://github.com/goNurture/agent-plugin", + "keywords": [ + "education", + "teaching", + "learning", + "formative-assessment", + "assessment-for-learning", + "feedback", + "hinge-questions", + "success-criteria", + "responsive-teaching", + "dylan-wiliam", + "edtech" + ] +} diff --git a/skills/analysing-student-work/SKILL.md b/skills/analysing-student-work/SKILL.md new file mode 100644 index 0000000..505f8ee --- /dev/null +++ b/skills/analysing-student-work/SKILL.md @@ -0,0 +1,118 @@ +--- +name: analysing-student-work +description: Read this when you have more than one learner's work in front of you and the question is what to do about it. Use it when given a class set, assignment or quiz results, a mark sheet or response CSV, or when asked what a class got wrong, which items were poor, who needs intervention, or what to teach next. It produces a next teaching move with a built-in check, not a report. +--- + +# Analysing student work — responsive teaching + +The point of collecting evidence is the decision that follows. The deliverable +here is what happens in the next lesson. + +## What done looks like + +``` +CLASS: learners · · + +WHAT THE EVIDENCE SHOWS + + +WHAT'S BEHIND IT + + +THE MOVE — next lesson + Starter ( min): + Then: + Check: + +DIFFERENT NEEDS + Secure (): Group (): + Individual (teacher-only names): + +ITEMS TO FIX + + +WHAT THIS EVIDENCE CANNOT TELL YOU + +``` + +## Constraints + +- **Wrong answers before scores.** The score distribution is nearly + information-free; the pattern of *specific* wrong answers is the data. +- **Cluster by cause, not by surface error.** Six different wrong answers may come + from one misconception. `references/error-taxonomy.md` sorts them into slip / + misconception / gap / misread-the-task / not-attempted — five categories with + five different responses. Reteaching a slip wastes everyone's time. +- **Findings carry numbers**, not impressions. +- **Every misconception is named in plain language**, never "struggled with X". +- **A check is built in** that will show whether the move worked. +- **Broken items are separated from genuine difficulty.** Roughly one broken item + per assessment is normal, and a broken item's "findings" are noise. +- **No claim rests on fewer items than it can bear.** Five items cannot support a + claim about an individual; one class cannot support a claim about a cohort. + +## Reading the numbers + +Run `scripts/item_analysis.py` on any set of objective responses — facility, +distractor frequencies, discrimination, and flags for items that are broken rather +than hard. `python3 item_analysis.py --help`. + +**The single highest-value finding is a distractor a third of the class chose.** +That's a nameable misconception with an obvious response, and it's what turns a +set of results into a lesson. + +Other patterns: facility above 0.9 means the item isn't diagnostic here; below 0.3 +means check the item before concluding anything about the class; wrong answers +spread evenly means guessing, so go back further than you think; negative +discrimination means the key or wording is wrong, so don't act on the item at all. + +Caveats to state every time: discrimination is unstable below ~30 learners — with +one class it's a hint, not a finding; item statistics describe behaviour *in this +group*, not quality in the abstract. + +## Reading open responses + +Don't score first. **Sort.** Read the whole set once without marking, then sort +into 3–5 piles by what each response *does* — narrates instead of explaining; +explains but no evidence; evidence but no link back to the question; secure. Name +each pile in a sentence a learner would understand: that sentence is your +whole-class teaching point. The biggest pile determines the next lesson. Pull two +anonymised extracts — modal error and secure — for a best-of-two starter. + +Sorting is faster than marking and produces something you can teach from, which +marking usually doesn't. + +## Choosing the grain of response + +| Share of class | Response | +|---|---| +| More than about a third | Whole-class reteach and one feedback sheet. Not thirty comments | +| Roughly 10–30% | Targeted group during the next task while others extend | +| A handful | Individual feedback with a specific task | +| One learner, repeatedly, across unrelated topics | Not a topic problem. Flag prerequisites, attendance or access needs to the teacher — say "worth checking", don't diagnose | + +## Working with connected data + +If a Nurture server is available, go from the class down to the **actual +responses** — aggregate statistics hide exactly the distractor patterns you need. +Look across assignments to tell a topic problem from a persistent one, and read +learners' own reflections, which often identify the sticking point faster than the +work does. Discover the available tools rather than assuming names. + +Session analytics can locate *when* comprehension dropped, but see the warning in +`eliciting-evidence` about treating engagement data as evidence of understanding. + +## Privacy and proportionality + +- Individual learners are named only in teacher-facing output. Anything shown to a + class is anonymised. +- Never rank learners by attainment. +- Never infer ability, home circumstances, effort or character from results. Report + what the work shows and stop. +- Don't carry a learner's data between contexts the teacher didn't intend. + +## References + +- `scripts/item_analysis.py` — item analysis from a CSV. +- `references/error-taxonomy.md` — the five error categories, how to sort a class + set fast, the diagnostic questions to ask a learner, and what not to conclude. diff --git a/skills/analysing-student-work/references/error-taxonomy.md b/skills/analysing-student-work/references/error-taxonomy.md new file mode 100644 index 0000000..358a704 --- /dev/null +++ b/skills/analysing-student-work/references/error-taxonomy.md @@ -0,0 +1,102 @@ +# Clustering wrong answers by cause + +Six wrong answers can have one cause; two identical wrong answers can have two. +Sorting by *what the learner did* rather than by *how wrong they were* is what +turns a pile of errors into a teaching move. + +## The five categories + +Sort every wrong response into one of these. The response differs for each. + +### 1. Slip +Knows the method, executed it wrong. Sign errors, transcription, a dropped +term, misread digit. **Marker:** the learner spots it instantly when asked to +check. **Response:** not reteaching — a checking routine. Estimating before +calculating, re-reading the question against the answer, working backwards. +Reteaching a slip wastes everyone's time and tells the learner you think they +don't understand. + +### 2. Misconception +Consistent, rule-governed, and *reasonable* given a restricted diet of examples. +"Multiplication makes bigger" is true for every whole number greater than one. +**Marker:** the same error appears across items and the learner defends it. +**Response:** direct confrontation. Predict–observe–explain with a case where the +rule fails, then rebuild the boundary. Reteaching the correct method *without* +addressing the wrong one leaves the misconception intact underneath, and it +returns under pressure. + +### 3. Gap +The prerequisite isn't there. **Marker:** the error is in a *sub*-step, not the +target step. A learner failing an algebra item on arithmetic. **Response:** go +back to the prerequisite. Repeating the target lesson louder will not work. + +### 4. Misreading the task +Understood the content, answered a different question. **Marker:** the response is +competent but off-target. **Response:** not content at all — question decoding. +Underlining the command word, restating the question, checking the answer against +what was asked. Frequently a reading-load fault in the item; check that first. + +### 5. Not attempted +**Marker:** blank. Distinguish, by asking: ran out of time / didn't know where to +start / didn't think it was worth attempting. Three different causes, three +different responses. A high omission rate on late items is a timing or stamina +problem, not a knowledge problem — and a high omission rate on *one* item usually +means the item is inaccessible. + +## Sorting a class set fast + +For a set of 30 objective responses, use the skill's +`scripts/item_analysis.py`. + +For constructed responses, sort physically or in a list: + +1. **First pass, no marking.** Read everything to establish the range. +2. **Second pass, sort into piles** by category above. For piles of type 2 + (misconception), split further by *which* misconception. +3. **Count.** The counts determine the grain of response — see the table in + SKILL.md. +4. **Name each pile in one sentence a learner would recognise.** "You explained + what the source says instead of how much we can trust it." That sentence is + the teaching point, and it goes straight onto the whole-class feedback sheet. +5. **Pull two anonymised extracts** — the modal error and a secure response — for + a best-of-two comparison starter. + +Sorting is faster than marking and produces something you can teach from, which +marking usually doesn't. + +## The diagnostic questions to ask a learner + +When the written work isn't enough to tell which category you're in, two minutes +of talk usually settles it: + +- "Talk me through what you did." → separates slip from misconception; a slip is + spotted mid-sentence. +- "Would that always work?" → exposes over-generalisation. +- "How could you check?" → tells you whether any self-monitoring exists. +- "What was the question asking?" → separates misreading from content failure. +- "Where did you get stuck?" → for non-attempts, and worth more than any guess. + +## Patterns across a whole class + +| What you see | Likely cause | Response | +|---|---|---| +| One distractor takes a third of the class | A single shared misconception | Whole-class reteach of that specific idea | +| Errors spread evenly across all options | Guessing; the concept isn't there | Go back further than you think | +| Strong learners wrong, weak learners right | The key or the wording is wrong | Fix the item; don't act on it | +| Right answers, wrong reasoning (two-tier item) | Surface pattern-matching | Vary the surface features; teach the structure | +| Fine in class, poor in the assessment | Was scaffolded, hasn't transferred | Fade the scaffold deliberately, don't remove it | +| Fine last term, poor now | Not consolidated | Spaced retrieval, not reteaching | +| One learner wrong across unrelated topics | Not a topic problem | Flag to the teacher: prerequisites, attendance, access needs. Say "worth checking" — do not diagnose | + +## What not to conclude + +- **Don't infer effort, attitude or ability from an error pattern.** You cannot + see any of those in the work, and the guess is usually wrong and always + unhelpful. +- **Don't treat a low score as a measure of the learner.** It is a measure of the + distance between this task and this learner today, which is a fact about the + teaching sequence as much as about the learner. +- **Don't act on an item you haven't checked.** Roughly one broken item per + assessment is normal, and a broken item's "findings" are noise. +- **Don't scale conclusions beyond the evidence.** Five items cannot support a + claim about an individual; one class cannot support a claim about a cohort. diff --git a/skills/analysing-student-work/scripts/item_analysis.py b/skills/analysing-student-work/scripts/item_analysis.py new file mode 100755 index 0000000..8c3a493 --- /dev/null +++ b/skills/analysing-student-work/scripts/item_analysis.py @@ -0,0 +1,394 @@ +#!/usr/bin/env python3 +"""Item analysis for a class set of objective responses. + +Reads a CSV of learner responses and reports, per item: facility, the frequency +of every option chosen, and a discrimination index. Flags items whose statistics +suggest the item is broken rather than the learners. + +Standard library only; no dependencies. + +INPUT FORMAT +------------ +A CSV whose first column identifies the learner and whose remaining columns are +items. Blank cells are treated as omissions. + + student,Q1,Q2,Q3,Q4 + Alex,B,C,A,D + Bo,B,A,A,C + ... + +The key is supplied either as a row in the CSV whose identifier is "KEY" +(case-insensitive), or with --key. + + student,Q1,Q2,Q3,Q4 + KEY,B,A,A,D + Alex,B,C,A,D + +USAGE +----- + python3 item_analysis.py responses.csv + python3 item_analysis.py responses.csv --key B,A,A,D + python3 item_analysis.py responses.csv --json + python3 item_analysis.py responses.csv --id-column 0 --key-row KEY + +INTERPRETING THE OUTPUT +----------------------- +facility Proportion correct. >0.90 = not diagnostic here; <0.30 = check + the item before concluding anything about the class. +discrimination Upper-lower index: (facility in the top group) - (facility in the + bottom group), grouped by total score. Near zero or negative means + the item does not separate learners who know the material from + those who do not -- do not act on its results. +distractors Proportion choosing each option. A distractor taken by 30%+ of the + class is the most useful thing in this report: it is a specific, + shared misconception with an obvious teaching response. + +Caveats worth repeating to any teacher reading this: discrimination is unstable +below about 30 learners, and with a single class it is a hint, not a finding. +Item statistics describe how the item behaved in THIS group, not its quality in +the abstract. +""" + +from __future__ import annotations + +import argparse +import csv +import json +import math +import sys +from collections import Counter + +# Fraction of the cohort placed in the upper and lower groups for the +# discrimination index. 0.27 is the conventional split (Kelley, 1939), which +# maximises the stability of the index for normally distributed scores. +GROUP_FRACTION = 0.27 + +HIGH_FACILITY = 0.90 +LOW_FACILITY = 0.30 +WEAK_DISCRIMINATION = 0.20 +NOTABLE_DISTRACTOR = 0.30 +DEAD_DISTRACTOR = 0.05 +SMALL_COHORT = 30 + + +class InputError(Exception): + """Raised for malformed input, reported without a traceback.""" + + +def read_csv(path, id_column, key_row_label): + """Return (item_names, [(learner_id, [responses])], key_or_None).""" + try: + if path == "-": + rows = list(csv.reader(sys.stdin)) + else: + with open(path, newline="", encoding="utf-8-sig") as handle: + rows = list(csv.reader(handle)) + except OSError as exc: + raise InputError(f"could not read {path}: {exc}") from exc + + rows = [row for row in rows if any(cell.strip() for cell in row)] + if len(rows) < 2: + raise InputError("need a header row and at least one response row") + + header = rows[0] + if id_column >= len(header): + raise InputError( + f"--id-column {id_column} is out of range for {len(header)} columns" + ) + + item_names = [ + name.strip() or f"item{i}" + for i, name in enumerate(header) + if i != id_column + ] + if not item_names: + raise InputError("no item columns found") + + key = None + learners = [] + for row in rows[1:]: + # Tolerate short rows; treat missing trailing cells as omissions. + padded = list(row) + [""] * (len(header) - len(row)) + identifier = padded[id_column].strip() + responses = [ + padded[i].strip() for i in range(len(header)) if i != id_column + ] + if identifier.upper() == key_row_label.upper(): + key = responses + continue + learners.append((identifier, responses)) + + if not learners: + raise InputError("no learner rows found (only a key row?)") + return item_names, learners, key + + +def normalise(value): + """Case-fold a response so 'b' and 'B' are the same answer.""" + return value.strip().upper() + + +def score_matrix(learners, key): + """Return a list of per-learner correctness lists (1/0), aligned to the key.""" + scored = [] + for _, responses in learners: + scored.append( + [ + 1 if normalise(response) == normalise(expected) and expected else 0 + for response, expected in zip(responses, key) + ] + ) + return scored + + +def discrimination(scored, totals, item_index): + """Upper-lower discrimination index for one item. + + Learners are ranked by total score; the index is the item's facility in the + top GROUP_FRACTION minus its facility in the bottom GROUP_FRACTION. Returns + None when the cohort is too small to split. + """ + n = len(scored) + group_size = int(math.floor(n * GROUP_FRACTION)) + if group_size < 1: + return None + + order = sorted(range(n), key=lambda i: totals[i]) + lower = order[:group_size] + upper = order[-group_size:] + + upper_correct = sum(scored[i][item_index] for i in upper) / group_size + lower_correct = sum(scored[i][item_index] for i in lower) / group_size + return upper_correct - lower_correct + + +def analyse(item_names, learners, key): + if len(key) != len(item_names): + raise InputError( + f"key has {len(key)} answers but there are {len(item_names)} items" + ) + + # A blank key cell would otherwise mark every response for that item wrong, + # producing a facility of 0.00 and a "check the item" flag that looks like a + # finding about the class. Reject it rather than report it. + blank = [item_names[i] for i, answer in enumerate(key) if not answer.strip()] + if blank: + raise InputError( + "the answer key is incomplete — no correct answer given for: " + + ", ".join(blank) + ) + + scored = score_matrix(learners, key) + totals = [sum(row) for row in scored] + n = len(learners) + + items = [] + for index, name in enumerate(item_names): + responses = [normalise(r[index]) for _, r in learners] + counts = Counter(response or "(omitted)" for response in responses) + correct = sum(row[index] for row in scored) + + options = [ + { + "option": option, + "count": count, + "proportion": count / n, + "is_key": option == normalise(key[index]), + } + for option, count in sorted( + counts.items(), key=lambda kv: (-kv[1], kv[0]) + ) + ] + + items.append( + { + "item": name, + "key": normalise(key[index]), + "n": n, + "correct": correct, + "facility": correct / n, + "discrimination": discrimination(scored, totals, index), + "options": options, + "flags": [], + } + ) + + for item in items: + item["flags"] = flag_item(item) + + return { + "n_learners": n, + "n_items": len(item_names), + "mean_score": sum(totals) / n, + "score_distribution": dict(sorted(Counter(totals).items())), + "items": items, + } + + +def flag_item(item): + flags = [] + facility = item["facility"] + disc = item["discrimination"] + key_suspect = disc is not None and disc < 0 + + if facility > HIGH_FACILITY: + flags.append( + "near-ceiling: not diagnostic in this group; keep only as a " + "prerequisite or confidence check" + ) + if facility < LOW_FACILITY: + flags.append( + "very low facility: check the item itself before concluding " + "anything about the class" + ) + if disc is not None and disc < 0: + flags.append( + "NEGATIVE discrimination: stronger learners did worse — the key or " + "the wording is probably wrong. Do not act on this item" + ) + elif disc is not None and disc < WEAK_DISCRIMINATION: + flags.append( + "weak discrimination: the item does not separate learners who know " + "the material from those who do not" + ) + + for option in item["options"]: + if option["is_key"] or option["option"] == "(omitted)": + continue + if option["proportion"] >= NOTABLE_DISTRACTOR: + if key_suspect: + # Don't send a teacher off to reteach a misconception that may + # be an artefact of a mis-keyed item. Resolve the key first. + flags.append( + f"distractor {option['option']} taken by " + f"{option['proportion']:.0%}, but discrimination is negative " + f"— check the key before treating this as a misconception" + ) + else: + flags.append( + f"distractor {option['option']} taken by " + f"{option['proportion']:.0%} — a shared misconception; name " + f"it and reteach it" + ) + elif option["proportion"] <= DEAD_DISTRACTOR: + flags.append( + f"distractor {option['option']} attracted " + f"{option['proportion']:.0%} — doing no diagnostic work, replace it" + ) + + omitted = next( + (o for o in item["options"] if o["option"] == "(omitted)"), None + ) + if omitted and omitted["proportion"] > 0.10: + flags.append( + f"{omitted['proportion']:.0%} omitted — check timing, position on the " + f"paper, or reading demand" + ) + + return flags + + +def bar(proportion, width=24): + filled = int(round(proportion * width)) + return "█" * filled + "·" * (width - filled) + + +def report(result): + out = [] + out.append( + f"{result['n_learners']} learners · {result['n_items']} items · " + f"mean {result['mean_score']:.1f}/{result['n_items']}" + ) + if result["n_learners"] < SMALL_COHORT: + out.append( + f"NOTE: fewer than {SMALL_COHORT} learners. Treat discrimination as a " + "hint, not a finding, and do not draw individual-level conclusions " + "from a small number of items." + ) + out.append("") + + for item in result["items"]: + disc = item["discrimination"] + disc_text = "n/a (cohort too small)" if disc is None else f"{disc:+.2f}" + out.append( + f"{item['item']} key={item['key']} " + f"facility={item['facility']:.2f} discrimination={disc_text}" + ) + for option in item["options"]: + marker = "✔" if option["is_key"] else " " + out.append( + f" {marker} {option['option']:<12} {bar(option['proportion'])} " + f"{option['proportion']:>5.0%} ({option['count']})" + ) + for flag in item["flags"]: + out.append(f" ! {flag}") + out.append("") + + priorities = [ + (item["item"], flag) + for item in result["items"] + for flag in item["flags"] + if "shared misconception" in flag or "NEGATIVE" in flag + ] + if priorities: + out.append("ACT ON THESE FIRST") + for name, flag in priorities: + out.append(f" {name}: {flag}") + out.append("") + + out.append( + "Reminder: these statistics describe how the items behaved in this group. " + "The teaching decision comes from the misconceptions behind the " + "distractors, not from the scores." + ) + return "\n".join(out) + + +def main(argv=None): + parser = argparse.ArgumentParser( + description="Item analysis for a class set of objective responses.", + formatter_class=argparse.RawDescriptionHelpFormatter, + epilog=__doc__.split("USAGE")[1] if "USAGE" in __doc__ else None, + ) + parser.add_argument("csv", help="path to the responses CSV, or - for stdin") + parser.add_argument( + "--key", + help="comma-separated correct answers, one per item " + "(overrides a KEY row in the file)", + ) + parser.add_argument( + "--key-row", + default="KEY", + help="identifier marking the key row in the CSV (default: KEY)", + ) + parser.add_argument( + "--id-column", + type=int, + default=0, + help="zero-based index of the learner identifier column (default: 0)", + ) + parser.add_argument( + "--json", action="store_true", help="emit JSON instead of a text report" + ) + args = parser.parse_args(argv) + + try: + item_names, learners, file_key = read_csv( + args.csv, args.id_column, args.key_row + ) + key = [k.strip() for k in args.key.split(",")] if args.key else file_key + if not key: + raise InputError( + "no answer key found — add a row identified as " + f"'{args.key_row}' or pass --key" + ) + result = analyse(item_names, learners, key) + except InputError as exc: + parser.exit(2, f"error: {exc}\n") + + print(json.dumps(result, indent=2) if args.json else report(result)) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/skills/designing-assessments/SKILL.md b/skills/designing-assessments/SKILL.md new file mode 100644 index 0000000..c34f0df --- /dev/null +++ b/skills/designing-assessments/SKILL.md @@ -0,0 +1,103 @@ +--- +name: designing-assessments +description: Read this before generating any quiz, test, task or item — assessment written item-first drifts to recall and produces a score nobody can act on. Use it when asked to build a quiz, test, exam paper, question set, retrieval practice or knowledge check, to write or fix items and distractors, or to judge whether an existing assessment measures what it claims. It produces a blueprint, items whose wrong answers are diagnostic, and a guide to what each result means. +--- + +# Designing assessments that elicit evidence of learning + +Wiliam's second key strategy. An assessment is a machine for producing an +inference about what learners know. Design it backwards from the inference. + +## The first question, always + +**What decision will this inform, and what would make you decide differently?** + +If the answer is "it tells me who's doing well", stop — a ranking supports almost +no instructional decision. Push for a real one: + +| Decision | Shape | +|---|---| +| Move on or reteach, now | One hinge question → `hinge-questions` | +| Which of several misconceptions is in the room | Diagnostic MCQ, distractors mapped to misconceptions | +| Are the prerequisites there before I start | 5–8 items, a week ahead | +| Who needs which intervention | Item-level diagnostic, clean subscales | +| Has it stuck since last term | Spaced retrieval, no new content | +| A grade for reporting | Summative — say so plainly, design for reliability | + +Never blur the last row into the others. The same test can be used either way, but +the design trade-offs differ: formative wants diagnostic richness per item, +summative wants reliability across the whole. Optimise for both, get neither. + +## What done looks like + +1. **Blueprint** — outcomes × cognitive demand × item count. Written before any item. +2. **The items.** +3. **Mark scheme**, with acceptable alternatives spelled out. +4. **Interpretation guide** — per item: distractor → misconception → action. This + is what separates a formative instrument from a score generator. +5. **Timing**, and where it sits in the sequence. +6. **What this assessment cannot tell you** — one honest paragraph, always included. + It's what stops the scores being over-read. + +## Constraints + +- **Blueprint first.** Items written first drift to recall, because recall items + are the easy ones to write. The grid exposes it immediately. +- **Every distractor has a named misconception behind it.** Write the misconception + in plain language, work the problem as a learner holding it, and the answer they + get is the distractor. If you can't name it, you haven't designed the item. +- **No item is answerable correctly without the target knowledge**, or incorrectly + because of something irrelevant to it. +- **The interpretation guide exists.** No guide, not shipped. +- **Item count is defensible for the claims made.** Six items cannot support an + individual placement decision. Say so before the scores exist. + +## Validity is a property of inferences, not of tests + +Correct this whenever someone asks "is this a valid test?". State the inference, +then attack it on three fronts: + +- **Construct under-representation** — does it cover the domain, or the + easily-testable corner? A "scientific enquiry" test made of recall items doesn't. +- **Construct-irrelevant variance** — does anything *else* move the scores? Reading + demand in a maths problem, cultural assumptions in a comprehension text, time + pressure on a test that isn't about fluency. These are design faults, not learner + deficits. Run `references/fairness-check.md` over every item. +- **Reliability** — same score tomorrow, or with a different marker, or on a + parallel form? Short assessments are unreliable by construction. + +## Choosing the response format + +Per outcome, not per assessment; mixed-format is normal. Need to know *which* wrong +idea → diagnostic MCQ. Need to see reasoning → short constructed response. Need +procedural fluency → several similar items. Need transfer → novel context, same +structure. Need a performance → task plus rubric, and accept the reliability cost. + +`references/item-types.md` has the full comparison, including **two-tier items** +(answer + reason), which are the only cheap way to catch right-answer-wrong-reason +— the failure mode conventional marking hides. + +## Your specific failure modes + +- **Distractor padding.** Three plausible-sounding wrong options with no diagnostic + value. Name the misconception or replace the option. +- **All-recall drift.** Recall is cheapest for you to generate. Check the grid. +- **Answer leakage.** Longest option is the key; the key is grammatically + consistent with the stem and the others aren't; "all of the above" as filler; + key clustered on B and C. Check the distribution. +- **Context creep.** Elaborate scenarios that raise reading load without adding + construct relevance. +- **Fabricated stimulus.** Never invent data, sources, quotations or historical + detail and present them as real. Mark invented stimulus as invented. +- **Curriculum overconfidence.** Specifications differ by jurisdiction, board and + year. Ask which; if you can't, state the assumption at the top. + +Say plainly that AI-generated items contributing to a grade need checking against +the specification first. + +## References + +- `references/item-writing.md` — stems, options, distractor construction, item + types that reveal reasoning, retrieval-practice design. +- `references/item-types.md` — response format selection; two-tier items. +- `references/fairness-check.md` — construct-irrelevant variance, item by item. diff --git a/skills/designing-assessments/references/fairness-check.md b/skills/designing-assessments/references/fairness-check.md new file mode 100644 index 0000000..14e93f3 --- /dev/null +++ b/skills/designing-assessments/references/fairness-check.md @@ -0,0 +1,80 @@ +# Fairness check — construct-irrelevant variance + +Run every item past this list. Each hit is a design fault that will systematically +depress the scores of a group of learners for reasons unrelated to what you meant +to measure. + +## Language and reading + +- [ ] Is the reading demand higher than the construct requires? (Unless you are + assessing reading, the text should be as plain as possible.) +- [ ] Any idiom, colloquialism or culturally specific phrasing? ("Off the bat", + "a fair innings", "left field".) +- [ ] Any vocabulary that isn't the target vocabulary but is needed to parse the + question? +- [ ] Would a learner using English as an additional language be slowed by + syntax rather than by content? (Nested clauses, passive constructions, + conditionals.) +- [ ] Is any technical term used before it has been taught in that sense? + +## Context and background knowledge + +- [ ] Does the item assume experience not shared by all learners? (Holidays + abroad, particular sports, home internet access, a car, a garden, specific + religious or seasonal practice, particular family structure.) +- [ ] Does the context advantage learners from a particular region or country? + (Currency, units, place names, school system.) +- [ ] Would a learner who has never encountered this context still be able to + show the target knowledge? + +## Format and access + +- [ ] Does the item depend on colour discrimination alone? +- [ ] Is any diagram legible in black and white, and does it work with a screen + reader or a description? +- [ ] Does the layout require fine motor control (tiny answer boxes, complex + tables) not relevant to the construct? +- [ ] Is the font size and spacing accessible? +- [ ] Does the item require a specific device, plug-in, or interaction that not + every learner will have? + +## Timing + +- [ ] Is time pressure part of the construct? If not, is the time allowance + generous enough that speed doesn't determine the score? +- [ ] Are extra-time arrangements possible within the format as designed? + +## Prior attainment interactions + +- [ ] Does the item require a prerequisite skill that isn't the target and isn't + secure for everyone? (E.g. an algebra item that fails on arithmetic; a + history source item that fails on unfamiliar dating conventions.) +- [ ] If so — is that intentional and stated, or an accident? + +## Representation + +- [ ] Across the whole assessment, do names, contexts and people reflect the + range of learners taking it? +- [ ] Is any group represented only in a stereotyped role? +- [ ] Would any item cause distress in a way unrelated to its purpose? + (Bereavement, family separation, violence, food insecurity, migration — + these appear in "realistic" word problems more often than authors notice.) + +## Response + +- [ ] Is the expected form of the answer unambiguous? +- [ ] Would a correct but unexpected answer be marked wrong by the mark scheme + as written? +- [ ] Can a learner who knows the content but writes slowly still demonstrate it? + +## The overall read + +After the item-level pass, ask the two whole-assessment questions: + +1. **Construct under-representation** — does this set cover the domain, or the + convenient corner of it? Compare the blueprint against the actual curriculum + outcomes, not against the items you found easy to write. +2. **What will the scores be used for, and is the assessment strong enough to + carry that use?** Six items cannot support an individual placement decision. + A single essay cannot support a reliable grade. Say so before the scores exist, + not after someone has acted on them. diff --git a/skills/designing-assessments/references/item-types.md b/skills/designing-assessments/references/item-types.md new file mode 100644 index 0000000..755ef26 --- /dev/null +++ b/skills/designing-assessments/references/item-types.md @@ -0,0 +1,42 @@ +# Choosing the response format + +Choose per learning outcome, not per assessment. Mixed-format assessments are +normal and usually better. + +| Format | Shows you | Costs | Use when | +|---|---|---|---| +| **Multiple choice (diagnostic)** | Which misconception is present, across the whole class, in seconds | High design cost; can't show reasoning | You need whole-class diagnosis fast; misconceptions are known and enumerable | +| **Multiple choice (recall)** | Whether a fact is retrievable | Low value formatively | Retrieval practice, prerequisite checks | +| **Short constructed response** | Reasoning, in the learner's own terms | Marking time; marker variance | The reasoning *is* the outcome | +| **Show-your-working problems** | Where in a procedure it breaks | Marking time | Procedural outcomes with multiple steps | +| **Extended writing / essay** | Synthesis, argument, structure | Low reliability; heavy marking | The construct genuinely requires it — never as a default | +| **Performance / practical** | Skill in real conditions | Very low reliability; high logistics | The outcome is a performance | +| **Ranking / sorting / matching** | Discrimination between near-cases | Guessable if few items | Classification and category outcomes | +| **Concept map** | Structure of the learner's knowledge, including missing links | Hard to mark reliably | You want to see organisation, not recall | +| **Confidence-weighted response** | Where learners are confidently wrong (the dangerous case) | Needs teaching; can penalise the cautious | Diagnosing entrenched misconceptions | +| **Two-tier item** (answer + reason) | Right answer for the wrong reason | Doubles the item count | Any topic where guessing is plausible | + +## Two-tier items — worth calling out + +Tier 1: the answer. Tier 2: "which of these best explains your answer?" +Four response patterns, four different actions: + +| Tier 1 | Tier 2 | Meaning | Action | +|---|---|---|---| +| ✔ | ✔ | Secure | Move on | +| ✔ | ✘ | Right for the wrong reason — invisible to conventional marking | Reteach the mechanism | +| ✘ | ✔ | Slip or procedural error, understanding intact | Practice, not reteaching | +| ✘ | ✘ | Not yet | Reteach from the start | + +The ✔/✘ cell is the one that makes two-tier items worth the extra design effort: +it is the failure mode ordinary tests systematically hide. + +## When *not* to assess + +Some things are better observed than assessed: + +- Collaboration, resilience, curiosity — assessment instruments for these are + mostly measuring compliance or confidence. Observe and describe; do not score. +- Very early-stage understanding — assessing before there is anything to assess + produces a score that says only "this hasn't been taught yet". +- Anything where the result would change nothing. Cut it and reclaim the time. diff --git a/skills/designing-assessments/references/item-writing.md b/skills/designing-assessments/references/item-writing.md new file mode 100644 index 0000000..9f52452 --- /dev/null +++ b/skills/designing-assessments/references/item-writing.md @@ -0,0 +1,108 @@ +# Item-writing rules + +## Stems + +1. **Put the whole question in the stem.** A learner should be able to answer + before seeing the options. Test: cover the options — is it still a question? +2. **No negatives unless unavoidable**, and if unavoidable, emphasise: "Which is + **not** a…". Double negatives are a fairness fault. +3. **Strip decorative context.** Every word of scenario that isn't doing + construct work is reading load taxing the weakest readers hardest. +4. **One idea per item.** Compound items ("explain X and evaluate Y") make the + response uninterpretable — you cannot tell which part failed. +5. **Avoid clue words**: "always", "never", "all" in a distractor is a giveaway; + "usually", "may" in the key is a giveaway the other way. + +## Options + +6. **Parallel in length, structure and grammar.** The longest, most-qualified + option is the key far too often. Balance them deliberately. +7. **3–4 options is enough.** A fifth plausible distractor is rarely findable; + padding with an implausible one just raises the guess rate on the remaining set. +8. **No "all of the above" / "none of the above".** Both are answerable by partial + knowledge and neither is diagnostic. +9. **Options mutually exclusive and homogeneous** — same category, same grain. +10. **Key position randomised** across the set. Check the distribution; models + over-produce option B and C. + +## Distractors — the part that actually matters + +A distractor is a **hypothesis about how a learner might be wrong**. Build it in +this order: + +1. Write the misconception in plain language: *"Learners think multiplication + always makes a number bigger."* +2. Work the item *as a learner holding that misconception would*. +3. That result is the distractor. +4. Record the mapping in the interpretation guide. + +Sources of real misconceptions, in order of preference: + +- **Learners' own prior work.** Mine wrong answers from previous cohorts — if the + Nurture MCP is connected, `get_submissions` / `get_submission_detail` give you + actual wrong answers, which beat any invented distractor. +- **Published diagnostic banks** (e.g. the Diagnostic Questions collections in + maths and science, awarding-body examiner reports, national assessment + frameworks' common-error sections). +- **Research literature on misconceptions** in the specific domain. +- **Derived from the procedure**: enumerate the steps and generate the error at + each — sign error, wrong order of operations, off-by-one, unit not converted. + This is systematic and beats free-associating "plausible" wrong answers. + +Test each distractor: **who would choose this, and what do they believe?** If you +cannot answer, replace it. + +## Worked distractor construction — example + +*Item target:* dividing by a fraction. + +`What is 6 ÷ ½ ?` + +| Option | Misconception behind it | +|---|---| +| 3 | "Dividing makes things smaller" → learner halves instead | +| 12 | **KEY** | +| 6.5 | Reads the operation as addition of the fraction | +| 2 | Divides by 2 and then... (weak — replace) | + +The fourth is doing no work. Better: `1/12` — learner inverts the *dividend* +rather than the divisor, a real and common slip. Now every wrong answer names a +specific thing to reteach. + +## Constructed-response items + +11. **Specify the response length and form** in the stem: "in one sentence", + "state two reasons", "show your working". Ambiguity about form produces + variance that isn't about knowledge. +12. **Write the mark scheme at the same time as the item.** If you cannot write + the mark scheme, the item is ambiguous. +13. **Include acceptable alternatives.** Mark schemes that list only the expected + answer punish learners who are right in an unexpected way. +14. **Ask for the reasoning where the reasoning is the target.** "Which is + larger, 3/7 or 4/9? How do you know?" — the second sentence is the item. + +## Items that reveal thinking + +The highest-value formative items make the *reasoning* visible, not the answer: + +- **Justify a choice**: "Sam says X. Do you agree? Explain." +- **Spot the error**: give worked reasoning with one flawed step; the learner + identifies and repairs it. +- **Two things the same, one different**: "Which is the odd one out, and why?" + (multiple defensible answers — the reasoning is assessed, not the choice.) +- **Always / sometimes / never true**: "Squaring a number makes it larger." +- **What's the question?** Give the answer, ask for a question it answers. +- **Compare two responses**: "Which of these two answers is better, and why?" +- **Predict then check**: commits the learner before the evidence arrives, which + is what makes a demonstration change anyone's mind. + +## Retrieval practice sets + +If the purpose is durable memory rather than diagnosis, the design rules change: + +- Low stakes, frequent, cumulative — draw from all prior units, not the last one. +- Space the repetitions; interleave related-but-distinguishable topics. +- Answers checked and corrected immediately. +- No new content, and no time pressure that turns it into a test. +- Frame explicitly as learning, not surveillance — that framing also determines + whether learners answer honestly. diff --git a/skills/eliciting-evidence/SKILL.md b/skills/eliciting-evidence/SKILL.md new file mode 100644 index 0000000..5a952e3 --- /dev/null +++ b/skills/eliciting-evidence/SKILL.md @@ -0,0 +1,113 @@ +--- +name: eliciting-evidence +description: Read this when the question is how to find out what a class is thinking during a lesson, rather than what to ask them. Use it when asked to improve questioning, check understanding mid-lesson, run a discussion, involve quieter students, or design an exit ticket or do-now — and when a teacher says the class "seems to get it" but the results say otherwise. It produces routines where every learner commits to an answer the teacher can read at a glance. +--- + +# Eliciting evidence of learning + +Wiliam's second key strategy, delivery side. `designing-assessments` and +`hinge-questions` cover *what* to ask; this covers *how*, so the answer tells you +about the class and not about the three learners with their hands up. + +## The problem + +Teacher asks → volunteers raise hands → teacher picks one → teacher evaluates. That +routine is a machine for producing the *appearance* of understanding. It samples +the confident, rewards speed over thought, and lets most of the room opt out of +thinking. Wiliam's phrase for the alternative: **basketball, not ping-pong**. + +Every question should serve one of two purposes, and you should be able to say +which: **to cause thinking**, or **to give the teacher information**. Serving +neither makes it filler. + +## What done looks like + +A routine specified tightly enough to run tomorrow: what's asked, how every +learner responds, how long the think time is *in seconds*, what the teacher does +at each outcome, and what the whole thing costs in minutes. + +## Constraints + +- **Every learner produces a response**, not just volunteers. +- **Everyone commits before seeing anyone else's answer.** Cards at chest height, + boards up on a count, written commitment before movement. +- **The teacher reads the room at a glance.** If you have to walk around to see + the answers, that's circulation — useful, but not an all-student response system. +- **Think time is stated in seconds.** Rowe (1974): teachers typically wait under + one second; three seconds lengthens responses and increases how many learners + respond at all. The neglected half is wait time *after* the answer. +- **Nothing here is graded.** The moment a check carries marks, learners answer to + look right rather than to reveal what they think, and the instrument stops working. +- **The time cost is stated and proportionate.** A twelve-minute routine in a + fifty-minute lesson has to earn it. +- **Self-reports are labelled as self-reports.** Traffic lights, thumbs and + reaction spectra measure how learners *feel*, which correlates weakly with what + they know — worst for the learners furthest behind, who don't know they don't + know. Use them to start a conversation, never as evidence of understanding. Say + this every time you recommend one. + +## The four levers + +**Wait time** — free, immediate, benefits everyone. For teachers who find the +silence unbearable: announce it. "I'll ask, then we all wait ten seconds." + +**No hands up** — selection, random or deliberate, instead of volunteering. Two +companion rules or it fails: "I don't know" is not an exit (come back to them +after two others and have them choose between the answers), and never select +without think time first — that's an ambush, and the least confident learners +experience it as one. + +**All-student response** — mini whiteboards, ABCD cards, hand signals, digital +polling, everybody-writes. The biggest single change available to most classrooms. + +**Structured discussion** — think–pair–share (don't skip the silent stage, or the +faster partner does the thinking), pose–pause–pounce–bounce, say-it-better, +agree/disagree/build, random reporter named *after* the discussion. + +`references/technique-bank.md` has each with its setup cost and its specific +failure mode, plus a cost-to-value ordering for teachers who can only change one +thing. + +## Exit tickets + +The highest-value five minutes in a lesson — *if* they're read before the next +one. One or two questions on the lesson's core idea; sortable into piles in under +three minutes for a class of thirty; **the pile determines tomorrow's starter**. If +the result won't change the next lesson, don't collect it. + +## Judgement calls + +- **What do you need to find out, and what would you do differently either way?** + If there's no branch, the routine is theatre. +- **Is this a move-on-or-reteach decision?** Then it's a hinge question — route. +- **One technique, not five.** Recommend a single routine and let it bed in for a + term. Five new routines at once produces none of them. + +## Live session data + +If a Nurture Signal server is connected, a live session exposes engagement +reactions, comments, questions and participation. Discover the available tools +rather than assuming names. + +What it licenses: + +| Data | Tells you | Does **not** tell you | +|---|---|---| +| Engagement reactions | Self-reported energy, in real time, from everyone including the silent | Whether anyone understands | +| Comments | What some learners chose to say | What the non-commenters think | +| Questions | Where confusion is surfacing, and its shape | How widespread it is | +| Quiz / poll responses | Actual understanding, if the items are diagnostic | Anything the items didn't ask | +| Participation counts | Who is responding | Who is learning | + +Use reactions and comments to decide **where to look**; use a diagnostic question +to find out **what's going on**. Reporting a reaction distribution as a measure of +understanding is the standard way engagement dashboards mislead — say so. + +## Failure modes + +- A great question with a poor sampling method. The commonest one. +- Think time announced but not honoured. Three seconds feels like thirty. +- "Any questions?" — elicits nothing. Ask something specific, or "write down the + one thing you're least sure about". +- Rhetorical questions counted as checks. "Everyone happy with that? Good." +- Collecting evidence and not acting on it. diff --git a/skills/eliciting-evidence/references/technique-bank.md b/skills/eliciting-evidence/references/technique-bank.md new file mode 100644 index 0000000..c612a54 --- /dev/null +++ b/skills/eliciting-evidence/references/technique-bank.md @@ -0,0 +1,146 @@ +# Technique bank + +Drawn largely from Wiliam & Leahy, *Embedding Formative Assessment* (2015). +Each entry: what it is, the setup cost, and the failure mode to warn about. + +Recommend **one or two** of these to a teacher, not the list. Wiliam's own advice +is that a teacher should adopt a small number of techniques and run them for a +term before adding more. + +--- + +## Getting a response from everyone + +**Mini whiteboards.** Everyone writes, everyone holds up on a count. +*Setup:* boards, pens, erasers; a routine for distribution. +*Fails when:* learners hold up late and copy. Fix with "3–2–1–show". + +**ABCD cards.** Laminated letter cards or a four-way fan. +*Setup:* one set per learner, kept in books. +*Fails when:* held overhead so everyone sees everyone. Chest height. + +**Four corners.** Each corner of the room is an option; learners move. +*Setup:* none. Good for opinion questions and for waking a class up. +*Fails when:* learners follow friends. Require a written commitment first. + +**Everybody writes.** Two minutes of silent writing before any discussion. +*Setup:* none. The cheapest upgrade to any discussion in this list. +*Fails when:* the teacher shortens it because the room is quiet. + +**Exit ticket.** One or two questions on a slip at the end. +*Setup:* slips; three minutes of lesson time. +*Fails when:* not read before the next lesson. Then it is pure cost. + +**Entry ticket / do-now.** Same, at the start, on prior content. +*Setup:* none. Doubles as retrieval practice. +*Fails when:* it becomes a settling activity rather than a diagnostic. + +--- + +## Choosing who speaks + +**No hands up.** Selection by the teacher, random or deliberate. +*Setup:* a randomiser — lolly sticks, cards, an app. +*Fails when:* used without think time. That is an ambush, not a question. + +**Random reporter.** Groups discuss; reporter named afterwards. +*Setup:* none. Very high accountability for group work. +*Fails when:* the same confident learner briefs the group anyway. Rotate roles. + +**Phone a friend, then evaluate.** A stuck learner may nominate someone, but must +then say whether they agree and why. +*Setup:* none. Keeps "I don't know" from being an exit. + +**Hot seat.** Sustained questioning of one learner, class summarises after. +*Setup:* none, but needs a secure classroom climate. +*Fails when:* the rest of the class disengages. The summarise-after step is what +prevents that — don't drop it. + +--- + +## Improving the answer + +**Pose–pause–pounce–bounce.** Ask, wait, select, then pass the answer to someone +else to extend or challenge. +*Fails when:* the bounce is only ever to a stronger learner, which the class notices. + +**Say it better.** A second learner restates a roughly-right answer more precisely. +*Fails when:* it reads as correcting the first learner. Frame it as a class +drafting exercise, and use it on good answers too. + +**Agree / disagree / build.** Every contribution must open with one of the three +and reference the last speaker. +*Setup:* display the three stems. +*Fails when:* it becomes ritual. Rotate out after a few weeks and back later. + +**Wait time on both sides.** Three seconds after the question; three after the answer. +*Fails when:* only the first half is done. + +**Statistics on wrong answers.** "Last year 40% chose X. Why?" +*Fails when:* the number is invented. Only use real data. + +--- + +## Self-report (use with care) + +**Traffic lights.** Red / amber / green against a criterion. +*Use:* to pair greens with reds for peer explanation — that pairing is the value. +*Warning:* it is a confidence report, not evidence of understanding, and it is +least accurate for the learners furthest behind. Never plan from it alone. + +**Fist to five.** 0–5 fingers of confidence. Same warning. + +**Learning-intention self-check.** Learners rate their work against the criteria +before submitting. +*Use:* the gap between their rating and the teacher's is itself diagnostic — a +learner who is confidently wrong needs something different from one who knows +they are stuck. + +--- + +## Making thinking visible over a lesson + +**Two-column notes / prediction log.** Learners commit to a prediction before a +demonstration. Committing first is what makes anyone update their beliefs when the +result contradicts them. + +**Concept map at the start and end.** The difference is the evidence. + +**Question shells.** Reusable frames the class applies to any content: "What would +change if…?", "How is this the same as… and different from…?", "What would have to +be true for this to be false?", "Which of these is the odd one out?" +*Use:* teaches the class to generate questions, not only to answer them. + +**Learners write the next question.** After a topic, learners write the exam +question they think is coming, plus the mark scheme. Reveals what they think the +important thing is, which is frequently not what you taught. + +--- + +## Whole-class routines with a long payoff + +**The secret student.** A randomly chosen, unnamed learner's conduct determines a +class reward; the name is revealed only if they earned it. Accountability without +public exposure. + +**Homework help board.** Learners post problems; other learners answer. Turns +homework into peer instruction and gives the teacher a live map of the sticking +points. + +**C3B4ME.** "See three before me" — three sources consulted before asking the +teacher. Frees teacher attention for genuine blocks and builds resourcefulness. +*Fails when:* used to deflect learners who genuinely need the teacher. Exempt +anyone who is stuck at the first step. + +--- + +## Cost-to-value ordering + +If a teacher can only do one thing, in rough order of value per unit of effort: + +1. **Wait time** — free, immediate, everyone benefits. +2. **Mini whiteboards** — cheap, transforms what the teacher can see. +3. **No hands up** with think time — free, changes who does the thinking. +4. **Exit tickets that determine the next starter** — five minutes, closes the loop. +5. **Hinge question at one decision point per lesson** — high design cost, highest + diagnostic value. diff --git a/skills/formative-assessment/SKILL.md b/skills/formative-assessment/SKILL.md new file mode 100644 index 0000000..3c0604c --- /dev/null +++ b/skills/formative-assessment/SKILL.md @@ -0,0 +1,101 @@ +--- +name: formative-assessment +description: Read this when a teaching request touches assessment, feedback, or checking understanding, and you need to know which of the eight Nurture skills applies. Use it when the request is broad ("help me plan this unit", "how do I know they've got it?"), spans several skills, or when the user asks about assessment for learning or responsive teaching generally. It routes to the right skill and carries the shared Dylan Wiliam framework the others assume. +--- + +# Formative assessment (Dylan Wiliam) + +The router for this plugin. Read the routing table, go to the skill that does the +work. Come back here only for the shared framework. + +## Routing + +| The request | Skill | +|---|---| +| Comment on / mark / respond to student work; feedback or marking policy | `giving-feedback` | +| Build a quiz, test, task, item bank; check an assessment measures what it claims | `designing-assessments` | +| A mid-lesson question that decides whether to move on or reteach | `hinge-questions` | +| Learning intentions, success criteria, rubrics, exemplars | `success-criteria` | +| Questioning, all-student response, exit tickets, discussion | `eliciting-evidence` | +| Peer feedback, self-assessment, reflection prompts, student ownership | `peer-and-self-assessment` | +| Interpret a class set and decide what to teach next | `analysing-student-work` | + +Spanning request ("plan this unit formatively")? Work the strategies in order — +goals, then evidence, then feedback, then peer and learner agency — invoking each +skill as you reach it. Don't try to do their work here. + +## The framework the other skills assume + +Formative assessment is not a kind of test. It is **evidence about learning being +elicited, interpreted and acted on to make better decisions about what to do next** +than would have been made without it (Black & Wiliam, 2009). + +Wiliam's own caveat, worth repeating to sceptical teachers: "formative assessment" +is a badly chosen name because people hear "a test". He prefers **responsive +teaching**. + +Three questions × three agents: + +| | Where the learner is going | Where they are now | How to get there | +|--------------|---------------------------|--------------------|------------------| +| **Teacher** | Clarifying learning intentions and success criteria | Engineering discussions and tasks that elicit evidence | Providing feedback that moves learners forward | +| **Peer** | ← Understanding intentions and criteria → | ← Learners as instructional resources for one another → | | +| **Learner** | | ← Learners as owners of their own learning → | | + +Collapsed: the **five key strategies**. One big idea underneath — *use evidence of +learning to adapt teaching to meet learner needs*. + +Most classrooms are strong on the teacher row and weak on the other two. When +diagnosing a practice, check the rows before the columns. + +## What good output looks like, across every skill + +1. **Evidence changes a decision.** If nothing would be done differently depending + on the result, it isn't formative — say so and cut it. +2. **The learner has something specific to do**, and time allocated to do it. +3. **Wrong answers are diagnostic** — they say *which* wrong idea is held. +4. **The learning is named, not the activity.** +5. **Everyone is sampled, not the volunteers.** +6. **Claims are made at the strength the evidence supports** — including out loud, + to the teacher, when it's weaker than they'd like. + +## Recommend one thing + +Wiliam's practical counsel is that a teacher should adopt one or two techniques +and run them for a term, not adopt fifteen. When someone asks "what should I +change?", answer with **one** change and say why that one. A list of twelve +techniques reads as thorough and produces nothing. + +## Ask before assuming + +You will usually be missing something only the teacher has. Ask for it rather than +inventing it: + +- What are the success criteria, and have the learners seen them? +- What's the age group, subject, and curriculum or specification? +- What decision is this for — and what would you do differently either way? +- What have you already tried? + +If the teacher is mid-flow, take your best guess, **state the assumption in one +line**, and carry on. Don't block on a question you can reasonably answer yourself. + +## Real learner work beats invented examples + +If a Nurture MCP server is connected, it can reach classes, assignments, +submissions and reflections; the Nurture Signal server reaches live-session +reactions, comments, questions and analytics. Discover the tools actually +available rather than assuming names — they change. + +Pull the real work before advising. Then: + +- Never name a learner in anything a class will see. Aggregate or anonymise. +- Say what the data can't tell you. Engagement reactions measure self-reported + energy, not understanding — never report them as evidence of learning. +- If nothing is connected, say so once and work from the teacher's description. + Never fabricate class statistics. + +## References + +- `references/five-strategies.md` — each strategy expanded, with its techniques. +- `references/evidence-base.md` — the studies, what they found, and where the + common overclaims are. Read before citing a number. diff --git a/skills/formative-assessment/references/evidence-base.md b/skills/formative-assessment/references/evidence-base.md new file mode 100644 index 0000000..3767563 --- /dev/null +++ b/skills/formative-assessment/references/evidence-base.md @@ -0,0 +1,140 @@ +# Evidence base — what the studies actually found + +Use this when a user asks "what's the evidence for this?" or challenges a claim. +State findings at the strength the study supports, and volunteer the limits. An +overclaim here costs credibility with exactly the audience that matters. + +--- + +## Sadler (1989) — the theoretical spine + +Sadler, D. R. "Formative assessment and the design of instructional systems", +*Instructional Science* 18, 119–144. + +Three conditions for a learner to improve: possess a concept of the goal or +standard; compare their actual performance with that standard; engage in action +that closes the gap. This is the reason strategy 1 comes first and the reason +feedback without a shared standard fails. + +--- + +## Black & Wiliam (1998) — *Inside the Black Box* / "Assessment and Classroom Learning" + +Review of 250+ studies. Reported effect sizes in the range 0.4–0.7 for formative +assessment interventions. + +**Be honest about this one.** Wiliam himself has repeatedly said the 0.4–0.7 range +should not be treated as a precise estimate — it aggregates very heterogeneous +interventions and outcome measures, and he has described quoting it as a number as +a mistake. Cite it as "a substantial body of evidence that classroom assessment +practices affect learning", not as "formative assessment has an effect size of +0.7". If a user quotes 0.7 at you, correct it gently. + +--- + +## Kluger & DeNisi (1996) — feedback is not reliably good + +*Psychological Bulletin* 119(2), 254–284. 131 experiments, 607 effect sizes, +23,663 observations. + +- Mean effect of feedback interventions: **d ≈ 0.41**. +- In **38%** of cases, the feedback intervention **lowered** performance. +- Their Feedback Intervention Theory: performance degrades when feedback moves + attention away from the task and toward the self. + +This is the single most important result for the `giving-feedback` skill. The +practical consequence: praise, grades, comparative ranking and anything that +invites "how good am I?" rather than "what do I do next?" are live risks, not +neutral additions. + +Caveat: largely non-classroom studies (lab and workplace tasks). The mechanism +generalises; the magnitude should not be transplanted uncritically. + +--- + +## Butler (1988) — grades crowd out comments + +"Enhancing and undermining intrinsic motivation: the effects of task-involving and +ego-involving evaluation on interest and performance", *British Journal of +Educational Psychology* 58, 1–14. 132 Year 5–6 students, three conditions. + +- **Comments only** → roughly 30% gain in subsequent performance, and interest + maintained. +- **Grades only** → no gain. +- **Grades and comments** → no gain; interest patterned like grades-only. + +The finding that changed practice: adding a grade to a comment does not add +information, it subtracts attention. + +**Caveat worth stating.** Guskey (2019) has criticised how widely this single +small study is generalised: the comparison was task-involving comments against +ego-involving grades, so form and substance are confounded, and the sample was +narrow. The defensible claim is that *ego-involving evaluation undermines +task-focused response to feedback*, not that "grades are always harmful". +Comment-only marking is well-supported as a default, not as a law. + +--- + +## Hattie & Timperley (2007) — the four levels of feedback + +*Review of Educational Research* 77(1), 81–112. Feedback operates at four levels: + +1. **Task** — how well the task was performed. Effective, especially for novices. +2. **Process** — the strategies used. Most powerful for deepening learning. +3. **Self-regulation** — the learner's monitoring and direction of their own + effort. Powerful where learners are capable of it. +4. **Self** — praise directed at the person ("you're clever"). Generally + ineffective for learning, and dilutes any task feedback bundled with it. + +Practical rule: aim at process and self-regulation; use task-level for novices; +keep self-level out of the feedback entirely (which does not mean being cold — +warmth belongs in the relationship, not in the feedback text). + +--- + +## Nyquist (2003) — stronger feedback beats weaker feedback + +Meta-analysis of 185 studies, ordering feedback types by strength: + +- weakest: knowledge of results (right/wrong) only +- then: knowledge of correct response +- then: elaborated feedback explaining why +- strongest: feedback plus an explicit opportunity and requirement to act on it + +Effect sizes rise across that ordering. This is the empirical case for building +response time into the lesson rather than treating the written comment as the +deliverable. + +--- + +## Rowe (1974) — wait time + +Mean teacher wait time after a question is under one second. Extending to ~3 +seconds increases response length, the number of learners volunteering, the +incidence of learner-to-learner exchange, and the frequency of speculative +reasoning. Wait time *after* a learner's answer is the more neglected of the two. + +--- + +## Roediger & Karpicke (2006) and the testing effect + +Retrieval practice produces more durable learning than restudy. Relevant here +because it means low-stakes questioning has a double payoff: it gives the teacher +evidence *and* it strengthens memory. Frame frequent low-stakes quizzing as +learning, not surveillance — that framing also determines whether learners answer +honestly. + +--- + +## Where the evidence is thin + +Be candid about these when asked: + +- **Peer assessment quality** varies enormously with the protocol and the age of + the learners; the strong results come from tightly-structured protocols with + explicit criteria, not from "swap books and comment". +- **Self-assessment accuracy** is poor without exemplars, and correlates with + prior attainment — the learners who most need it are worst at it. +- **Learning styles** have no support; if a user raises them, say so plainly. +- **Digital dashboards** measure engagement proxies (clicks, reactions, time) + which are not measures of learning. Say what the metric can and cannot license. diff --git a/skills/formative-assessment/references/five-strategies.md b/skills/formative-assessment/references/five-strategies.md new file mode 100644 index 0000000..319d935 --- /dev/null +++ b/skills/formative-assessment/references/five-strategies.md @@ -0,0 +1,131 @@ +# The five key strategies, expanded + +Source framing: Wiliam, D. (2011/2018) *Embedded Formative Assessment*; Wiliam, D. +& Leahy, S. (2015) *Embedding Formative Assessment*; Black, P. & Wiliam, D. (2009) +"Developing the theory of formative assessment", *Educational Assessment, +Evaluation and Accountability* 21(1). + +--- + +## Strategy 1 — Clarifying, sharing and understanding learning intentions and success criteria + +**The problem it solves.** Sadler (1989) argued that for a learner to improve, +three conditions must hold: they must possess a *concept of the goal*, be able to +*compare their current level* against it, and be able to *take action to close the +gap*. Strategy 1 is condition one. Without it, feedback has nothing to attach to. + +**Core distinctions.** + +- *Learning intention* vs *context*: "to understand how writers use setting to + create mood" is the intention; "in the opening of *Great Expectations*" is the + context. Learners who only meet the contextualised version tend to encode the + context and fail to transfer. +- *Product criteria* ("your conclusion states which hypothesis the data support") + vs *process criteria* ("you checked each calculation with an estimate first"). + Novices usually need process criteria; product criteria alone leave them + guessing at the route. +- *Criteria* vs *exemplars*: for anything with genuine quality (writing, argument, + design), criteria are necessarily vague — "well structured" means nothing to + someone who cannot yet do it. Exemplars carry the meaning that words cannot. + +**Techniques.** Sharing work samples of varying quality and asking learners to +rank them; co-constructing criteria from those samples; "what would a really good +one look like"; choosing the best of two anonymous responses and justifying; +learners generating the rubric before attempting the task; "WALT/WILF" framings +(use sparingly — they often degrade into restating the task). + +--- + +## Strategy 2 — Engineering effective discussions, tasks and activities that elicit evidence of learning + +**The problem it solves.** Teachers routinely overestimate understanding because +they sample from volunteers, and because questions are pitched to confirm rather +than to probe. + +**Core moves.** + +- *Question design before question delivery.* Wiliam's point is that the quality + of the question matters more than the technique used to ask it. Two purposes + only: to **cause thinking**, and to **provide information to the teacher**. +- *All-student response systems.* Mini whiteboards, ABCD/multiple-choice cards, + hand signals, exit tickets, digital polling. The design constraint is that the + teacher can read the whole room at once. +- *No hands up (except to ask a question).* Random or deliberate selection so the + sample is not self-selecting. +- *Wait time.* Rowe's work found teachers typically wait under a second after + asking; extending to around three seconds lengthens and improves responses, and + increases the number of learners who respond at all. Wait time *after* the + learner's answer matters as much as before. +- *Hinge questions.* A single diagnostic item at a decision point in the lesson. + See the `hinge-questions` skill. + +See the `eliciting-evidence` skill for the full technique bank. + +--- + +## Strategy 3 — Providing feedback that moves learners forward + +**The problem it solves.** Feedback is not reliably helpful. Kluger & DeNisi's +1996 meta-analysis found an average effect of d ≈ 0.41 across 607 effect sizes, +but that feedback interventions *lowered* performance in around 38% of cases — +most often when feedback directed attention to the self rather than the task. + +**The design rules.** + +1. Feedback should be **more work for the recipient than the donor**. +2. Feedback should **cause thinking** — it is a recipe for future action, not a + post-mortem. +3. Feedback should focus on the **task and the process**, not the person. +4. Feedback must be **timed so the learner still has room to act**, and lesson + time must be allocated for acting on it. +5. Grades and comments together behave like grades alone (Butler, 1988) — the + grade crowds out the comment. Where the system permits, separate them in time. + +See the `giving-feedback` skill. + +--- + +## Strategy 4 — Activating learners as instructional resources for one another + +**The problem it solves.** One teacher cannot give thirty learners timely +feedback. Peers can — and the act of assessing someone else's work against +criteria is itself a powerful way to internalise the criteria. + +**Conditions that make it work.** Peer assessment against *explicit criteria*, +directed at *the work* not the person, with a *protocol* that constrains the +response format. Ungoverned peer feedback degrades into praise or personal +comment. + +**Techniques.** Two stars and a wish; pre-flight checklist (a peer signs off the +checklist before work is handed in); C3B4ME ("see three before me"); homework help +board; peer editing with a specified focus; error-spotting in a deliberately +flawed exemplar; group reporting-back where the reporter is chosen at random after +the discussion. + +--- + +## Strategy 5 — Activating learners as owners of their own learning + +**The problem it solves.** The learner is the only person present for all of their +own learning. Self-regulation is the highest-leverage and slowest-building target. + +**Techniques.** Traffic lights / red-amber-green self-rating against criteria +(then pair greens with reds); learning portfolios; "plus, minus, equals" progress +tracking against one's own prior work; self-assessment before submission; learning +logs and structured reflection prompts; the "secret student" (a randomly chosen, +unnamed student whose conduct determines a class reward — accountability without +exposure). + +**Guardrail.** Self-assessment is unreliable when learners lack a concept of the +goal — which returns you to strategy 1. Self-assessment without exemplars mostly +measures confidence. + +--- + +## Sequencing advice + +Wiliam's practical counsel to schools is deliberately unambitious: teachers should +adopt **one or two techniques**, run them for a term, and only then add more. +Teacher learning communities that meet monthly and hold each other to a specific +committed change outperform one-off training. When advising, resist the urge to +hand over the whole list. diff --git a/skills/giving-feedback/SKILL.md b/skills/giving-feedback/SKILL.md new file mode 100644 index 0000000..e521dd8 --- /dev/null +++ b/skills/giving-feedback/SKILL.md @@ -0,0 +1,110 @@ +--- +name: giving-feedback +description: Read this before writing any comment, mark, or response on a piece of student work — your default behaviour here is the documented failure mode. Use it when asked to mark, comment on, respond to or improve feedback on student work, a submission, or a class set, and when designing a marking or feedback policy. It produces feedback that names one gap and gives the learner a specific task, instead of prose that describes the work. +--- + +# Giving feedback that moves learners forward + +Wiliam's third key strategy. Feedback exists to change what the learner does next. +Not to describe the work, not to justify a grade. + +## Why the default is wrong + +Feedback is not reliably beneficial. Kluger & DeNisi (1996) found feedback +interventions *lowered* performance in 38% of cases, principally when attention +moved from the task to the self. Fluent, comprehensive, encouraging feedback is +the shape that fails. + +## What done looks like + +A learner reads it and knows exactly what to do, and doing it takes them longer +than writing it took you. + +Concretely, the output has four parts: + +``` +**Against the criteria:** + +**The gap:** + +**Your task:** + +**Time needed:** +``` + +## Constraints — non-negotiable + +1. **More work for the recipient than the donor.** If your comment is longer than + the learner's required response, invert it. +2. **Cause thinking, not compliance.** Once you supply the answer, the learning is + gone. Locate the problem; don't correct it. Escalate specificity only as far as + this learner needs — "three agreement errors on this page, find them" beats + "line 14, check the verb", and both beat fixing it for them. +3. **Task and process, never the person.** Cut every clause about ability, effort, + attitude or character — *including the positive ones*. "You're really good at + this" is self-level feedback (Hattie & Timperley) and it competes with the task + feedback beside it. +4. **At most three points.** Usually one. The learner cannot act on twelve and will + act on none. Choose by leverage. +5. **Against criteria the learner has already seen.** If none exist, say what you + inferred; offer `success-criteria`. Never mark against a private standard. +6. **Comments, not comments-and-a-grade.** Butler (1988): grades plus comments + performed like grades alone. Where a grade is required, release it after the + learner has responded. +7. **Response time is allocated**, or the feedback is decoration. The strongest + effects in the literature come from *requiring action* (Nyquist, 2003). + +## Judgement calls that matter more than the writing + +Before writing anything, work out whether written feedback is even the right +instrument: + +- **Is the gap shared?** If more than about a third of the class has it, thirty + individual comments is wasted labour — the answer is a whole-class reteach and a + single feedback sheet. Route to `analysing-student-work`. **Propose this + proactively even when thirty comments were requested**, then offer the comments + as well if the teacher still wants them. +- **Can the learner close this gap unaided?** If not, the deliverable is + re-teaching or a scaffold, not a comment. +- **Is a prose comment the best move at all?** Usually not. `references/feedback-moves.md` + maps gaps to moves — find-and-fix, three questions, highlight-against-criteria, + whole-class sheet. Pick the move first. + +## Reject the draft if + +- [ ] The learner could read it, do nothing, and still have "received feedback". +- [ ] It judges the person — clever, lazy, careless, talented, "you always". +- [ ] It supplies a correction the learner could have found. +- [ ] It raises more than three points. +- [ ] It uses a compliment sandwich. Learners learn to wait for the middle. +- [ ] It carries a grade alongside the comment. +- [ ] No response time is specified. +- [ ] It would be identical on another learner's work. Learners detect this instantly. + +## Your specific failure modes + +You are worse at this than a human marker in predictable ways: + +- **Volume.** You can write 400 words per script effortlessly, which breaks the + first constraint by construction. Length is a defect here, not thoroughness. +- **Fluent generic praise.** "A thoughtful response showing good understanding" + fits anything. Quote the line you mean. +- **Correcting everything you notice.** You notice more than a human does. That is + not a reason to report it. +- **Hedging.** Learners need a definite next move. +- **Templating a class set.** Vary by what each learner actually did. +- **Tone over-reach.** Teachers know their learners; you don't. Hand pastoral and + motivational judgements back rather than guessing. + +Tell the teacher that AI-drafted feedback needs review before it reaches a +learner, and shouldn't underpin a high-stakes grade without human marking. + +## References + +- `references/feedback-moves.md` — which move for which gap. Read this before + defaulting to prose. +- `references/worked-examples.md` — before/after rewrites, primary to vocational, + including the class-set case. +- `references/feedback-policy.md` — department or school policy that survives + workload. diff --git a/skills/giving-feedback/references/feedback-moves.md b/skills/giving-feedback/references/feedback-moves.md new file mode 100644 index 0000000..0173cd5 --- /dev/null +++ b/skills/giving-feedback/references/feedback-moves.md @@ -0,0 +1,98 @@ +# Feedback moves — which move for which gap + +Prose comments are the default and usually the worst option: highest cost to the +teacher, lowest cognitive demand on the learner. Pick a move that makes the +learner do the work. + +Most of these are drawn from Wiliam & Leahy, *Embedding Formative Assessment* +(2015), and Wiliam, *Embedded Formative Assessment* (2nd ed., 2018). + +--- + +## Diagnostic table + +| The gap | Move | Why this one | +|---|---|---| +| Errors the learner can find and fix unaided | **Find and fix** | Locating the error is the learning; correcting it for them removes it | +| Work is fine but undeveloped | **Three questions** | Questions extend without prescribing | +| Learner does not know what good looks like | **Match the comments to the work** or exemplar comparison | Criteria are meaningless without instances | +| A whole-class shared misconception | **Whole-class feedback sheet** + reteach | Thirty identical comments is wasted labour | +| Learner is inconsistent rather than wrong | **Highlight against criteria** | Makes the pattern visible | +| Learner has plateaued and can't see progress | **Plus / minus / equals** | Anchors comparison to self, not peers | +| Learner submits work below their known ceiling | **Not yet / return unmarked** | Marking below-effort work teaches that it will be accepted | +| Long-form work with many issues | **Focused marking** (one criterion only) | Concentrates attention; the rest is noise | +| Learner needs to internalise the standard | **Peer pre-flight checklist** | Moves quality control before submission | + +--- + +## The moves + +### Find and fix +State the *number* and the *region*, never the location. "There are five spelling +errors in this paragraph — find and correct them." Scale the region to the +learner: whole page → paragraph → line. Works for calculation slips, referencing, +agreement, unit errors, unsupported claims. + +### Three questions +Write three questions in the margin — no statements. The learner answers them in +writing and returns the work. Good questions: "What evidence would convince a +reader who disagreed with you?", "Why did you choose this method rather than +substitution?", "Which step here would break if the number were negative?" + +### Match the comments to the work +Give the class a set of anonymised extracts and a separate set of comments; the +learners match them. Forces engagement with criteria and with the reasoning behind +each judgement. Very effective as a lesson starter before returning their own work. + +### Highlight against criteria +Two colours: green where a criterion is met, amber where it is attempted but not +secured. No prose. The learner writes the improvement themselves. Cheap for the +teacher, and the learner must construct the diagnosis. + +### Plus / minus / equals +The learner compares this piece with their own previous piece: what is better (+), +what is worse (−), what is unchanged (=). Anchors comparison to self rather than +to peers, which keeps the feedback task-involving rather than ego-involving. + +### Whole-class feedback sheet +Instead of individual comments, one sheet: common strengths, common errors, three +misconceptions with worked corrections, spelling/vocabulary to fix, two model +extracts, and named "see me" list. Read at the start of the response lesson. This +is the highest workload-to-impact ratio move available and the one to recommend +when a teacher raises marking workload. + +### Not yet +Return work that falls below the learner's demonstrated capability unmarked, with +the criterion it failed and a resubmission deadline. Requires the teacher to have +credible evidence of the ceiling, and consistency — one exception collapses it. + +### Focused marking +Announce in advance the single criterion the work will be assessed against, and +assess only that. The rest genuinely goes unmarked. Reduces workload and increases +the learner's attention on the target. + +### Improvement time (DIRT) +Not a feedback move but the container all of them need: dedicated lesson time, +usually 10–20 minutes, in which the only activity is responding to feedback, in a +distinguishable pen or a separate section so the response is visible. Without +this, none of the above changes anything. + +### Coded marking +A shared code — `Sp` spelling, `//` new paragraph, `?` unclear, `E` needs +evidence, `^` word missing. The learner decodes and acts. Fast to write, forces +the learner to diagnose. Requires the code to be taught and displayed. + +--- + +## Moves to avoid, and what to say instead + +| Common practice | Problem | Replace with | +|---|---|---| +| Compliment sandwich | Learner discards the bread and mistrusts the praise | Criterion statement + gap + task | +| "Good work!" / ticks | Knowledge of results only — weakest feedback type (Nyquist) | Highlight against criteria | +| Grade plus comment | Grade crowds out the comment (Butler 1988) | Comment, response, then grade later | +| Correcting every error | Learner acts on none of them | Focused marking on one criterion | +| Rewriting the sentence for them | Removes the learning | Locate the region, require the rewrite | +| "Try harder" / "more effort" | Self-level; not actionable | Name the process step that was skipped | +| Comparative ranking, displayed | Ego-involving; harms the bottom half and the top | Plus / minus / equals against self | +| Feedback at end of unit | No opportunity to act | Feedback at a point where the learning is still live | diff --git a/skills/giving-feedback/references/feedback-policy.md b/skills/giving-feedback/references/feedback-policy.md new file mode 100644 index 0000000..cbb6811 --- /dev/null +++ b/skills/giving-feedback/references/feedback-policy.md @@ -0,0 +1,91 @@ +# Designing a feedback policy that survives contact with workload + +Use when a user asks about a department or whole-school marking policy, or says +some version of "the marking is killing us". + +## The framing to lead with + +Wiliam's position: the question is not "how much marking?" but "what proportion of +the feedback a learner receives actually changes what they do?" A policy that +mandates volume optimises the wrong variable. Every hour a teacher spends writing +comments no learner acts on is an hour not spent designing the next lesson — and +lesson design has the larger effect. + +So the design principle is: **minimise donor effort per unit of recipient action.** + +## The five policy decisions + +**1. Frequency and depth are traded, not stacked.** +Specify *deep marking* on a small number of pieces per term (2–3), and *light +verification* on everything else. Do not specify "every piece, every two weeks" — +that is how policies produce compliance marking. + +**2. Response time is mandated, marking format is not.** +The one thing worth making non-negotiable is that feedback is followed by +timetabled response time, visibly evidenced. Leave teachers free to choose whether +that feedback arrived as a comment, a code, a whole-class sheet, or verbally. + +**3. Grades are separated from comments.** +Where reporting requires grades, the policy should specify that the comment is +issued first and the grade released after the learner has responded. If grades +must appear on the work, they go in the mark book, not the margin. + +**4. Whole-class feedback is explicitly permitted.** +Many policies implicitly ban it by requiring individual annotation. Name it as a +legitimate form of feedback, with a template, and it becomes the default for +class-wide gaps — which is where most of the workload goes. + +**5. What is *not* marked is stated as clearly as what is.** +Focused marking only works if the teacher can announce "I am assessing paragraph +structure only" and be backed by policy when a book scrutiny finds unmarked errors. + +## Template — one page + +``` +Feedback policy: + +Purpose + Feedback exists to change what learners do next. Any feedback that does not + produce a visible learner response is a cost with no return. + +What teachers do + - Deep feedback on pieces per term, on assessment points listed in the + scheme of work. + - Every piece of deep feedback is followed by <20> minutes of timetabled + improvement time, in the same or next lesson. + - Feedback names at most three things and specifies the learner's task. + - Feedback is against criteria the learners have already seen. + - Whole-class feedback sheets are used where a gap is shared by a third or more + of the class. + - Grades are recorded centrally, not written on the work. + +What learners do + - Respond in so the response is visible. + - Self-assess against the criteria before submitting. + +What we do not do + - Comment on effort, attitude or ability. + - Correct errors learners can find themselves. + - Mark for the benefit of an observer. + +How we know it's working + - Book looks focus on the learner's response, not the teacher's annotation. + - Termly: sample 6 learners, ask them what their last feedback asked them to do. + If they cannot say, the policy is failing regardless of how the books look. +``` + +## The accountability trap + +Most bad marking policies exist to produce evidence for inspection rather than +learning for students. If a user is constrained by this, say so honestly and help +them design for both: visible learner response in a distinguishable pen is +*better* evidence of impact than dense teacher annotation, and can be defended as +such. Do not pretend the constraint isn't real. + +## Verbal feedback + +The most timely feedback is spoken, in the lesson, while the work is still being +made. Policies that only recognise written feedback push teachers away from the +higher-impact channel. If a policy needs verbal feedback evidenced, evidence it +through the learner's subsequent action, never through a "verbal feedback given" +stamp — that stamp is pure compliance cost. diff --git a/skills/giving-feedback/references/worked-examples.md b/skills/giving-feedback/references/worked-examples.md new file mode 100644 index 0000000..9833b90 --- /dev/null +++ b/skills/giving-feedback/references/worked-examples.md @@ -0,0 +1,130 @@ +# Worked examples — feedback rewrites + +Each pair shows a typical comment and a rewrite that satisfies the five rules. +Note in each case that the rewrite is *shorter* and the learner's job is *longer*. + +--- + +## 1. Secondary English — analytical paragraph + +**Criterion:** *Supports interpretation with precise textual evidence and explains +how the writer's choices create the effect.* + +> ❌ "A really thoughtful response — you clearly understand the character well. +> You've made some lovely points about Curley's wife. Try to use more quotations +> to back up what you say, and explain them in more detail. Also watch your +> paragraphing. Well done, keep it up! **6/10**" + +Problems: self-level opening and closing; four separate demands; "more detail" is +not actionable; the grade will absorb all the attention; the learner does nothing. + +> ✅ **Against the criteria:** Your second paragraph states an interpretation +> clearly — that loneliness makes her cruel. +> +> **The gap:** You assert the effect but never point at the words that create it. +> +> **Your task:** Choose the one line in the passage that best proves loneliness +> drives her cruelty. Rewrite paragraph 2 so the quotation is embedded mid-sentence, +> and add one sentence beginning "Steinbeck's choice of ___ makes the reader…". +> +> **Time needed:** 12 minutes, Thursday's improvement slot. + +--- + +## 2. Primary maths — subtraction with exchange + +**Criterion:** *Exchanges correctly across a zero.* + +> ❌ "5/8. You've made mistakes in questions 3, 6 and 8 — remember to borrow +> properly! See the example on the board." + +Problems: score first; "borrow properly" names no process; errors located exactly, +so nothing to find; no response required. + +> ✅ **Against the criteria:** Every question without a zero in the top number is +> correct — your method is secure there. +> +> **The gap:** Three answers are wrong. All three have a zero in the tens column. +> +> **Your task:** Find the three. For each one, write what the 0 becomes after you +> exchange, before you subtract. +> +> **Time needed:** 8 minutes, start of tomorrow's lesson. + +Note the deliberate move: the *pattern* is given (zero in the tens column) because +that is the learning; the *locations* are withheld because finding them is the work. + +--- + +## 3. Secondary science — practical write-up + +**Criterion:** *Conclusion is consistent with all the data collected.* + +> ❌ "Your conclusion is wrong — the rate did not increase at 60 °C, it decreased +> because the enzyme denatured. You should have said that. Otherwise a good +> write-up with a clear method." + +Problems: supplies the entire answer, so the learner has nothing left to think +about; "otherwise good" is unattached to any criterion. + +> ✅ **Against the criteria:** Your method is reproducible — someone else could run +> this from your description. +> +> **The gap:** Your conclusion fits four of your five results. +> +> **Your task:** Which result does your conclusion fail to explain? Write two +> possible explanations for that result, then say which one your data support and +> why. +> +> **Time needed:** 15 minutes. + +--- + +## 4. Vocational / coursework — portfolio evidence + +**Criterion:** *Evidence is mapped to the assessment criterion it claims to meet.* + +> ❌ "Not enough evidence here for criterion 2.3. Add more." + +Problem: the learner does not know what would count, which is the actual gap. + +> ✅ **Against the criteria:** Your photographs and witness statement fully cover +> 2.1 and 2.2. +> +> **The gap:** For 2.3 you have shown *that* you did the check, not *how you +> decided* it had passed. +> +> **Your task:** Look at the exemplar in the shared folder (marked "2.3 secure"). +> List the two things it does that yours doesn't, then add them. +> +> **Time needed:** 20 minutes in Friday's workshop. + +--- + +## 5. Class set — where individual feedback is the wrong tool + +A teacher asks for feedback on 28 essays. 19 have the same failure: they narrate +the source rather than evaluating its reliability. + +Do **not** produce 28 comments. Produce: + +> **Whole-class feedback (teacher use)** +> +> *Shared gap (19/28):* Source paragraphs summarise content instead of evaluating +> reliability. This is the reteach. +> +> *Proposed move:* 15-minute starter — two anonymised extracts on the board, one +> narrating and one evaluating. Class identifies which does what, and names the +> three moves the evaluating one makes. Then all 19 rewrite their source paragraph. +> +> *Different need — 5 learners:* Evaluate reliability but only from provenance; +> push toward corroboration. Give them the "cross-reference" sentence stem. +> +> *Different need — 4 learners:* Secure on this criterion. Extension: which of your +> sources would you drop entirely, and what would you lose? +> +> *Individual comments needed:* 2 (see list) — both misread the question. + +This is the correct answer to "mark my class set", and you should propose it +proactively rather than generating thirty comments because thirty were requested. +Say why, then offer the individual comments as well if the teacher still wants them. diff --git a/skills/hinge-questions/SKILL.md b/skills/hinge-questions/SKILL.md new file mode 100644 index 0000000..5377c89 --- /dev/null +++ b/skills/hinge-questions/SKILL.md @@ -0,0 +1,114 @@ +--- +name: hinge-questions +description: Read this whenever a teacher needs to know, mid-lesson, whether the class has got it. Use it when asked for a quick check for understanding, a diagnostic or multiple-choice question, an exit-ticket question, or when someone needs to decide between moving on and reteaching. It produces a single item where every wrong answer names a specific misconception, plus the plan for what to do at each result. +--- + +# Hinge questions + +A hinge question sits where the lesson forks: the point at which what you do next +should depend on what the class currently understands. Highest-value item type in +Wiliam's toolkit, and the hardest to write. + +## What done looks like + +An item plus a response plan, in this shape: + +``` +HINGE QUESTION — +Place: +Format: Time: min + + +A. … B. … C. … D. … + +Key: + +Diagnosis: A → → + B → KEY + C → → + +Response plan: ≥80% key → move on + 50–80% → pair a right with a wrong, re-poll + <50% → reteach + C dominant → reteach first +``` + +The response plan is not an extra. Without it you have written a quiz question. + +## Hard constraints + +Both must hold or it isn't a hinge question: + +1. **Every learner answers in under two minutes** — ideally under one. +2. **The teacher reads the whole class in about 30 seconds**, standing at the + front, mid-lesson. This is what forces multiple choice with a simultaneous + all-class response. + +And three design constraints: + +3. **Every wrong answer is diagnostic.** If you can't write "a learner who picks C + believes ___", the option isn't working. +4. **The key can't be reached by wrong reasoning.** The most-broken rule. Defend + it by allowing more than one correct option without saying how many (kills + elimination), by making surface features point the wrong way, or by adding a + second tier: "which reason explains your answer?" +5. **Everyone commits before seeing anyone else.** Otherwise you're measuring + conformity. + +## The part that actually takes the work + +**Name the misconceptions before writing any options.** Then work the problem *as +a learner holding each one*, and the answer they'd get is your distractor. +Distractors invented for plausibility get chosen by nobody, and an item where +everyone picks the key tells the teacher nothing. + +Real misconceptions, best first: your own learners' previous wrong answers (pull +them if a Nurture server is connected); examiner reports and diagnostic question +banks; the misconception literature for the domain; systematic derivation from the +procedure. `references/misconception-sources.md` covers all four plus +domain-by-domain starting points. + +If you had to invent them, **say so** — they're untested. + +Then try to break your own item: reach the key by elimination, by surface +pattern-matching, by applying last lesson's rule without understanding. If any +works, redesign. + +## Judgement calls + +- **Is there actually a hinge here?** Ask the teacher: at what point would you do + two genuinely different things depending on the response? If there's no such + point, they need a retrieval quiz or an exit ticket, not this. Say so. +- **What would make the next 20 minutes wasted?** That misunderstanding is the + hinge — not the thing that's easiest to ask. +- **Offer 2–3 alternatives** at different lesson points when the topic has more + than one plausible hinge. + +## Reject the item if + +- [ ] It can't be answered in 2 minutes or read in 30 seconds. +- [ ] Any distractor has no named misconception behind it. +- [ ] The key is reachable by elimination, guessing, or surface matching. +- [ ] There's no response plan. +- [ ] It's a disguised recall question — wrong answers just mean "didn't remember". +- [ ] It has five or six options. Unreadable at a glance. +- [ ] Reading demand is doing work the concept should be doing. + +## Two failures that happen after the item is written + +- **Asking it and not acting on it.** The commonest failure in practice: the + teacher polls, sees half the class is wrong, and moves on because the plan says + so. Put the response plan in the output and say this out loud. +- **Recording it as a grade.** Hinge questions are diagnostic. The moment they + carry marks, learners optimise for looking right rather than revealing what they + think, and the instrument stops working. + +Treat the percentage bands as sensible defaults, not thresholds with evidential +standing. + +## References + +- `references/misconception-sources.md` — where real misconceptions come from, + how to derive them systematically, domain starting points. +- `references/worked-examples.md` — hinge questions across maths, science, + English, history, primary and vocational, each with its misconception map. diff --git a/skills/hinge-questions/references/misconception-sources.md b/skills/hinge-questions/references/misconception-sources.md new file mode 100644 index 0000000..6840454 --- /dev/null +++ b/skills/hinge-questions/references/misconception-sources.md @@ -0,0 +1,121 @@ +# Finding real misconceptions + +A distractor invented for plausibility is usually not chosen by anyone. A +distractor derived from a real misconception splits the class and tells the +teacher what to do. Getting real ones is most of the work. + +## Ranked sources + +**1. Your own learners' previous wrong answers.** Unbeatable. If the Nurture MCP +server is connected, `get_assignments` → `get_submissions` → +`get_submission_detail` gives you actual responses from actual learners on actual +tasks. Mine the wrong answers, cluster them, and each cluster becomes a distractor. + +**2. Examiner reports and mark scheme "common errors" sections.** Awarding bodies +publish exactly this, per question, per year, with frequencies. + +**3. Published diagnostic banks.** Diagnostic Questions collections (strong in +maths and science), national curriculum support materials, and the misconception +inventories that exist for specific domains — e.g. the Force Concept Inventory in +mechanics, and its equivalents in genetics, chemistry and astronomy. + +**4. The domain's misconception research literature.** Most curriculum areas have +a body of work on characteristic learner errors. Cite it if you use it. + +**5. Systematic derivation.** When you have none of the above, derive rather than +imagine — see below. Flag derived distractors as untested. + +## Systematic derivation + +Enumerate the steps of the target procedure. At each step, generate the +characteristic error. This produces distractors that at least correspond to +mechanically possible reasoning. + +**Procedural errors:** sign flip; inverse operation applied; steps in wrong order; +one step omitted; a step applied twice; off-by-one; correct method on the wrong +quantity; unit not converted; rounding at the wrong stage. + +**Over-generalisation:** a rule that holds in the cases seen so far, applied +outside its range. "Multiplication makes bigger." "Adding an -s makes a plural." +"The subject comes first in a sentence." "Metals are shiny, so this shiny thing +is a metal." This is the richest single source — most misconceptions are +successful generalisations from a restricted diet of examples. + +**Prototype effects:** the learner has one canonical example and rejects +non-prototypical members. A triangle must sit on a horizontal base. A bird must +fly. A "solvent" must be water. An "argument" must be an argument between people. + +**Surface-feature matching:** the learner selects the method by what the problem +*looks* like rather than its structure. Any question with two numbers and the word +"altogether" gets addition. + +**Confusing correlated concepts:** heat and temperature; weight and mass; area and +perimeter; speed and acceleration; cause and correlation; theme and plot; date and +period; author and narrator. + +**Linguistic interference:** everyday meaning bleeding into technical meaning. +"Theory" as guess; "significant" as important; "energy" as vigour; "salt" as table +salt; "acute" as severe; "novel" as any book. + +**Reversal errors:** getting the direction of a relationship backwards. The +classic students-and-professors problem; confusing the antecedent and consequent +of a conditional; reading a graph's axes the wrong way round; mixing up which +variable is manipulated. + +## Testing your distractors + +Before shipping the item, for each distractor answer: + +- **Who picks this?** Describe a specific learner state, not "someone confused". +- **What would you teach them?** If the answer is the same for two distractors, + merge them and find a different misconception for the freed slot. +- **Is it more attractive than the key to someone who doesn't know?** A distractor + nobody finds attractive is decoration. + +After use, ask the teacher for the actual distribution. An option chosen by nobody +should be replaced; that feedback loop is how an item bank gets good. + +## Domain starting points + +Non-exhaustive; use as prompts, not as a substitute for the sources above. + +**Number:** multiplication makes bigger / division makes smaller; longer decimal = +larger value (2.35 > 2.5); fractions compared by numerator or by denominator +alone; 0 is neither positive nor a number "you can have"; negative signs treated +as subtraction; percentages as absolute amounts; equals sign read as "here comes +the answer" rather than as a relation. + +**Algebra:** letters as labels rather than variables (4a = "4 apples"); the +conjoining error (3 + 2n = 5n); reversal in translating word to symbol; equations +"solved" by operating on one side only. + +**Geometry & measure:** area and perimeter conflated; angle as the length of the +drawn arms; shapes only recognised in canonical orientation; scale factor applied +linearly to area. + +**Mechanics:** motion requires a continuing force; heavier objects fall faster; a +force acts *in* a moving object; action–reaction pairs applied to the same body. + +**Chemistry:** mass "lost" in combustion; dissolving is melting; atoms of a solid +expand when heated; conservation of matter breaking down in gas-phase reactions. + +**Biology:** plants take food from the soil; evolution as individual adaptation +within a lifetime; dominant allele means most common; "survival of the fittest" +as strongest. + +**Reading & literature:** narrator equated with author; theme stated as topic; +inference treated as guessing; quotation used as evidence with no explanation of +effect. + +**Writing:** paragraphing by length rather than by idea; connectives added without +a logical relation; formality confused with complexity. + +**History:** presentism (judging past actors by present norms); source reliability +equated with the source being written at the time; cause conflated with the last +event before the outcome; "bias" used as a blanket disqualifier. + +**Geography:** weather and climate; erosion and weathering; development as wealth +alone; maps read without attention to projection or scale. + +**Computing:** assignment read as mathematical equality; a variable holding all its +past values; loops thought to execute in parallel; equality of objects vs identity. diff --git a/skills/hinge-questions/references/worked-examples.md b/skills/hinge-questions/references/worked-examples.md new file mode 100644 index 0000000..9281a51 --- /dev/null +++ b/skills/hinge-questions/references/worked-examples.md @@ -0,0 +1,159 @@ +# Worked hinge questions + +Each shows the misconception map and the response plan. Note that in several, +*more than one option is correct and the learners are not told how many* — this +is the main defence against elimination strategies. + +--- + +## Maths, ~age 11 — comparing decimals + +**Place:** After introducing decimal place value, before independent practice on +ordering. +**Format:** ABCD cards. **Time:** 1 min. + +> Which of these numbers is the largest? +> +> A. 0.62 B. 0.7 C. 0.532 D. 0.5301 + +**Key:** B + +| Choice | Misconception | Reteach | +|---|---|---| +| A | "Longer string of digits after the point = larger" applied at 2 dp | Place value with a number line, tenths first | +| C | Same rule, more strongly held (3 dp beats 2 dp) | As above; start with tenths only | +| D | Strongest form — 4 dp read as largest | As above, plus concrete equivalence 0.7 = 0.70 = 0.700 | + +**Response plan:** ≥80% B → move on. Mixed A/C/D → the misconception is uniform +and specific; reteach with 0.7 vs 0.70 vs 0.700 equivalence, then re-poll with a +parallel item. Even spread including B → guessing; return to place value. + +**Why it works:** every distractor is longer than the key, so the misconception is +maximally attractive and the item cannot be passed by "pick the longest". + +--- + +## Science, ~age 14 — forces and motion + +**Place:** After the demonstration, before applying Newton's first law to problems. +**Format:** Mini whiteboards, letters only. **Time:** 90 s. + +> A puck slides across frictionless ice at a constant speed. Which forces act on it? +> +> A. A forward force keeping it moving +> B. Gravity downward +> C. The normal force from the ice upward +> D. A forward force larger than friction + +**Key:** B and C — and the class is *not told* how many are correct. + +| Choice | Misconception | Reteach | +|---|---|---| +| A | Impetus theory — motion requires a continuing force. The single most persistent misconception in mechanics. | Contrast constant velocity with acceleration explicitly; predict-then-observe | +| D | Same, with a net-force overlay | As A, then free-body diagrams | +| Missing B or C | Vertical forces ignored when attention is on horizontal motion | Free-body diagram routine on every problem | + +**Response plan:** Anyone selecting A or D holds impetus theory and will misread +every subsequent problem — this is a stop-and-reteach signal even at 30% of the +class, because it is known to be resistant. Pair discussion rarely shifts it; +prefer a predict–observe–explain demonstration. + +--- + +## English, ~age 13 — evidence in analytical writing + +**Place:** Before drafting an analytical paragraph. +**Format:** ABCD cards. **Time:** 2 min. + +> Which of these sentences uses evidence to support an interpretation? +> +> A. The writer uses the word "crept", which is a verb. +> B. The writer describes how the fog "crept" through the streets. +> C. The word "crept" makes the fog seem like a living thing that is hiding +> something. +> D. The fog is described as creeping, which creates a mysterious atmosphere. + +**Key:** C and D (again, count not disclosed). + +| Choice | Misconception | Reteach | +|---|---|---| +| A | Technique-spotting: naming the word class counts as analysis | Model the difference between *identifying* and *explaining effect* | +| B | Narration: retelling the text counts as evidence | Two-column exemplar comparison — "what it says" vs "what it does" | + +**Response plan:** A dominant → whole-class reteach with the "so what?" chain. +B dominant → the class can find evidence but not use it; this is a different and +easier fix — give a sentence stem. Mixture → pair one C/D learner with one A/B +learner and re-poll. + +--- + +## Primary, ~age 7 — the equals sign + +**Place:** Before number-sentence work. **Format:** Thumbs / cards. **Time:** 1 min. + +> What number goes in the box? `8 + 4 = ☐ + 5` + +Offered as A. 12 B. 7 C. 17 D. 9 + +**Key:** B + +| Choice | Misconception | Reteach | +|---|---|---| +| A | Equals read operationally: "= means write the answer" — the single biggest barrier to later algebra | Balance scales; "is the same as"; true/false number sentences | +| C | Adds everything visible | As A, plus attention to the structure of the sentence | +| D | Adjusts by the difference in the wrong direction | Nearly there — check with concrete materials | + +**Response plan:** A dominant is extremely common and worth a full reteach even in +a class that is otherwise fluent — it predicts later algebra failure. + +--- + +## History, ~age 15 — source evaluation + +**Place:** Before an independent source-evaluation task. **Format:** Poll. **Time:** 2 min. + +> A private letter written by a soldier during the battle, and a history book +> written 60 years later. Which is more useful for finding out what the battle was +> like for soldiers? +> +> A. The letter, because it was written at the time +> B. The letter, because it comes from someone who experienced it +> C. The book, because the author could see the whole picture +> D. It depends on what specifically you want to find out + +**Key:** B and D + +| Choice | Misconception | Reteach | +|---|---|---| +| A | Contemporaneity treated as automatically conferring usefulness | Counter-example: a contemporary propaganda poster | +| C | Hindsight treated as automatically superior | Counter-example: the book's account of individual experience | +| Rejecting D | Usefulness treated as a property of the source rather than of the question | Explicit teaching: usefulness is always *for a purpose* | + +**Response plan:** If most pick A, the class has learned a rule ("primary = +better") rather than a method. Reteach with paired counter-examples before any +independent work; source questions will otherwise all be answered by the rule. + +--- + +## Vocational — electrical safety + +**Place:** Before supervised practical. **Format:** Cards. **Time:** 1 min. + +> Before working on a circuit you have isolated, what must you do? +> +> A. Check the circuit is dead with an approved voltage indicator +> B. Check the voltage indicator on a known live source before and after testing +> C. Lock off the isolator and keep the key +> D. Put a warning notice on the isolator + +**Key:** A, B, C and D — all of them. + +The design point: the misconception is that safe isolation is *one* action rather +than a complete procedure. Any learner selecting fewer than four is not yet safe +to work unsupervised, which makes this a genuine hinge — the response plan is not +"reteach", it is "these learners do not proceed to the practical". + +**Response plan:** 100% four-of-four → proceed. Anything less → whole-class +re-run of the procedure, then re-poll; no practical work until clear. Note that +this is one of the rare cases where the hinge question gates an activity rather +than choosing between teaching moves. diff --git a/skills/peer-and-self-assessment/SKILL.md b/skills/peer-and-self-assessment/SKILL.md new file mode 100644 index 0000000..ffe0a44 --- /dev/null +++ b/skills/peer-and-self-assessment/SKILL.md @@ -0,0 +1,104 @@ +--- +name: peer-and-self-assessment +description: Read this when learners are going to assess work — their own or each other's — because unstructured peer feedback reliably degrades into praise and personal comment. Use it when asked about peer review, peer marking, self-assessment, reflection prompts, metacognition or student ownership, and when a teacher's feedback workload is the real constraint. It produces a scripted protocol with explicit criteria, a constrained response format, and allocated time to act on it. +--- + +# Peer and self-assessment + +Wiliam's fourth and fifth key strategies. Together they answer the arithmetic +problem at the centre of formative assessment: one teacher cannot give thirty +learners timely feedback, but thirty learners can. + +The deeper reason isn't workload. **Assessing someone else's work against criteria +is one of the most reliable ways to internalise the criteria** — the assessor +usually learns more than the assessed. So optimise the protocol for what it +teaches the *giver*. + +## What done looks like + +A script a teacher could run tomorrow: the exact steps, time per stage, the +sentence stems if the class needs them, what the receiver does with the result and +when, and how the teacher samples the quality of the feedback being given. + +## Conditions — all four, or it fails + +1. **Explicit criteria the learners already understand**, via exemplar work, not + handout. If `success-criteria` hasn't been done, peer assessment will not work. + Do that first. +2. **A protocol that constrains the response** — how many points, of what type, in + what form. "Give each other feedback" produces nothing. +3. **Directed at the work, never the person.** A class norm, stated and enforced. + The first unchallenged personal comment kills the protocol. +4. **The giver can't solve it for the receiver.** Peer feedback that supplies the + answer is copying with extra steps. + +Plus a sequencing rule that decides whether any of this works: **start on +anonymous work, not on each other's.** Two or three rounds on anonymised exemplars +first is the difference between a working protocol and "this is really good, maybe +check your spelling". Most classrooms attempt open peer review on day one, which +is why most peer assessment disappoints. `references/protocols.md` has the +six-step sequence for teaching it over a half-term. + +## Self-assessment — the harder one + +Unreliable exactly where it matters most: learners who lack the concept of the +goal can't judge their distance from it, and confidence correlates worst with +competence at the bottom. So **self-assessment needs exemplars or a prior piece, +not just criteria.** Anchored to nothing, it measures mood. + +What works: self-check against a checklist before submitting (shifts quality +control to the only point where the learner can still act cheaply); plus/minus/ +equals against your own previous piece (comparison to self, not peers, which keeps +it task-involving); the three-sentence gap statement; predict-the-mark, where the +*size and direction* of the error is itself the diagnosis. + +What doesn't: "how do you think you did?", end-of-lesson smiley faces, and +self-assigned grades that feed reporting — that last one creates an incentive to +misreport and destroys the diagnostic value. + +## Reflection prompts + +Generic prompts get generic answers. Make them specific and answerable: + +| Weak | Better | +|---|---| +| "What did you learn today?" | "What can you do now that you couldn't at the start?" | +| "How do you feel about this topic?" | "Which part would you struggle to explain to someone a year below?" | +| "What went well?" | "Which decision here are you least sure was right?" | +| "What will you do better next time?" | "Name the one step you skipped, and what you'll do instead." | + +If reflections already exist in a connected Nurture server, read what learners +actually wrote before designing new prompts — the failure is usually the prompt, +not the learners. + +## Guardrails + +- **Never let peer assessment contribute to a grade.** It changes the incentive + from helping to negotiating and puts learners in a position they shouldn't be in + with each other. +- **Pairings are not neutral.** Don't pair a learner with whoever is most likely to + embarrass them, and don't always pair strongest with weakest — that arrangement + teaches one of them their job is to be helped. +- **Protect learners who are behind.** Rank-ordering, displayed peer ratings and + "who got the most stars" convert a formative tool into a status contest. +- **Sensitive content.** Learners write about their own lives. Any protocol that + circulates personal writing needs a no-questions-asked opt-out. +- **Watch for the confident-and-wrong assessor.** A fluent but mistaken peer does + real damage. This is why the teacher samples the feedback given. + +## Judgement calls + +- **Is the prerequisite in place?** Criteria understood via exemplars. If not, + that's the work — route to `success-criteria` and say why. +- **Which protocol?** Match to the situation, not to fashion. + `references/protocols.md` has a selection table and full scripts — two stars and + a wish, pre-flight checklist, focused review, error-spotting, observation + schedule, C3B4ME, help board. +- **Does the receiver have time to act?** Peer feedback with no response time has + the same problem as teacher feedback with no response time. + +## References + +- `references/protocols.md` — full scripts with timings and sentence stems, the + self-assessment routines, and the half-term sequence for teaching a class to + give feedback. diff --git a/skills/peer-and-self-assessment/references/protocols.md b/skills/peer-and-self-assessment/references/protocols.md new file mode 100644 index 0000000..2876183 --- /dev/null +++ b/skills/peer-and-self-assessment/references/protocols.md @@ -0,0 +1,191 @@ +# Peer and self-assessment protocols — full scripts + +Each protocol gives the steps, the timing, and the exact wording where wording +matters. Adapt the sentence stems to the age group; keep the structure. + +--- + +## Two stars and a wish (10 min) + +The starting protocol. Run it on **anonymised exemplars** for the first two or +three outings before letting learners use it on each other's work. + +1. *(1 min)* Display the criteria. Read them aloud. +2. *(4 min)* Each learner writes, about the piece in front of them: + - **★** One thing that meets criterion ___, and where. + - **★** One more, against a different criterion. + - **✎** One wish: *"It would be even better if you…"* — must name a criterion + and a specific place in the work. +3. *(3 min)* Return and read. +4. *(2 min)* Receiver writes one sentence: what they will change, and where. + +**Rules to state out loud:** every star names a criterion and a location; the +wish is an instruction, not an opinion; no comment about the person. + +**Failure mode:** stars become "I like it" and the wish becomes "check your +spelling". Fix by requiring the criterion number in every line. + +--- + +## Pre-flight checklist (5 min, before submission) + +Borrowed from aviation, and the highest-value-per-minute protocol here. + +1. The teacher publishes a short checklist (4–6 binary items) with the task. +2. Before submitting, the learner self-checks. +3. A peer then checks the **same** list and signs it. +4. Work is only accepted with a signed checklist. + +**Why it works:** it moves quality control to before submission, and the signature +creates real accountability for the checker. Teachers report a sharp drop in +mechanical errors, which frees the actual feedback to be about the substance. + +**Rules:** the checker signs; if an item is later found unmet, both learners fix +it. The checklist must be binary — anything requiring judgement belongs elsewhere. + +--- + +## Focused peer review (20 min, drafting) + +1. *(2 min)* Teacher names **one** criterion — the same one for everyone. +2. *(5 min)* Learners read their partner's draft **twice**: once for sense, once + for the criterion. +3. *(5 min)* Reviewer writes, using stems: + - *"The strongest example of ___ is in paragraph ___ because ___"* + - *"The place where ___ is weakest is ___"* + - *"One question I had as a reader was ___"* +4. *(3 min)* Author reads and asks the reviewer **one** clarifying question aloud. +5. *(5 min)* Author revises that one criterion only. + +**Rules:** one criterion only — the discipline is the point; the reviewer's +question in step 3 must be a genuine reader's question, not a disguised +instruction. + +--- + +## Error-spotting in a flawed exemplar (15 min) + +Best where the class shares a misconception, and no learner has to own it. + +1. Teacher provides one piece containing **three** deliberate errors of a named + type (state the number, not the locations). +2. *(5 min)* Pairs find all three. +3. *(5 min)* For each: what's wrong, why someone might think it, and the fix. +4. *(5 min)* Class shares; teacher names the misconception explicitly. +5. Learners then check their own work for the same error. + +**Why this before peer review of real work:** it teaches the diagnostic move with +no social stakes. Step 5 is the transfer and is usually skipped — don't skip it. + +--- + +## Observation schedule (practical / performance) + +For practical, performance, oral or physical work, where a peer *judging* is +neither reliable nor comfortable. + +The peer **records**, rather than evaluates: + +``` +Observing: __________ Criterion: __________ + +Tally each time you see: + Checks the measurement before recording ▯▯▯▯▯ + States the reason for a choice out loud ▯▯▯▯▯ + Returns equipment to a safe position ▯▯▯▯▯ + +One thing you saw that isn't on this list: ______________ +``` + +The performer reads their own tallies and draws their own conclusion. The observer +learns the criteria by watching for them. Nobody has to deliver a verdict about a +peer, which removes the main source of friction. + +--- + +## C3B4ME — "see three before me" (ongoing routine) + +Before asking the teacher, consult three sources: your notes, the working wall or +model on display, and a peer. State which three, and display them. + +**Exemption, always stated:** anyone stuck at the very first step goes straight to +the teacher. Otherwise the routine punishes exactly the learners who need most +help, which inverts its purpose. + +--- + +## Homework help board (ongoing) + +A physical or digital board where learners post problems and other learners +answer. The teacher reads it before the lesson — it is a free map of the sticking +points across the class. + +**Rules:** post the specific difficulty, not "I don't get question 4"; answers +explain the route, never just give the answer; the teacher only intervenes on +something wrong or unanswered after 24 hours. + +--- + +## Self-assessment: the gap statement (5 min) + +Three sentences, against one named criterion. Sadler's three conditions, +compressed: + +``` +Where I am: +Where I need to be: +My next step: +``` + +**Rule:** the next step must be an action, not an aspiration. "Be more detailed" +is rejected; "add a second piece of evidence to paragraph 3" is accepted. + +--- + +## Self-assessment: plus / minus / equals (10 min) + +Learners compare this piece against their own previous piece on the same +criterion: + +- **+** what is better, and the specific evidence +- **−** what is worse +- **=** what has not changed + +Then: *"which of the − or = items will I target next, and how?"* + +**Why it matters:** comparison to self, not to peers. This is the version of +self-assessment that works for learners well behind the class — the reference +point is reachable. + +--- + +## Self-assessment: predict the mark (5 min) + +Before results are returned, learners predict their own mark against the criteria. +The teacher records the gap between predicted and actual. + +| Pattern | Meaning | Response | +|---|---|---| +| Predicts high, scores low | Doesn't yet hold the standard — the dangerous case | Exemplar work; this learner cannot self-correct yet | +| Predicts low, scores high | Standard held; confidence isn't | Evidence-based reassurance, specific and repeated | +| Accurate | Self-assessment is functioning | Push toward setting their own next targets | + +Recording this over a term shows whether self-assessment is developing — which is +the actual target of strategy 5, and one of the few things worth tracking +longitudinally. + +--- + +## Teaching the class to give feedback + +Peer feedback quality is a taught skill. A rough sequence over a half-term: + +1. Rank two anonymous pieces and justify. *(judgement)* +2. Match teacher comments to anonymous pieces. *(criteria in action)* +3. Two stars and a wish on an anonymous piece; compare with the teacher's. *(calibration)* +4. Pre-flight checklist on a partner's work. *(binary, low stakes)* +5. Focused review, one criterion, on a partner's draft. *(judgement, real stakes)* +6. Open peer review against full criteria. *(only after the above)* + +Most classrooms attempt step 6 on day one, which is why most peer assessment +disappoints. diff --git a/skills/success-criteria/SKILL.md b/skills/success-criteria/SKILL.md new file mode 100644 index 0000000..a794196 --- /dev/null +++ b/skills/success-criteria/SKILL.md @@ -0,0 +1,104 @@ +--- +name: success-criteria +description: Read this when a task needs a stated standard — feedback, peer assessment and self-assessment all fail silently without one. Use it when asked for learning objectives, lesson aims, WALT/WILF, success criteria, a rubric or mark scheme for learners, or a checklist for a task, and when feedback has nowhere to attach because no standard was ever shared. It produces a decontextualised learning intention, 3–5 learner-checkable criteria, and the exemplar activity that makes them mean anything. +--- + +# Learning intentions and success criteria + +Wiliam's first key strategy, and the prerequisite for the other four. Sadler +(1989): a learner cannot close a gap they cannot perceive. + +## What done looks like + +``` +LEARNING INTENTION + We are learning to . + +CONTEXT (today's vehicle, not the learning) + + +SUCCESS CRITERIA (3–5, learner-checkable, learner voice) + □ … □ … □ … + +HOW THESE GET SHARED (the activity, not the handout) + + +EXEMPLAR + +``` + +## The two distinctions to get right + +**Intention vs context.** The intention is the transferable learning; the context +is the vehicle. "To write a letter to the council about the car park" conflates +them; "to write persuasively for a specific audience" *in the context of* that +letter separates them. Learners who only meet the contextualised version encode +the context — they learn car-park letters, not persuasion, and don't transfer. + +**Product vs process criteria.** Product: "your conclusion states which hypothesis +the data support." Process: "you checked each result against your prediction +before writing the conclusion." Novices usually need process criteria — product +criteria describe the destination without the route. Say which you wrote and why. + +## Why exemplars matter more than wording + +For anything with real quality in it, criteria are irreducibly vague. "Well +structured", "sophisticated analysis" mean nothing to a learner who can't yet do +it — they only become meaningful once you can already recognise the thing. This +is not fixable by better wording; more words make rubrics longer, not clearer. + +So for extended writing, practical work, design, or oral work: deliver the rubric +asked for **and** the exemplar activity, and say the second matters more. +Comparison beats description — "which of these two is better and why?" is a far +easier judgement than rating one against a scale, and the criteria learners +generate that way are ones they can already apply. + +## Constraints + +- **3–5 criteria.** Twelve is a mark scheme, not success criteria. +- **Every criterion is one a learner could honestly self-check.** "Uses + sophisticated vocabulary" fails. "Every claim I make is followed by evidence + from the text" passes. +- **Observable verbs only.** Not "understand", "know about", "be aware of", + "appreciate", "explore" — those name no observable. +- **No criterion restates the task** ("completes all questions", "writes two + pages") or measures compliance, effort or presentation unless that is genuinely + the learning. +- **No comparative language** in learner-facing criteria. "More sophisticated + than" — than what? They can't check against a comparison they can't see. +- **Sharing means the learners have used them**, not received them. Specify the + activity. + +## Judgement calls + +- **Should the criteria be shared at all?** For genuine problem-solving, + investigation or design, telling learners the product criteria in advance tells + them the answer. Share the *process* criteria up front and build the product + criteria with the class afterwards from what they produced. Flag this rather + than mechanically producing an up-front list. +- **Rubric or checklist?** Most teachers reach for an analytic grid where a + checklist or a single-point rubric would serve better and cost less. + `references/rubric-patterns.md` picks by situation and is honest about what each + format costs. +- **Is the language actually the learners'?** Read it aloud as a learner of that + age. Examiner language cannot be self-checked; keep the mark scheme separate. + +## Failure modes + +- **Objectives that are tasks.** "To complete the worksheet on ratios." The verb + describes an activity, not a capability. +- **Criteria that are quantity.** "Write three paragraphs" measures nothing. +- **Rubric bloat.** Four bands × six strands = 24 cells nobody reads. Bands are for + reporting; criteria are for working. +- **WALT/WILF theatre.** Writing the objective on the board, having learners copy + it, never referring to it again — five minutes a lesson for nothing. If that's + the practice, say so and propose the mid-lesson self-check instead. +- **Flattening creative work.** Where divergence is the point, constrain process + and quality, not the product's shape, or you get thirty identical pieces. + +## References + +- `references/rubric-patterns.md` — checklist, single-point, analytic, comparative + judgement, progression: which to use and what each costs. +- `references/exemplar-protocols.md` — the activities that make criteria land, + with timings, plus how to source and anonymise exemplars. diff --git a/skills/success-criteria/references/exemplar-protocols.md b/skills/success-criteria/references/exemplar-protocols.md new file mode 100644 index 0000000..4b1522c --- /dev/null +++ b/skills/success-criteria/references/exemplar-protocols.md @@ -0,0 +1,118 @@ +# Exemplar protocols + +Sharing criteria means learners have *used* them. These are the activities that do +that. All are short — 10–20 minutes — and all belong *before* the learners attempt +the task, not after. + +--- + +## Best of two (10 minutes) — the default + +Two anonymised responses to the same task, one clearly stronger. + +1. Learners read both, decide which is better, and write **one sentence** saying why. +2. Pairs compare reasons. +3. The class's reasons are collected on the board — these become the success + criteria, in the learners' own words. +4. The teacher adds at most one criterion the class missed. + +Why it works: choosing between two is far easier than judging one against an +abstract scale, and the criteria that emerge are ones learners can already apply. + +**Set-up rule:** both responses must answer the *same* task, and differ in quality +rather than topic or approach — otherwise learners debate preference, not quality. + +--- + +## Rank the four (20 minutes) + +Four responses spanning the range. Groups rank them and must justify each +adjacent pair. + +Adjacent-pair justification is the point: it forces learners to articulate the +specific difference at each step, which is where the criteria live. Ranking alone +produces "this one's best" with no reasoning. + +Follow with: "which pair was hardest to separate, and what finally decided it?" + +--- + +## Find the criterion (15 minutes) + +Give the criteria list and one strong response. Learners annotate the response, +marking where each criterion is met. + +Use when criteria already exist and must be used (an exam board's, a +department's). The annotation converts abstract wording into recognisable moves. +Any criterion learners cannot locate in the exemplar is one you need to reword or +teach explicitly. + +--- + +## Improve the weak one (20 minutes) + +One deliberately mid-quality response. In pairs, learners make **three** specific +improvements — not a rewrite. + +The constraint to three forces prioritisation, which is the same judgement they +will need on their own draft. Collect the improvements; the most-chosen ones are +the criteria that matter most to this class right now. + +--- + +## What did the marker see? (15 minutes) + +Give responses and comments separately; learners match them. (The `giving-feedback` +skill lists this as a feedback move — it works equally well as goal-setting before +the task.) + +Strong for exam classes: it makes the marker's reasoning visible rather than +leaving it as an opaque verdict. + +--- + +## Co-construction (25 minutes, once per unit) + +1. Learners attempt a short version of the task cold, with no criteria. +2. Anonymised samples of their own work go on the board. +3. Class discusses what distinguishes the stronger ones. +4. The class writes the criteria. +5. Those criteria govern the real task. + +Highest-investment, highest-return version. Using the class's *own* work rather +than exemplars from elsewhere makes the standard feel reachable — the strongest +piece is by someone in the room. + +**Handle with care:** anonymise properly. Learners recognise handwriting and +phrasing; retype the samples. Never use a weak sample identifiable to its author. + +--- + +## Sourcing exemplars + +In order of preference: + +1. **Previous cohorts' work**, retyped and anonymised, with permission per your + school's policy. Most credible, most calibrated to what your learners actually + do. If the Nurture MCP is connected, `get_submissions` on a previous + assignment gives you a real range to draw from. +2. **Awarding body exemplar material** — annotated and moderated, though often + only at band boundaries and often unrepresentative of ordinary work. +3. **Teacher-written** at deliberately different quality levels. Fast, but + teacher-written "weak" work is characteristically wrong — it contains the + errors teachers *expect*, cleanly, and rarely the messy, partly-right work + real learners produce. Use as a stopgap. +4. **AI-generated.** Use only for the strong exemplar, and even then check it + against the actual specification. AI-written "weak" exemplars are the least + realistic of all: they are fluent-but-shallow, whereas real weak work is + usually disfluent-but-thinking. If you generate one, mark it as generated and + tell the teacher to swap it for real work as soon as they have some. + +## Anonymisation rules + +- Retype; do not photograph handwriting. +- Strip names, dates, and any personal detail in the content itself — learners + write about their own lives more than authors of exemplars expect. +- Never display work identifiable to a learner in the room as the weak example. +- Get consent per your setting's policy before using a learner's work with + another class or in staff training. diff --git a/skills/success-criteria/references/rubric-patterns.md b/skills/success-criteria/references/rubric-patterns.md new file mode 100644 index 0000000..6c985c3 --- /dev/null +++ b/skills/success-criteria/references/rubric-patterns.md @@ -0,0 +1,129 @@ +# Rubric patterns + +Pick the lightest structure that supports the decision. Rubric complexity buys +apparent precision and costs usability, and the precision is usually illusory. + +--- + +## 1. Checklist — binary, learner-facing + +``` +□ Every claim is followed by evidence from the text +□ Each quotation is embedded inside my own sentence +□ I have explained the effect on the reader, not just named the technique +□ My final paragraph answers the question directly +``` + +**Use for:** self-check before submission, peer pre-flight checklists, process +criteria, anything where the learner needs to act during the work. +**Strength:** learners can genuinely apply it. **Limit:** no gradation — cannot +distinguish adequate from excellent. + +This is the default. Most teachers reach for an analytic grid when a checklist +would do the job better. + +--- + +## 2. Single-point rubric + +One column describing proficiency; blank columns either side for the marker or +peer to note where the work falls short or exceeds it. + +| Not yet | **The standard** | Beyond | +|---|---|---| +| | Argument is sustained across the whole response, with each paragraph advancing it | | +| | Evidence is selected because it is the strongest available, not the first found | | + +**Use for:** extended writing, project work, anything where the interesting +information is *how* a piece departs from the standard. +**Strength:** avoids the fiction that all shortfalls are the same shortfall; +generates specific comments naturally; far quicker to write than a full grid. +**Limit:** doesn't produce a score. + +Strong default when a checklist is too coarse and a full grid is overkill. + +--- + +## 3. Analytic rubric (the grid) + +Strands × levels, descriptors in every cell. + +**Use for:** summative moderation, multiple markers, when a defensible score is +required. +**Costs, which should be stated when you deliver one:** +- Descriptor language is comparative and therefore vague ("some", "detailed", + "sophisticated") — the words carry meaning only for people who already share + the standard, which is markers, not learners. +- Halo effects: markers who place a piece in a band on one strand place it there + on the others. +- It fragments quality into strands that don't independently vary in real work. +- Learners optimise for the cells, which flattens the work. + +If you must produce one: keep to 3–4 strands and 3–4 levels, make the descriptor +differences *substantive* (what changes) rather than *quantitative* ("more +detail"), and pair every band with an annotated exemplar. + +--- + +## 4. Comparative judgement + +No rubric. Assessors are shown pairs of responses and choose the better one; +many judgements are aggregated into a scale. + +**Use for:** hard-to-rubric constructs — writing quality, mathematical +problem-solving, design. Reliability is typically high because "which is better" +is a much easier judgement than "which band is this". +**Cost:** produces a rank, not a diagnosis. It tells you the order, not the reason, +so it needs pairing with something diagnostic for formative use. Needs enough +judges or enough judgements. + +Mention this when a teacher is struggling to get consistent marking on extended +writing across a department — it usually solves that specific problem. + +--- + +## 5. Progression / learning ladder + +An ordered sequence describing what typically comes next, rather than levels of +quality on one task. + +``` +… can identify the variables in a described experiment +… can identify which variable must be controlled and say why +… can design a fair test for a given question +… can identify the limitation of their own design +``` + +**Use for:** long-run skill development across a year, self-assessment against +"what's my next step", and for making the *next* target obvious rather than the +current *level*. +**Strength:** points forward, which is what learners need. +**Limit:** progressions are approximations — real learners skip and regress. Say +so; a ladder presented as a fixed sequence becomes a ceiling. + +--- + +## Choosing + +| Situation | Pattern | +|---|---| +| Learner needs to check work in progress | Checklist | +| Feedback needs to be specific and individual | Single-point | +| Score needed, multiple markers, moderation | Analytic grid + exemplars | +| Extended writing, consistency across a department is the problem | Comparative judgement | +| "What's my next step?" over a term or year | Progression | +| Creative or divergent work | Process checklist only; construct product criteria after | + +## Descriptor writing rules + +If you are writing band descriptors: + +- Name **what changes**, not how much: "selects evidence to counter an + anticipated objection" beats "uses more sophisticated evidence". +- Avoid quantity words as the sole discriminator — "some", "several", "a range + of" produce marker disagreement and teach learners to pad. +- One idea per descriptor. Compound descriptors ("clear and detailed and + accurate") make work that is clear but not detailed unplaceable. +- Keep parallel structure across bands so the difference is visible. +- Write the top band first, then the bottom, then interpolate — writing upward + produces bands that are all "the same but more".