Assessing Learning in the Age of AI
Assessment should produce trustworthy evidence that students achieved the learning outcomes. Generative AI changes which tasks can provide that evidence, but it does not imply that every assessment must prohibit AI or incorporate it.
Start with the capability students must demonstrate, then choose the role of AI. Do not begin with a preferred tool or an attempt to make a task “AI-proof”.
Choose the role of AI deliberately
| Level | Use when | Possible assessment evidence |
|---|---|---|
| AI-free | Independent performance is the outcome | supervised performance, oral explanation, practical demonstration, or time-bounded task |
| AI-limited | Some support is compatible with the outcome | final product plus specified planning, editing, or accessibility support |
| AI-permitted | Tool choice is not central to the outcome | final product, source verification, and proportionate disclosure |
| AI-integrated | Critical AI use is part of the outcome | interaction decisions, evaluation, revision, disciplinary reasoning, and reflection |
Use Setting course and assignment AI rules to turn the selected level into concrete instructions. A course can—and usually should—use different levels for different assessments.
Redesign from the outcome
Define the claim the assessment should support
Write what a successful student must know or be able to do. Separate outcomes that require independent fluency from those that involve selecting, using, or evaluating tools.
Decide what authentic resources are allowed
Real professional work may include references, calculators, software, colleagues, or AI. Match the assessment conditions to the intended capability without allowing a resource to perform the capability being assessed.
Collect more than generic product quality
Ask for evidence that reveals disciplinary decisions: application to specific course material, justification of choices, source use, revision, oral defence, performance, or transfer to a new case. Use multiple forms of evidence for high-consequence decisions.
Check burden, access, and fairness
Oral follow-ups, portfolios, and staged submissions consume staff and student time. Required AI may introduce cost, account, language, privacy, or accessibility barriers. Choose a design the institution can support consistently and provide an equivalent alternative where needed.
Pilot and calibrate
Try the task with and without allowed AI support. Check whether it elicits the target thinking, whether the instructions are interpretable, and whether the rubric rewards student learning rather than polish. Review results across different student groups and revise before high-stakes use.
Useful assessment patterns
Staged work with purposeful reflection
Students can submit a proposal, evidence plan, draft, feedback response, and final product. Ask them to annotate a small number of consequential changes and explain how evidence or feedback affected their decisions. Iteration can support self-regulated learning, but the resulting record is evidence of process—not proof of authorship.1
Oral or performance follow-up
A short, structured conversation can test whether students can explain, transfer, or defend important choices.2 Publish the questions or question types, use a rubric, account for disability and language needs, and avoid turning an informal conversation into an inconsistent penalty.
Portfolio
A curated portfolio can show development across tasks when it includes selection rationale, revision, feedback use, and reflection. Portfolios require clear purpose and support; collecting more artefacts does not automatically produce better evidence.3
AI critique or comparison
Provide an AI-generated explanation, solution, or artefact and ask students to identify claims, verify evidence, find omissions, and produce a justified revision. Prefer an educator-vetted artefact so the activity does not depend on a live model producing a useful mistake.
AI-supported production
Where AI is allowed, assess the student’s framing, source selection, verification, disciplinary decisions, and final accountability. A reflective commentary should address specific decisions and failures, not merely list prompts.
Multimodal or scenario-based performance
Students can interpret data, produce a diagram, respond to an unfolding case, or explain the same idea in several modes. Multiple modes can broaden the evidence available, but they must remain accessible and aligned with the construct being assessed.4
Process evidence without surveillance
Drafts, notes, use records, and version history can support learning when they are proportionate and announced in advance. Do not require full AI transcripts, screenshots, keystroke histories, or continuous monitoring by default. These may collect personal or irrelevant data, disadvantage some workflows, and still do not prove authorship.
If work raises an integrity concern, apply the published rule and the normal institutional process. See Why AI detection does not prove authorship.
Checklist before release
- Every task maps to an approved learning outcome.
- The AI level follows from that outcome.
- Allowed tools and actions are concrete.
- The evidence can be assessed reliably within available time.
- Students receive practice in any unfamiliar format.
- Access, accessibility, privacy, and non-AI alternatives are addressed.
- The rubric rewards knowledge, reasoning, and decisions—not model fluency.
- The brief explains disclosure, questions, and integrity procedures.
- High-stakes decisions use more than one appropriate source of evidence.
References & Footnotes
Footnotes
-
Nicol, D. J., & Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: a model and seven principles of good feedback practice. Studies in Higher Education, 31(2), 199–218. https://doi.org/10.1080/03075070600572090 ↩
-
Joughin, G. (1998). Dimensions of oral assessment. Assessment & Evaluation in Higher Education, 23(4), 367–378. https://doi.org/10.1080/0260293980230404 ↩
-
Driessen, E. W., van Tartwijk, J., van der Vleuten, C. P. M., & Wass, V. (2007). Portfolios in medical education: why do they meet with mixed success? A systematic review. Medical Education, 41(12), 1224–1233. https://doi.org/10.1111/j.1365-2923.2007.02944.x ↩
-
Ross, J., Curwood, J. S., & Bell, A. (2020). A multimodal assessment framework for higher education. E-Learning and Digital Media, 17(4), 290–306. https://doi.org/10.1177/2042753020927201 ↩