Skip to Content
đźš§ This site is under development; its content is not final and may change at any time. đźš§

Assessing Learning in the Age of AI

Assessment should produce trustworthy evidence that students achieved the learning outcomes. Generative AI changes which tasks can provide that evidence, but it does not imply that every assessment must prohibit AI or incorporate it.

Start with the capability students must demonstrate, then choose the role of AI. Do not begin with a preferred tool or an attempt to make a task “AI-proof”.

Choose the role of AI deliberately

LevelUse whenPossible assessment evidence
AI-freeIndependent performance is the outcomesupervised performance, oral explanation, practical demonstration, or time-bounded task
AI-limitedSome support is compatible with the outcomefinal product plus specified planning, editing, or accessibility support
AI-permittedTool choice is not central to the outcomefinal product, source verification, and proportionate disclosure
AI-integratedCritical AI use is part of the outcomeinteraction decisions, evaluation, revision, disciplinary reasoning, and reflection

Use Setting course and assignment AI rules to turn the selected level into concrete instructions. A course can—and usually should—use different levels for different assessments.

Redesign from the outcome

Define the claim the assessment should support

Write what a successful student must know or be able to do. Separate outcomes that require independent fluency from those that involve selecting, using, or evaluating tools.

Decide what authentic resources are allowed

Real professional work may include references, calculators, software, colleagues, or AI. Match the assessment conditions to the intended capability without allowing a resource to perform the capability being assessed.

Collect more than generic product quality

Ask for evidence that reveals disciplinary decisions: application to specific course material, justification of choices, source use, revision, oral defence, performance, or transfer to a new case. Use multiple forms of evidence for high-consequence decisions.

Check burden, access, and fairness

Oral follow-ups, portfolios, and staged submissions consume staff and student time. Required AI may introduce cost, account, language, privacy, or accessibility barriers. Choose a design the institution can support consistently and provide an equivalent alternative where needed.

Pilot and calibrate

Try the task with and without allowed AI support. Check whether it elicits the target thinking, whether the instructions are interpretable, and whether the rubric rewards student learning rather than polish. Review results across different student groups and revise before high-stakes use.

Useful assessment patterns

Staged work with purposeful reflection

Students can submit a proposal, evidence plan, draft, feedback response, and final product. Ask them to annotate a small number of consequential changes and explain how evidence or feedback affected their decisions. Iteration can support self-regulated learning, but the resulting record is evidence of process—not proof of authorship.1

Oral or performance follow-up

A short, structured conversation can test whether students can explain, transfer, or defend important choices.2 Publish the questions or question types, use a rubric, account for disability and language needs, and avoid turning an informal conversation into an inconsistent penalty.

Portfolio

A curated portfolio can show development across tasks when it includes selection rationale, revision, feedback use, and reflection. Portfolios require clear purpose and support; collecting more artefacts does not automatically produce better evidence.3

AI critique or comparison

Provide an AI-generated explanation, solution, or artefact and ask students to identify claims, verify evidence, find omissions, and produce a justified revision. Prefer an educator-vetted artefact so the activity does not depend on a live model producing a useful mistake.

AI-supported production

Where AI is allowed, assess the student’s framing, source selection, verification, disciplinary decisions, and final accountability. A reflective commentary should address specific decisions and failures, not merely list prompts.

Multimodal or scenario-based performance

Students can interpret data, produce a diagram, respond to an unfolding case, or explain the same idea in several modes. Multiple modes can broaden the evidence available, but they must remain accessible and aligned with the construct being assessed.4

Process evidence without surveillance

Drafts, notes, use records, and version history can support learning when they are proportionate and announced in advance. Do not require full AI transcripts, screenshots, keystroke histories, or continuous monitoring by default. These may collect personal or irrelevant data, disadvantage some workflows, and still do not prove authorship.

If work raises an integrity concern, apply the published rule and the normal institutional process. See Why AI detection does not prove authorship.

Checklist before release

  • Every task maps to an approved learning outcome.
  • The AI level follows from that outcome.
  • Allowed tools and actions are concrete.
  • The evidence can be assessed reliably within available time.
  • Students receive practice in any unfamiliar format.
  • Access, accessibility, privacy, and non-AI alternatives are addressed.
  • The rubric rewards knowledge, reasoning, and decisions—not model fluency.
  • The brief explains disclosure, questions, and integrity procedures.
  • High-stakes decisions use more than one appropriate source of evidence.

References & Footnotes

Footnotes

  1. Nicol, D. J., & Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: a model and seven principles of good feedback practice. Studies in Higher Education, 31(2), 199–218. https://doi.org/10.1080/03075070600572090  ↩

  2. Joughin, G. (1998). Dimensions of oral assessment. Assessment & Evaluation in Higher Education, 23(4), 367–378. https://doi.org/10.1080/0260293980230404  ↩

  3. Driessen, E. W., van Tartwijk, J., van der Vleuten, C. P. M., & Wass, V. (2007). Portfolios in medical education: why do they meet with mixed success? A systematic review. Medical Education, 41(12), 1224–1233. https://doi.org/10.1111/j.1365-2923.2007.02944.x  ↩

  4. Ross, J., Curwood, J. S., & Bell, A. (2020). A multimodal assessment framework for higher education. E-Learning and Digital Media, 17(4), 290–306. https://doi.org/10.1177/2042753020927201  ↩

Last updated on