AI as a Craft Assessment Guide
Purpose of This Guide
This guide covers how to design assessments that involve AI in ways that are sound for learning and honest with students. The companion Teaching Guide covers preparation, delivery, and reflection, and AI as a Craft introduces the approach both are built on.
Assessment design follows the same five stages as everything else in the craft: Define, Decide, Design, Evaluate, Reflect. The aim isn't a checklist to complete once. It's for these stages to become second nature, so that by the time you write an assessment brief you have already answered the questions students will ask about it.
One principle above all the others: never design an assessment that rewards students for removing themselves from their own work. The whole craft is about putting yourself in the work. Assessment should measure how well they do it.
Define: The Cognitive Work
Before you consider AI at all, be crystal clear about what students are actually learning to do. If you can't answer these questions clearly, your assessment needs redesigning regardless of AI.
Question 1: What are students learning to do?
- Technical skill execution?
- Critical analysis and evaluation?
- Creative problem-solving?
- Strategic decision-making?
- Synthesis of complex information?
Question 2: Why does this assignment exist?
- To demonstrate procedural knowledge?
- To show understanding of concepts?
- To apply theory to novel situations?
- To develop professional judgement?
Question 3: What evidence would prove they can do it?
- Observable performance?
- A tangible artefact?
- Articulation of reasoning?
- Iterative development process?
If you can't answer these clearly, STOP. Your assessment is broken regardless of AI.
A Note on Bloom's Taxonomy in Relation to AI
The bottom four levels (Remember, Understand, Apply, and Analyse) are increasingly within AI's capability. A student who only needs to recall facts, summarise concepts, complete standard problems, or identify patterns can do all of that with AI assistance in ways that are genuinely difficult to detect or prevent.
The top two levels are different in kind, not just degree. Evaluate requires students to make and justify judgements: to critique, defend a position, and reason under genuine ambiguity. Create requires students to produce something original reflecting their own thinking and choices.
If your assessments don't reach Evaluate and Create, you're testing what AI can already do. Put another way, you're asking "is it right?" when the questions worth asking of a student's work are the craft's pair: "Is it right and is it you?"
Decide: Choose the Door
Based on your Define answers, choose which of the four doors makes sense for this assessment. The door is a deliberate design choice, and you will explain it to students, so choose one you can defend.
AUTOMATE: AI performs defined tasks
Appropriate when: You want students to demonstrate task delegation, quality control, and efficiency.
Examples: Using AI to generate first drafts, data summaries, basic code snippets.
Assessment focus: Selection judgement, output evaluation, refinement skill.
COLLABORATE: Human-AI thinking partnership
Appropriate when: You want students to demonstrate iterative thinking, creative development, complex problem-solving.
Examples: Co-writing analysis, collaborative coding, research synthesis.
Assessment focus: Steering the collaboration, critical engagement, value-adding beyond AI capability.
AGENTIC: AI acting independently
Appropriate when: You want students to demonstrate systems thinking, design judgement, ethical consideration, or professional decision-making under realistic conditions.
Examples: Students building agents (chatbots, AI tutors, interactive experiences), or students working inside a simulation you have built. Both directions are covered below.
Assessment focus: Behavioural design and responsibility for outcomes, or judgement demonstrated within the simulation.
NO AI: Use your own creativity
Appropriate when: The cognitive work requires demonstrating personal capability independent of AI. This is not a retreat from the modern world. It is a positive design choice: some capabilities must be verifiably the student's own before it means anything to amplify them.
Examples: Timed exams, presentations, practical demonstrations, portfolios with verifiable process evidence.
Assessment focus: Authentic performance that establishes human capability.
Check yourself: Are you choosing this door because it's sound for learning, or because traditional assessment is familiar?
Design: Your Assessment Task and Rubric Weighting
The Skeleton Underneath: The 4 Ds
The rubrics in this guide assess four competencies from the AI Fluency Framework (Dakan & Feller, 2025). These map directly onto the stages of the craft:
Delegation (developed at Define and Decide): making sound decisions about whether, when, and through which door to use AI, and what stays human.
Description (developed at Design): communicating with AI effectively. Who you are, what you want, why, refined through iteration, or configured into agent behaviour.
Discernment (developed at Evaluate): critically evaluating AI outputs and processes. Is it right, and is it you? Catching errors, biases, and the flat average-of-everyone voice.
Diligence (developed at Reflect): using AI responsibly, transparently, and answerably. Fact-checking, acknowledging AI use, owning the outcome.
Automate Assessments
Students use AI to independently perform clearly defined tasks. Focus is on efficiency, quality control, and appropriate task selection.
Typical task: Give students a defined outcome and allow them to choose and use AI tools to complete it, plus a brief reflection on tool choices and quality evaluation.
Suggested weighting: Delegation 30% | Description 20% | Discernment 40% | Diligence 10%
Example
"Use AI tools of your choice to create a competitive analysis of three companies in the renewable energy sector. Submit: (1) The analysis report, (2) A 300-word reflection explaining which tools you chose, why, and how you evaluated output quality."
Collaborate Assessments
Students work iteratively with AI as a thinking partner. The process involves back-and-forth dialogue where both human and AI contribute, and the student's own judgement is visible in the steering.
Typical task: Give students a complex, open-ended problem requiring iterative development, with process documentation.
Suggested weighting: Delegation 10% | Description 30% | Discernment 30% | Diligence 30%
Example
"Collaborate with AI to develop a theoretical framework for analysing gender representation in video games. Submit: (1) Your framework paper, (2) Conversation log documenting your theoretical development process, (3) 500-word critical reflection on where AI helped you explore theoretical connections and where your disciplinary expertise shaped the analysis."
Agentic Assessments, Direction One: Students Build the Agent
Students configure AI systems to act independently and interact with others. Focus is on designing AI behaviour, not just getting outputs. The six boxes (Purpose, Nature, Subject, Setting, Communication, Mechanisms) give students a design language for this.
Typical task: Design and configure an AI agent for a specific purpose, with documentation and ethical reflection.
Suggested weighting: Delegation 25% | Description 25% | Discernment 10% | Diligence 40%
Example
"Design an AI tutor that helps students learn molecular orbital theory through Socratic questioning. Submit: (1) The configured AI tutor, (2) Testing documentation showing how it responds to common misconceptions, (3) Ethical impact statement addressing risks of reinforcing errors, accessibility, and transparency about AI limitations."
No AI Assessments
The 4 Ds framework does not apply. Use traditional disciplinary assessment criteria. However, consider: Can you verify the work is the student's own? Does the format test what you want to test? Are you excluding AI because it's sound for learning, or because traditional assessment is familiar?
Quick Reference: Suggested Rubric Weightings
| Door | Delegation | Description | Discernment | Diligence |
|---|---|---|---|---|
| Automate | 30% | 20% | 40% | 10% |
| Collaborate | 10% | 30% | 30% | 30% |
| Agentic (students build) | 25% | 25% | 10% | 40% |
| Agentic (simulation) | Primarily disciplinary criteria; see below | |||
| No AI | Traditional criteria | |||
These weightings are starting points. Adjust based on your specific disciplinary context and learning outcomes.
Agentic Assessments, Direction Two: Students Inside Your Simulation
The second direction reverses the roles: you build the agent, and students work inside it. Design an agent that plays a part (a patient presenting with ambiguous symptoms, a hostile client, a community objecting to a planning proposal, an archive with gaps and contradictions) and load in a scenario. Students undertake the encounter, applying taught theory to a situation that responds, resists, and occasionally misleads.
This is assessment at the top of Bloom's, where this guide has been pointing all along. A simulation cannot be answered from recall. It demands judgement under ambiguity, justified decisions, and live application: Evaluate and Create, performed rather than described.
Building One
Use the six boxes from AI as a Craft to design the character once. The Setting and Mechanisms boxes are where scenarios load: change the presenting symptoms, the client's demands, the year the archive burns down, and the same simulation carries through a whole module. Test it the way you'd assess a student doing the same: run it against common misconceptions, check where it breaks, document what it cannot do.
Design honesty into the ambiguity: a good simulation can include unreliable information, because reality does. The patient misremembers, the client exaggerates, the sources conflict. Spotting this becomes part of what you assess, and discernment becomes disciplinary skill.
What Students Submit
- The transcript of the encounter
- A decision log justifying each significant judgement against taught theory and available evidence
- A reflection on what they would do differently, and on where the simulation itself fell short of reality
How to Weight It
Assess primarily on disciplinary criteria: the quality of the decisions, not the typing. The simulation is the venue, not the subject. Where AI fluency competencies appear, they appear as disciplinary judgement (discernment about the agent's unreliable information) rather than as separate rubric lines. That final reflection on the simulation's limits is worth marks: a student who can articulate where the simulation diverges from reality is demonstrating exactly the professional judgement the assessment exists to develop.
Example
"You will conduct a consultation with a simulated patient presenting with ambiguous symptoms consistent with several conditions covered this semester. The patient will answer questions in character and may misremember details. Submit: (1) Your consultation transcript, (2) A decision log justifying your questions, working diagnosis, and recommended next steps against the evidence available to you, (3) A 500-word reflection on what you would do differently and where the simulation differed from a real consultation."
One Honest Caution
A simulation is teaching material like any other, so the corridor test applies to it: you must understand the scenario deeply enough to debrief it, defend its design, and answer for what the agent said in your name. Pilot it yourself before students touch it, and always give students a way to flag when the simulation behaves strangely. You are answerable for your agent's behaviour, exactly as your students are answerable for theirs.
Communicating with Students
Before presenting your assessment, help students understand that developing their own craft with AI builds genuine advantage. Students need to see that deferring their thinking to AI makes them less valuable, not more efficient. Work that could have been anyone's is worth what anyone can make it for: nothing.
The Reality Students Are Entering
What employers want: People who know when to use AI and when not to, who can get better results from AI than their competitors, who can add value beyond what AI can do alone, and who take responsibility for outcomes.
What they don't want: People who can't think without AI, who can't tell good AI output from rubbish, who produce work indistinguishable from everyone else's AI-generated mediocrity, or who hide behind "the AI made a mistake."
For Automate Assessments
Your advantage comes from: judgement in knowing which tasks to hand over, quality control in spotting errors before they damage your reputation, efficiency in getting good results faster, and accountability in vouching for your work.
For Collaborate Assessments
Your advantage comes from: iterative thinking that explores ideas you wouldn't reach alone, creative leverage letting you focus on the harder thinking, critical engagement knowing when AI improves your work, and an authentic voice producing work that is recognisably yours.
For Agentic Assessments
Your advantage comes from: systems thinking understanding how AI behaviour affects people, design judgement configuring AI that actually helps, ethical foresight anticipating potential harm, and responsibility for what you create. In simulations, your advantage is simpler still: the ability to think on your feet when the scenario talks back.
For No AI Assessments
Your advantage comes from: genuine expertise that's yours not borrowed, authentic judgement based on your own thinking, confidence in performing without tools, and trust built on personal capability.
The bottom line: Treating AI as a shortcut makes you ordinary. Anyone can type a prompt. Genuinely skilled people know when to use AI, how to use it well, and when to rely on their own expertise. That is why AI is a craft, not just a skill. Your hand should be guiding it into producing something that presents your own ingenuity.
Academic Integrity with AI
Suspected Unauthorised AI Use
Don't: Immediately accuse or penalise based on suspicion alone.
Do: Invite the student to a conversation to discuss their work. Ask them to explain their process and key decisions, probe their understanding of specific content, ask them to elaborate on points or apply concepts to new scenarios, and give them opportunity to demonstrate authentic understanding. This is the corridor test, applied to students: can they go deeper without the artefact in front of them?
Important: If a student can demonstrate genuine understanding and engagement with the work, the process they used to get there is less critical than the learning achieved.
Prevention is better than detection: If you're frequently suspecting AI misuse, your assessment design may need revisiting. An assessment that can be hollowed out invisibly is an assessment testing the wrong things.
Agentic Assessments: When AI Agents Fail
Distinguish between technical failure (the system doesn't work as intended) and design failure (the system works but is poorly designed).
Don't penalise for: Technical glitches outside their control, single-instance failures, platform limitations they couldn't reasonably overcome.
Do assess on: Quality of configuration and instructions, evidence of iterative testing, understanding of why failures occurred, quality of documentation and ethical considerations.
The same fairness applies in reverse for simulations: if your simulation misbehaves during an assessment, that is your technical failure, not the student's. Have a contingency, and never let a student's mark depend on your agent having a bad day.
Essential Assessment Brief Elements
- Which door this assessment uses: Automate, Collaborate, Agentic, or No AI
- What evidence is required (process documentation, AI use statements, conversation logs, decision logs)
- How to demonstrate authentic engagement with the work
- What constitutes misconduct in the context of this specific assessment
Generic statements like "appropriate use of AI" are insufficient. Specify exactly what you expect students to document and submit.
Further Support
For assistance with assessment design, contact your Learning Designer. For understanding AI risk levels, talk to Prof. Akeyo. To think a design through with a companion, meet Keith, and for the approach itself, start with AI as a Craft.
This guide is designed as a living document. The approach will evolve as AI technologies and teaching practices develop.
This guide was written and developed by David Calum Millar, refined and evaluated with the assistance of a custom agent called Morna, running on Claude (Anthropic) and trained on the Creative Compass framework by David Calum Millar. It applies the AI as a Craft approach (Millar, D.C., 2026), building on the AI Fluency Framework (Dakan, Rick and Feller, Joseph. "Framework for AI Fluency (Practical Summary Document)," Version 1.1, Ringling.edu/ai/, 2025) along with Bloom's Taxonomy (Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001)) and the Learning Design Compass framework (Millar, D.C., 2026). Special thanks to Dr. Ross McKenzie who helped to review and provide feedback to an earlier version of this guide.