Overview
Policies are usually evaluated by reading them. This project evaluates them by
using them: a fixed set of fictional student AI-use cases is applied to
different university policies, and the resulting decisions are compared. If the
same behaviour is misconduct at one institution, acceptable at another and
undecidable at a third, that inconsistency becomes visible and measurable.
Problem
Universities assert that their generative-AI rules are clear and fair. But
clarity is only testable at the point of decision. There is little evidence on
whether different institutional policies, applied to identical facts, converge
on the same outcome — or whether outcomes depend more on where a student
happens to be enrolled than on what they actually did.
My role
Sole researcher: designed the vignette set, selected the policy corpus,
performed the structured application of each policy to each case, and coded
the outcomes. [Add collaborator details if applicable.]
Research question
Do different university generative-AI policies, applied to the same student
behaviour, produce consistent decisions — and where they diverge, what features
of the policies explain the divergence?
Method
- Construct a set of fictional but realistic student AI-use vignettes — 75
scenario-based cases built to date (see the Case Portfolio for the
published examples).
- For each policy in the corpus, answer a fixed sequence of questions about
each vignette using only the policy text.
- Record the outcome for every policy–case pair using four categories.
- Analyse patterns of agreement and divergence across institutions and case
features.
Outcome categories
Each policy–case pair is classified as one of:
- Acceptable — the policy clearly permits the behaviour
- Misconduct — the policy clearly prohibits the behaviour
- Disclosure problem — the use itself is permitted but the acknowledgement requirements were breached
- Unclear — the policy does not determine an outcome for these facts
Vignette design, structured decision protocol, spreadsheet-based outcome
matrix, qualitative memo-writing on divergent pairs.
Process
Vignettes are held constant; only the policy varies. Every judgement must cite
the specific policy clause relied on, and pairs where no clause decides the
case are recorded as unclear rather than resolved by intuition — the
inability of a policy to decide is itself a finding. I log every coding
decision against its source clause, producing an audit record a second
reviewer could check.
Findings
Findings are reported only once verified against the completed outcome matrix.
- [Add verified finding here]
- [Add verified finding here]
Outputs
- Policy–case outcome matrix — [in preparation]
- The fictional case set, published in the Case Portfolio section of this site
- [Add paper or report details when confirmed]
Impact
- [Add verified impact here]
Limitations
- Decisions are made from policy text alone; real decision-makers use context,
precedent and discretion the study cannot capture.
- The vignette set cannot cover the full space of student behaviour.
- Single-analyst application of policies limits inter-rater reliability claims;
the protocol is designed so the exercise can be replicated.
Lessons learned
- Forcing every decision to cite a clause exposes how much everyday integrity
decision-making rests on unwritten interpretation.
- [Add further lessons as the project concludes]
- [Add related publication when available]
Downloadable materials
- [Add the decision protocol to public/downloads/ and list it in the frontmatter
downloads field]