A polished essay no longer tells an assessor enough about how a student reached the final result. It may represent careful independent work, substantial AI assistance, or a mixture that cannot be reconstructed from the document alone.
That uncertainty has made AI in university assessments a curriculum problem, not merely an academic-integrity problem. Blanket bans are difficult to enforce outside controlled settings. Unrestricted use can also let students bypass the reasoning or practice an assignment was supposed to develop.
Universities now have to decide which abilities students must demonstrate independently, where AI may assist them, and when responsible AI use should itself become part of the assessment.
A finished submission is no longer sufficient evidence
Higher education has long placed substantial weight on finished artifacts: essays, reports, coding projects, presentations, research summaries, and take-home examinations. Generative AI can now contribute to all of them.
Changing an essay topic or requesting a personal reflection does not solve the problem. AI systems can produce plausible reflections, revise code, reorganize arguments, imitate a supplied style, and adapt an answer to a local scenario. Adding “use critical thinking” to the instructions offers no assurance that the student did the thinking being assessed.
A stronger starting question is: What evidence would demonstrate that this student has achieved the learning outcome?
The answer changes by discipline. Foundational knowledge may require a supervised test. Clinical judgment may need observation during a simulation or placement. A design project could include prototypes, testing records, revisions, and an explanation of rejected choices. A strong final product remains valuable, but it becomes one part of the evidence rather than the whole case.
Australia’s Tertiary Education Quality and Standards Agency frames the issue around two responsibilities: preparing students to participate ethically in a society where AI is widely available and making trustworthy judgments about what they have learned.
Universities are still working out how to meet both responsibilities. A 2025 UNESCO survey collected 400 responses from UNESCO Chairs and UNITWIN Networks in 90 countries. Nineteen percent said their institutions already had a formal AI policy, while 42% reported that a framework was under development. The sample was limited to UNESCO-affiliated networks, so the figures should not be treated as a census of global higher education. They do, however, show how unfinished the policy work remains.
Set the AI rule according to what the task must prove

One rule cannot sensibly cover a closed-book mathematics test, a media-production project, and a postgraduate research proposal. The more practical approach is to classify assessments by purpose.
| Assessment mode | How AI is treated | Where it fits |
|---|---|---|
| Secure or controlled | AI is prohibited or restricted under enforceable conditions | Foundational knowledge, essential calculations, live performance, safety-critical decisions |
| AI-assisted | Defined assistance is allowed, but the student remains responsible for the work | Brainstorming, language review, coding support, formative feedback, preliminary analysis |
| AI-integrated | AI use is required because evaluating or directing it is part of the learning outcome | Auditing outputs, comparing systems, testing reliability, documenting human oversight |
UCL uses a similar three-category structure: AI cannot be used, AI may play an assistive role, or AI has an integral role. A useful qualification is sometimes missed when institutions copy this model: UCL presents the categories as guidance, not as a single stand-alone policy. The staff member setting an assessment determines the conditions, and the instructions in the assessment brief take precedence.
The University of Sydney uses a two-lane framework. Secure assessments provide evidence that students have achieved important outcomes. Open assessments support learning with contemporary technologies, including generative AI where relevant.
Neither model divides an entire degree into “AI” and “no AI.” Both recognize that assessment serves different purposes at different points.
Start with the program, not isolated assignments
Redesigning one assignment at a time often produces contradictory rules. A student may be told to avoid AI completely in one module, disclose every prompt in another, and use it freely in a third without any explanation of acceptable practice.
Program teams should map where each important outcome is introduced, practised, and finally demonstrated. They can then select the points where independent competence needs to be verified.
A business degree, for example, might permit AI-assisted case work throughout several modules. Students could use approved tools to generate scenarios, organize evidence, or review a draft recommendation. At selected stages, they might complete a supervised analysis or defend a recommendation before an assessor.
This is a better use of secure assessment than moving large parts of a degree back into examination halls. Controlled tasks are worth the cost when they protect an essential capability. They add little when introduced simply because staff feel uneasy about unsupervised work.
Assessment designs that expose the student’s reasoning
No format can prove authorship perfectly. Better designs collect evidence from more than one stage or require students to explain, revise, test, or apply what they submit.
Build a connected sequence
Instead of collecting one final report, an instructor might assess a proposal, an initial evidence set, a response to feedback, and the finished submission.
The stages need to affect one another. If every component can be generated independently shortly before the deadline, the extra paperwork reveals little. Asking students to explain why new evidence changed a decision is more useful than requiring five loosely connected documents.
Version histories and process notes can support this approach, but they are easy to overrate. Mandatory screenshots of every prompt create marking work and may disadvantage students using different platforms or assistive technologies. A prompt archive is not proof of learning.
Use oral checks where they add evidence
A short oral component can test whether a student understands submitted work. An engineering student might explain a design constraint. A programmer could modify a function after a requirement changes. A history student might defend the interpretation of a primary source.
These checks should measure a published learning outcome, not operate as surprise interrogations. Students need to know the format and criteria in advance. Universities must also provide appropriate adjustments or equivalent formats for students who need them.
Adding a viva to every assignment would create serious scheduling and staffing problems. Oral assessment is most useful at major checkpoints, in smaller advanced courses, or when an established academic process requires clarification of a student’s work.
Make AI output the object of analysis
Where AI literacy belongs in the curriculum, students can be asked to inspect an AI response rather than submit it as their answer.
A law student might verify whether cited cases exist and apply to the issue. A statistics student could identify invalid assumptions in an AI-generated analysis. A media student might examine provenance, bias, disclosure, and ownership questions surrounding a generated image.
The marks should reward verification, correction, and disciplinary judgment. Clever prompting is often overvalued. Outside courses where prompt design is an intended capability, the student’s decisions before and after an AI response matter more than the wording used to obtain it.
Tie work to specific evidence
Generic assignments are easier to outsource. More useful tasks involve a defined dataset, laboratory result, archival collection, field observation, client constraint, or decision made during class.
Specificity does not prevent AI use; students may still upload materials to an AI service. Its value lies in giving assessors a clearer way to examine how evidence supports the student’s judgment.
There is an immediate privacy limit. Patient information, personal data, unpublished research, confidential client material, and restricted assessment content should not be entered into public AI services without explicit institutional approval.
Change the rubric before adding more policing
Traditional rubrics often reward structure, grammar, fluent prose, and presentation. Those qualities still matter, but AI can improve all four without demonstrating much understanding.
More weight may need to move toward:
- Correct use of disciplinary concepts
- Quality and traceability of evidence
- Justification of significant decisions
- Recognition of uncertainty
- Verification of AI-generated claims
- Response to feedback
- Ability to explain or modify the work
- Accurate disclosure of permitted AI assistance
Assessors should not treat awkward writing as evidence of authenticity or polished writing as evidence of misconduct. Both assumptions are unreliable.
TEQSA’s 2026 work on AI-integrated learning places particular emphasis on evaluative judgment, critical thinking, and ethical reasoning. Those capabilities are more durable than expertise with a particular chatbot interface.
Put the AI rules beside the assignment
A university-wide policy cannot tell a student what is permitted in a particular laboratory report, translation exercise, or design project. Each assessment brief should make the local rule explicit.
Students need to know whether AI is prohibited, optional, or required; which activities are allowed; how assistance must be disclosed; and what information must not be uploaded. The brief should also identify approved services and provide a contact for questions before submission.
“AI may be used appropriately” is too vague. So is “all AI use must be cited.” A useful policy distinguishes between acknowledging assistance and citing evidence. AI might help organize a draft, but its output should not replace the journal articles, legislation, datasets, or primary documents supporting an academic claim.
Policies based on behaviour are easier to maintain than lists of product names. AI features can appear inside search engines, office suites, learning platforms, coding environments, and specialist applications with little notice.
A detection score is not a verdict
AI-detection software may provide a signal, but it cannot establish authorship by itself. Turnitin’s documentation acknowledges that its model may misidentify human-written, AI-generated, and AI-paraphrased text. The company states that the result should not be the sole basis for adverse action.
UCL has taken a firmer position. Its guidance says the university is not investing in generative-AI detection software and that staff must not upload student work to such tools because of intellectual-property and personal-data concerns.
A fair investigation still requires human judgment and the institution’s established procedure. Relevant evidence might include drafts, source use, previous work, version history, a conversation with the student, or a demonstration of the underlying skill. Treating a percentage as proof is both weak assessment practice and a procedural risk.
Equity and access cannot be an afterthought
An AI-integrated assignment may become unfair when some students have paid subscriptions, faster internet access, newer devices, or access to services unavailable in their country.
If a particular AI capability is required, the university should provide an approved tool or a genuinely equivalent route through the assessment. Students should not have to purchase premium access to remain competitive or create personal accounts with unclear data practices.
The same care applies to accessibility. A secure assessment may prohibit generative AI while still allowing approved assistive technology. An oral assessment may need an adjusted or equivalent format. Disclosure requirements should not force students to reveal disability-related information.
UNESCO’s guidance places privacy, inclusion, equity, human agency, and ethical validation at the center of educational AI use. These are design requirements, not issues to handle after an assignment has been launched.
A workable route from policy to practice
Universities do not need to redesign every task in one academic year. The sensible first step is to identify heavily weighted take-home work, high-enrolment modules, professional accreditation requirements, and outcomes currently judged through one unverified artifact.
Program teams can then classify each task, protect the most important independent checkpoints, and redesign selected open assessments. Briefs and rubrics should be revised before new enforcement measures are introduced.
Pilots should track more than suspected AI use. Staff workload, accessibility problems, student understanding, appeals, marking consistency, and the quality of learning evidence all matter. A redesign has failed if it produces mountains of documentation, reduces a rich subject to frequent memory tests, or leaves instructors with an impossible volume of oral examinations.
The strongest plans also involve students, educational designers, librarians, disability services, and data-governance staff. Rules that appear clear in a committee document can become confusing once applied across different disciplines and student circumstances.
Final Thoughts
Effective AI in university assessments begins with a clear decision about what evidence matters. For most institutions, program-level mapping is the best place to start—not detector procurement and not a blanket return to closed examinations.
Secure a limited number of important demonstrations, redesign open tasks around evidence and judgment, and state the AI rules beside every assignment. That approach protects the meaning of a qualification while preparing students to use technology critically, even as the tools themselves continue to change.





