What is ‘Appropriate’ AI use in higher education? maybe we agree more than I thought

So, last week I wrote about the increasing discord, intolerance, frustration and shaming in ‘debates’ about the place of AI in higher education and the prickly path I and my colleagues are trying to navigate. I was looking back on some of the sessions we offered during our university development week and found the slides from an event that might indicate that in the quieter middle there might be more consensus than the noisy, exhausting arguments suggest. One of the things I have become increasingly wary of in conversations about AI in higher education is the phrase ‘that’s clearly (in) appropriate use’, because while any given use/ application’s appropriateness can sound straightforward/ obvious it can of course connote enormous complexity about what we are actually asking students to do, why we are asking them to do it and what the heck we think they should be learning from doing it. Baselines in university regs and guidance and rules and policies read something like: Students should use AI appropriately. Staff should teach students how to use AI appropriately. Assessment guidance should distinguish appropriate from inappropriate use (I’m not even going to get into traffic lights here). At a casual glance it sounds perfectly reasonable until somebody asks pretty obvious questions: appropriate according to whom or appropriate for what?

My colleague, Jonathan Tulloch and I ran a session with colleagues in which we tried to pick our way through the complexity. Academic misconduct is often the starting point but we looked at very specific activities and drew anonymous (then discussed) standpoints from colleagues from across our 5 schools.  We began with some context: Noting a lack of clarity in the catch-all term ‘AI’. We outlined and collected a range of persistent issues and controversies ( hallucination, representation, bias, privacy, trust, environmental impact and more) before we even got to the considerable differences in experience, confidence and belief amongst staff. And, of course, there is disciplinary context. As I have witnessed first hand, applied AI processes can elicit an immediate ‘but this isn’t how we do history!’ response: substitute nursing, law, business, architecture or almost anything else according to taste.

There is also the not insignificant complication that students are already using it (whatever it is) in enormous numbers. The 2026 HEPI/Kortext survey reported that 95 per cent of students were using AI in at least one way for study and 94 per cent were using generative AI to help with assessed work. 68 per cent believed that AI skills were essential for thriving in today’s world, while fewer than half felt that their teaching staff were helping them to develop those skills.  So I maintain that ‘should students use AI? is not a particularly useful question any more though I accept that many argue it is. The more useful, interesting and necessarily nuanced questions concern what kinds of AI use support learning, what kinds begin to displace learning, what students need to be able to do independently and, crucially, whether those of us teaching them actually agree about where those boundaries lie.

Caution and reasoned discussion about these issues and tensions are a fundamental aspect of developing a critical literacy essential to make informed decisions according to the disciplinary and other contextual factors. As it is with all information literacy. The concern about cognitive offloading is not trivial but it’s also not (in my view) a straightforward reason to shut down discussion. Socrates, via Plato, worried that writing would ‘produce forgetfulness’ because those who used it would no longer practise their memory. GenAI gives that classical concern a contemporary revitalisation because the things we can potentially offload now include planning, summarising, drafting, calculating, coding, interpreting and perhaps increasingly quite substantial elements of thinking itself.

I do maintain though that AI definitely can enhance and enable learning and doing things. It’s really helped me get to grips with complexity of operating new machinery I had no previous exposure to and could make sense of the manuals which were pretty incomprehensibly translated from another language.  I have described before ways in which it really helps me personally engage with texts that are presented in unfriendly ways and to process thinking and ideas much more efficiently. It has convenience aspects that are hard to ignore. I spent years trying to get to grips with voice to text tools but never was fully able to use them fluidly and without frustration. This text was dictated into an LLM already pre-set to use British English spelling and with a prompt to remove only faltering elements and to italicise inconsistencies and repetition and then offer bulleted structural advice. The transcript I then cut and pasted into a Word document and I am now working through it and editing it using old school two finger typing (nb. The bulleted lists at the end are AI generated from the transcript which started to feel a bit long but I still wanted to include. My edits to them are minimal). I have spoken to colleagues and students who have found impressive ways to support their work, thinking and learning. It can explain something differently, provide opportunities for practice, help somebody organise what feels like an overwhelming task or lower barriers that have previously prevented participation. This becomes especially important when considering disabled students, given evidence that some students with dyslexia, ADHD and related conditions describe AI as more effective cognitive support than the formal adjustments they have experienced though even in that domain there are starkly contrasting experiences and opinions.  I think this is why I find blanket statements such as ‘AI should support your thinking, not do your thinking for you’ an appealing provocation or starting point but ultimately a little insufficient.

So what did colleagues think? We presented more than 60 colleagues with a series of scenarios and asked them to place each on a five-point continuum. A score of 1 represented complete agreement that the example constituted appropriate or balanced AI use, while 5 represented complete agreement that it constituted excessive use or over-reliance. The important point here is that I am less interested in treating the resulting numbers as precise measurements of how ‘good’ or ‘bad’ a particular practice is than in what their position on the scale tells us about congruence and conflict. Where responses cluster towards 1, there is relatively strong agreement that the practice is appropriate or balanced. Where they cluster towards 5, there is relatively strong agreement that AI is doing too much. It is the movement towards the middle, around 3, that becomes particularly interesting, because this is where we begin to see conflicting perspectives about what constitutes appropriate use.

Colleagues were strongly inclined to regard students generating AI-based quizzes or flashcards from lecture content as appropriate or balanced use, and similarly comfortable with students asking AI to re-explain difficult concepts using everyday examples. There was also strong agreement around a student drafting an essay themselves and then using AI to identify stylistic or grammatical issues.  There seems to be a reasonably coherent principle sitting underneath these responses. In each case AI is doing something useful for the student, sometimes something quite substantial, but it is not obviously taking over the intellectual work that we want the student to undertake. There is assumed grounding as a direct challenge to hallucination potential in each instance.

Planning is also an interesting area. My logic says (as someone with history background) that the planning is where the thinking happens even though many of the policies I have seen say something on the lines of ‘planning ok; writing support not’. Our  colleagues were broadly comfortable with students uploading deadlines and resources to an AI planning tool which then creates a week-by-week schedule, and were reasonably comfortable with AI helping a student define thematic clusters or task sequences at the beginning of a complex project. What they were much less comfortable with was a student asking AI to generate the research plan and then simply following it to the letter.  That latter point is a nuance I think we may have neglected a little and is worth discussion with students engaging in long form writing activities.

Where AI ‘helps me organise, explore or structure my thinking’, colleagues seem broadly comfortable. Where I hand over decisions about what I should think, investigate or do and simply follow what the machine tells me that’s obviously over-reliance. What, though, about a student who uses multiple prompts across multiple tools to assemble and edit a summative essay? Or somebody who asks AI for the first few steps of a multi-stage calculation and then completes the rest themselves? Or a student who uses an AI agent to create the base code for an interactive website and subsequently develops that code to produce a better and more accessible final product? These examples produced responses much closer to the middle of the continuum, which means that colleagues were not speaking with anything like the same degree of unanimity. I have discussed similar scenarios a lot over the last few years and I think it is fair to say the tendency is towards a more nuanced understanding and acceptance of complex workflows even though many policies in the wider academic sphere ask for a record of ‘prompts and outputs’ which is unworkable and, in my view, fails to understand such complexity. The continued disagreement shows that blanket policies are likely to remain unworkable even though may academic colleagues will tell you they just want to know what’s allowed and what’s not. My conclusion is that that is actually an impossible ask when you get to this level of granularity.

While I accept we do need a lot more in the way of foundations I don’t think we can ever happily say ‘here is the institutional guidance telling everybody which side of the line each example belongs on’. The disagreement may be telling us something much more fundamental about the inadequacy of trying to judge AI use without knowing the educational context in which it occurs. If I am trying to establish whether a student can independently identify and execute every stage of a multi-stage calculation, asking AI to get them started may undermine precisely the capability I am trying to assess. If, however, the calculation is merely one component of a much more complex professional problem and the important learning concerns what the student subsequently does with it, I might reach a very different conclusion. This subtlety can really only be determined by the disciplinary experts; context is everything but also problematic if those same staff lack confidence or core understanding of capabilities of even basic tools, let alone frontier models.

So perhaps one of the most useful conclusions from this exercise is that we are sometimes asking the wrong question. ‘Is this an appropriate use of AI?’ is probably not enough. ‘Is this an appropriate use of AI for this student, undertaking this activity, for this purpose, at this point in their learning? It’s better but a lot more fiddly. I STILL think we’ll talk a lot less about LLMs in a few years and reach a broad level of common understanding  and normalisation of GenAI tools (agents looming much larger on the threat scale) but we have to get there first. And before that we’ll likely have to navigate regulatory panic of some sort and a consolidation of ‘useful’ AI in the hands of fewer players (pretty safe prediction I think based on reading the news nothing more). We can start with where there is congruence though.  AI use that helps students organise, practise, understand, brainstorm or refine their own work is ok in a lot of contexts (I know it’s not without controversy but if we define specific parameters with, in particular, assessment validity firmly in mind) we might get somewhere that looks consistent.  Nothing is automatically appropriate in every context, but they seem to provide useful territory from which to begin developing shared principles.

Agency is key too as there’s that important distinction between interrogating with AI and obeying it, between using an output as a starting point and treating it as an answer, and between retaining responsibility for intellectual decisions and quietly handing those decisions over. Also, the amount of AI being used may actually be less important than what is being outsourced. A student might spend half an hour in quite an elaborate conversation with AI which causes them to question assumptions, consider alternatives and substantially improve their understanding. Another student might type a single prompt and receive exactly the intellectual work that the assessment was designed to elicit from them. Counting prompts, words or percentages of ‘AI contribution’ therefore tells us little. What our colleagues said suggests we should pay particular attention to the examples sitting towards the middle of the continuum, because these expose the assumptions we are making about learning. Rather than trying immediately to eliminate those disagreements through regulation, we should probably be discussing why reasonable colleagues see the same use differently.

So, as I work with colleagues supporting wider staff and student AI literacy work, what should we do?

I think at instituional level we need to:

  • Resist producing ever longer catalogues of permitted and prohibited AI behaviours, and instead establish clear principles around learning, agency, transparency, inclusion and accountability.
  • Allow those principles to be interpreted meaningfully within disciplines, recognising that appropriate use will differ according to what students are learning and why. We have gone with separate staff facing and student facing principles and we are testing that flex and complexity as we go.
  • Retain clear boundaries where they are necessary, particularly in relation to academic integrity, while avoiding the assumption that one institutional definition of appropriate use will work in every context.
  • Use realistic AI scenarios in staff development to explore areas of congruence and conflict, not to reveal the ‘correct’ answer, but to help colleagues articulate why they draw boundaries in different places.

I also think there has to be discussion at course/ module team level. Much of this ground has been covered before but I’m finding it useful to pull it together in one place:

  • Review actual assessments and identify the intellectual or professional work that students are expected to demonstrate.
  • Ask explicitly which elements of that work AI can legitimately support and which students need to be capable of doing themselves.
  • Consider where AI use might increase rather than diminish authenticity because it reflects contemporary professional practice.
  • Recognise progression. AI use that constitutes inappropriate outsourcing at Level 4 might represent sensible professional practice at Level 6 or postgraduate level because what we expect students to know and do has changed.
  • Pay particular attention to those uses sitting in the contested middle, because disagreement about them may reveal different assumptions about the purpose of an assessment.

Individual lecturers could:

  • Explain why particular uses of AI are appropriate or inappropriate rather than simply telling students what they are allowed to do.
  • Connect AI guidance explicitly to learning. For example: ‘You can use AI to brainstorm possible approaches, but you need to rationale your choice of approach, be able to verbally defend it in labs and undertake the final analysis yourself because developing and demonstrating that analytical capability is the purpose of this task.’
  • Help students understand that the important question is often not how much AI they have used, but what intellectual work they have handed over to it.
  • Encourage students to interrogate, challenge and adapt AI outputs rather than simply accepting them.

For some things there is a broad congruence in thinking but this exists more at general, sweeping principles levels. The more complex things get the less likelihood of a visible line existing. I think we need to do a lot more work to make the reasons for drawing different lines visible, discussable and educationally defensible while making sure our students are meaningfully and deeply engaging with the sort of disruption connoted by these tech.

Leave a comment