When traffic lights need to be shades of grey

In my previous post, I reflected on a session Jon and I ran with more than 60 colleagues exploring what constitutes appropriate AI use in higher education. We found more agreement than I expected, but also some revealing differences when it came to complex workflows involving such things as planning work, complex calculations and summatively assessed work. My central argument was that appropriateness depends on what students are learning, what they need to demonstrate independently and where they are in their learning. The thing is, and I am told this often, it’s all very well drawing conclusions in theory but what does that mean in practice? As we work towards wholesale curriculum evaluation and practice-based (re)design of teaching, learning & assessment practices under the PMPC (Practice Makes Professional Curriculum) umbrella, staff want clarity, students need consistency and institutions need workable guidance that is cognisant of the complexity and nuance necessary when thinking about the whole AI thing. So, building on those discussions, here is some of what I am recommending at institutional, course, module and individual lecturer levels.

I think at instituional level we need to:

  • Resist producing ever longer catalogues of permitted and prohibited AI behaviours, and instead establish clear principles around learning, agency, transparency, inclusion and accountability.
  • Allow those principles to be interpreted meaningfully within disciplines, recognising that appropriate use will differ according to what students are learning and why. We have gone with separate staff facing and student facing principles and we are testing that flex and complexity as we go.
  • Retain clear boundaries where they are necessary, particularly in relation to academic integrity, while avoiding the assumption that one institutional definition of appropriate use will work in every context.
  • Use realistic AI scenarios in staff development to explore areas of congruence and conflict, not to reveal the ‘correct’ answer, but to help colleagues articulate why they draw boundaries in different places.

I also think there has to be discussion at course/ module team level. Much of this ground has been covered before but I’m finding it useful to pull it together in one place:

  • Review actual assessments and identify the intellectual or professional work that students are expected to demonstrate.
  • Ask explicitly which elements of that work AI can legitimately support and which students need to be capable of doing themselves.
  • Consider where AI use might increase rather than diminish authenticity because it reflects contemporary professional practice.
  • Recognise progression. AI use that constitutes inappropriate outsourcing at Level 4 might represent sensible professional practice at Level 6 or postgraduate level because what we expect students to know and do has changed.
  • Pay particular attention to those uses sitting in the contested middle, because disagreement about them may reveal different assumptions about the purpose of an assessment.

Individual lecturers could:

  • Explain why particular uses of AI are appropriate or inappropriate rather than simply telling students what they are allowed to do.
  • Connect AI guidance explicitly to learning. For example: ‘You can use AI to brainstorm possible approaches, but you need to rationalise your choice of approach, be able to verbally defend it in labs and undertake the final analysis yourself because developing and demonstrating that analytical capability is the purpose of this task.’
  • Help students understand that the important question is often not how much AI they have used, but what intellectual work they have handed over to it.
  • Encourage students to interrogate, challenge and adapt AI outputs rather than simply accepting them.
  • And, of course, model critical use and engagement

For some things there is a broad congruence in thinking but this exists more at general, sweeping principles levels. The more complex things get the less likelihood of a visible line existing. I think we need to do a lot more work to make the reasons for drawing different lines visible, discussable and educationally defensible while making sure our students are meaningfully and deeply engaging with the sort of disruption connoted by these tech.

None of this removes the need for clear boundaries, and I certainly don’t think every decision about AI use should be left to individual lecturers or students. But I do want to get away from blanket bans and decisions not discussed with peers. I think traffic light systems have had an important part to play in evolving thinking but we need to acknowledge the lack of fit in so many (and increasingly complex) circumstances.

I would love to be able to say that every student understands the rationale for a teaching approach, assessment design and core principles dictating specific, tailored recommendations in relation to most appropriate sources and tools (including AI) recommended for completing a task or engaging with a topic.  But perhaps we need to become more comfortable with the idea that consistency does not always mean uniformity. We can agree on shared educational principles while recognising that their application will look different across disciplines, levels and assessments. By making differences visible, through discussion and then justified in terms of learning we have a chance of seizing the moment to effect positive changes across curricula and assessments which (to bang one of my favourite drums) is overdue anyway.

What is ‘Appropriate’ AI use in higher education? maybe we agree more than I thought

So, last week I wrote about the increasing discord, intolerance, frustration and shaming in ‘debates’ about the place of AI in higher education and the prickly path I and my colleagues are trying to navigate. I was looking back on some of the sessions we offered during our university development week and found the slides from an event that might indicate that in the quieter middle there might be more consensus than the noisy, exhausting arguments suggest. One of the things I have become increasingly wary of in conversations about AI in higher education is the phrase ‘that’s clearly (in) appropriate use’, because while any given use/ application’s appropriateness can sound straightforward/ obvious it can of course connote enormous complexity about what we are actually asking students to do, why we are asking them to do it and what the heck we think they should be learning from doing it. Baselines in university regs and guidance and rules and policies read something like: Students should use AI appropriately. Staff should teach students how to use AI appropriately. Assessment guidance should distinguish appropriate from inappropriate use (I’m not even going to get into traffic lights here). At a casual glance it sounds perfectly reasonable until somebody asks pretty obvious questions: appropriate according to whom or appropriate for what?

My colleague, Jonathan Tulloch and I ran a session with colleagues in which we tried to pick our way through the complexity. Academic misconduct is often the starting point but we looked at very specific activities and drew anonymous (then discussed) standpoints from colleagues from across our 5 schools.  We began with some context: Noting a lack of clarity in the catch-all term ‘AI’. We outlined and collected a range of persistent issues and controversies ( hallucination, representation, bias, privacy, trust, environmental impact and more) before we even got to the considerable differences in experience, confidence and belief amongst staff. And, of course, there is disciplinary context. As I have witnessed first hand, applied AI processes can elicit an immediate ‘but this isn’t how we do history!’ response: substitute nursing, law, business, architecture or almost anything else according to taste.

There is also the not insignificant complication that students are already using it (whatever it is) in enormous numbers. The 2026 HEPI/Kortext survey reported that 95 per cent of students were using AI in at least one way for study and 94 per cent were using generative AI to help with assessed work. 68 per cent believed that AI skills were essential for thriving in today’s world, while fewer than half felt that their teaching staff were helping them to develop those skills.  So I maintain that ‘should students use AI?’ ‘is not a particularly useful question any more though I accept that many argue it is. The more useful, interesting and necessarily nuanced questions concern what kinds of AI use support learning, what kinds begin to displace learning, what students need to be able to do independently and, crucially, whether those of us teaching them actually agree about where those boundaries lie.

Caution and reasoned discussion about these issues and tensions are a fundamental aspect of developing a critical literacy essential to make informed decisions according to the disciplinary and other contextual factors. As it is with all information literacy. The concern about cognitive offloading is not trivial but it’s also not (in my view) a straightforward reason to shut down discussion. Socrates, via Plato, worried that writing would ‘produce forgetfulness’ because those who used it would no longer practise their memory. GenAI gives that classical concern a contemporary revitalisation because the things we can potentially offload now include planning, summarising, drafting, calculating, coding, interpreting and perhaps increasingly quite substantial elements of thinking itself.

I do maintain though that AI definitely can enhance and enable learning and doing things. It’s really helped me get to grips with complexity of operating new machinery I had no previous exposure to and could make sense of the manuals which were pretty incomprehensibly translated from another language.  I have described before ways in which it really helps me personally engage with texts that are presented in unfriendly ways and to process thinking and ideas much more efficiently. It has convenience aspects that are hard to ignore. I spent years trying to get to grips with voice to text tools but never was fully able to use them fluidly and without frustration. This text was dictated into an LLM already pre-set to use British English spelling and with a prompt to remove only faltering elements and to italicise inconsistencies and repetition and then offer bulleted structural advice. The transcript I then cut and pasted into a Word document and I am now working through it and editing it using old school two finger typing (nb: The bulleted lists in the follow up post are AI generated from the transcript which started to feel a bit long but I still wanted to include. My edits to them are minimal). I have spoken to colleagues and students who have found impressive ways to support their work, thinking and learning. It can explain something differently, provide opportunities for practice, help somebody organise what feels like an overwhelming task or lower barriers that have previously prevented participation. This becomes especially important when considering disabled students, given evidence that some students with dyslexia, ADHD and related conditions describe AI as more effective cognitive support than the formal adjustments they have experienced though even in that domain there are starkly contrasting experiences and opinions.  I think this is why I find blanket statements such as ‘AI should support your thinking, not do your thinking for you’ an appealing provocation or starting point but ultimately a little insufficient.

So what did colleagues think? We presented more than 60 colleagues with a series of scenarios and asked them to place each on a five-point continuum. A score of 1 represented complete agreement that the example constituted appropriate or balanced AI use, while 5 represented complete agreement that it constituted excessive use or over-reliance. The important point here is that I am less interested in treating the resulting numbers as precise measurements of how ‘good’ or ‘bad’ a particular practice is than in what their position on the scale tells us about congruence and conflict. Where responses cluster towards 1, there is relatively strong agreement that the practice is appropriate or balanced. Where they cluster towards 5, there is relatively strong agreement that AI is doing too much. It is the movement towards the middle, around 3, that becomes particularly interesting, because this is where we begin to see conflicting perspectives about what constitutes appropriate use.

Colleagues were strongly inclined to regard students generating AI-based quizzes or flashcards from lecture content as appropriate or balanced use, and similarly comfortable with students asking AI to re-explain difficult concepts using everyday examples. There was also strong agreement around a student drafting an essay themselves and then using AI to identify stylistic or grammatical issues.  There seems to be a reasonably coherent principle sitting underneath these responses. In each case AI is doing something useful for the student, sometimes something quite substantial, but it is not obviously taking over the intellectual work that we want the student to undertake. There is assumed grounding as a direct challenge to hallucination potential in each instance.

Planning is also an interesting area. My logic says (as someone with history background) that the planning is where the thinking happens even though many of the policies I have seen say something on the lines of ‘planning ok; writing support not’. Our  colleagues were broadly comfortable with students uploading deadlines and resources to an AI planning tool which then creates a week-by-week schedule, and were reasonably comfortable with AI helping a student define thematic clusters or task sequences at the beginning of a complex project. What they were much less comfortable with was a student asking AI to generate the research plan and then simply following it to the letter.  That latter point is a nuance I think we may have neglected a little and is worth discussion with students engaging in long form writing activities.

Where AI ‘helps me organise, explore or structure my thinking’, colleagues seem broadly comfortable. Where I hand over decisions about what I should think, investigate or do and simply follow what the machine tells me that’s obviously over-reliance. What, though, about a student who uses multiple prompts across multiple tools to assemble and edit a summative essay? Or somebody who asks AI for the first few steps of a multi-stage calculation and then completes the rest themselves? Or a student who uses an AI agent to create the base code for an interactive website and subsequently develops that code to produce a better and more accessible final product? These examples produced responses much closer to the middle of the continuum, which means that colleagues were not speaking with anything like the same degree of unanimity. I have discussed similar scenarios a lot over the last few years and I think it is fair to say the tendency is towards a more nuanced understanding and acceptance of complex workflows even though many policies in the wider academic sphere ask for a record of ‘prompts and outputs’ which is unworkable and, in my view, fails to understand such complexity. The continued disagreement shows that blanket policies are likely to remain unworkable even though may academic colleagues will tell you they just want to know what’s allowed and what’s not. My conclusion is that that is actually an impossible ask when you get to this level of granularity.

While I accept we do need a lot more in the way of foundations I don’t think we can ever happily say ‘here is the institutional guidance telling everybody which side of the line each example belongs on’. The disagreement may be telling us something much more fundamental about the inadequacy of trying to judge AI use without knowing the educational context in which it occurs. If I am trying to establish whether a student can independently identify and execute every stage of a multi-stage calculation, asking AI to get them started may undermine precisely the capability I am trying to assess. If, however, the calculation is merely one component of a much more complex professional problem and the important learning concerns what the student subsequently does with it, I might reach a very different conclusion. This subtlety can really only be determined by the disciplinary experts; context is everything but also problematic if those same staff lack confidence or core understanding of capabilities of even basic tools, let alone frontier models.

So perhaps one of the most useful conclusions from this exercise is that we are sometimes asking the wrong question. ‘Is this an appropriate use of AI?’ is probably not enough. ‘Is this an appropriate use of AI for this student, undertaking this activity, for this purpose, at this point in their learning? It’s better but a lot more fiddly. I STILL think we’ll talk a lot less about LLMs in a few years and reach a broad level of common understanding  and normalisation of GenAI tools (agents looming much larger on the threat scale) but we have to get there first. And before that we’ll likely have to navigate regulatory panic of some sort and a consolidation of ‘useful’ AI in the hands of fewer players (pretty safe prediction I think based on reading the news nothing more). We can start with where there is congruence though.  AI use that helps students organise, practise, understand, brainstorm or refine their own work is ok in a lot of contexts (I know it’s not without controversy but if we define specific parameters with, in particular, assessment validity firmly in mind) we might get somewhere that looks consistent.  Nothing is automatically appropriate in every context, but they seem to provide useful territory from which to begin developing shared principles.

Agency is key too as there’s that important distinction between interrogating with AI and obeying it, between using an output as a starting point and treating it as an answer, and between retaining responsibility for intellectual decisions and quietly handing those decisions over. Also, the amount of AI being used may actually be less important than what is being outsourced. A student might spend half an hour in quite an elaborate conversation with AI which causes them to question assumptions, consider alternatives and substantially improve their understanding. Another student might type a single prompt and receive exactly the intellectual work that the assessment was designed to elicit from them. Counting prompts, words or percentages of ‘AI contribution’ therefore tells us little. What our colleagues said suggests we should pay particular attention to the examples sitting towards the middle of the continuum, because these expose the assumptions we are making about learning. Rather than trying immediately to eliminate those disagreements through regulation, we should probably be discussing why reasonable colleagues see the same use differently.

This exercise stimulated a lot of discussion and I think that is the best starting point for everyone. And through the noise we see emergent agreement in some places (perhaps more than I might have assumed). But that agreement becomes less secure when we move from broad principles to the messy reality of actual student activities and assessments but even defining those scenarios and helping colleagues realise that there can be legitimacy in position yourself at different places… according to multiple variables in the context. Rather than seeing that as a problem to be eliminated through ever more detailed rules, I take it as an invitation to make our assumptions about curriculum content, teaching and learning more visible. What should institutions, course teams and individual lecturers actually do with all this? I’ll share a few conclusions/ recommendations in a follow up post.

Summer of not love

I just had a really nice email from a high school student in South Africa saying that he’d read one of my blog posts and it had inspired him in his own school-based research. He was chancing his arm a bit and wondering whether he could interview me. But, he appears to be someone rather like I was when I was his age, and left it to the last minute and wanted to interview me today! I offered to speak to him on Monday, but that was a little bit too late for him. Nevertheless, what it did was make me think: my God, I haven’t written anything on my blog for months. So I had a look and the last entry was in May. How is that even possible? I know we had a hot summer but it feels now like it lasted about 5 minutes.

Anyway, it made me think, firstly, why am I not writing on my blog? And on that question I refer you to my calendar, my grey hair and the anguished look behind my eyes. But it also made me think about where everything is going. Because even though I don’t really have that much time to get my thoughts onto paper, doing so actually really helps me. It helps me cohere ideas and learn for myself exactly what it is that I’m thinking, particularly when it’s messy and messy is a word that captures not only the World but a lot of HE right now.

I might not be writing much (beyond completing a couple of chapters and doing a few reviews; nearly forgot about those) but I  do nevertheless still read a lot of controversial posts, especially on LinkedIn where expatriates of Twitter now seem to have gone to argue and share pictures of their dinners (or is that just my algorithm?). It’s especially grim in the AI and education space. The digs and the jibes from all ‘sides’, the over exhuberant ‘influencer’ language effusiveness from others. The horror at our imminent destruction. The bafflement that we’re not taking it seriously enough, or not using it well enough. Everybody’s going to be stupid within the next five years! No, everybody will be able to do super-intelligent things within the next five years with all their free time! No, actually AI will kill us all!  All of this kind of stuff. Actually it gets to me quite a bit.  Some of it is certainly ill-informed. But even when it isn’t ill-informed, it is very often incredibly bile-filled and increasingly acrimonious. Like a mirror to the polarised political car crash unfolding across the world we just don’t seem to have preserved or value reasoned debate.

I say this all the time but we have to be a little bit more nuanced, particularly from the perspective of those of us working in academia. This is unlike any phenomenon that has come before. And I increasingly find myself wanting to find a language for talking about it. A language of compromise, perhaps. A language of mutual respect. A language that enables us to take some of the heat out of the vitriolic, the dismissive, the scornful, the abusive. The shouting prevents people from allowing themselves the space to ask questions that they worry might be perceived as stupid. It makes people pick a side. And, while here I am picking the fence/ shelf/ tightrope (select preferred metaphor), I’m increasingly convinced that picking sides is one of the least useful things we can do.

A lot and not very much has happened since I last posted. One thing I think has genuinely shifted is that the big tech companies have recognised the necessity of, at least superficially, challenging the very real threat of overdependence, cognitive dependency or excessive cognitive offloading. And cognitive offloading is an interesting example of precisely the problem I’m talking about. As many people with superior intellects to mine have pointed out, we have been cognitively offloading in all sorts of ways, and with all sorts of tools, since the dawn of civilisation. Whether it was chipping symbols into rocks to keep tabs on stock or to confirm contracts, all the way through to everything we have done in the digital era and continue to do now. Cognitive offloading isn’t inherently a bad thing; it’s a thing humans do. It is the way in which you do it, when you do it and the tools with which you do it that are the issue. And even that relatively simple discrimination, that simple distinction, seems sometimes to have been lost on people who are very, very intelligent themselves and very well educated and should know bloomin’ better.

Yes, the big tech companies are obviously doing their best to muscle in on the learning space. There has been a noticeable reframing of AI away from the idea of the glorified search engine, which it never was anyway, and the quick-answer machine, towards the idea of the learning machine. Something that can supposedly support cognitive development rather than replace it. Whether that is actually having much impact on student behaviour is another matter.I suspect it is having limited, perhaps minimal, impact so far. All the evidence I see seems to suggest that sophisticated workflows are still very much in the minority. Relatively simplistic uses and applications of free versions of tools remain the norm. And that, I think, is where some of the work now lies. Can we properly help students understand where the affordances lie and where the issues lie?

We have some brilliant colleagues here at UEL doing fantastic work around defining, or helping define for students, what it means to be an AI critic. And the more we can do of that, the better. But we also increasingly need to acknowledge that there are a lot of students who are either vocally taking a stand against AI or, if not exactly taking a stand, are reluctant in their compliance when asked to use AI tools. And often quite cynical and sneery about it.

My own daughter is 15 and talks about AI slop, AI garbage, AI propaganda and AI misinformation. Its rep is poor amongst the youth, I think it is fair to say. That matters because while all of this is going on, AI itself is changing. Big tech is steering towards learning utilities framings – from beyond that idea of a quick-answer machine towards something that can supposedly help develop synaptic firing and the cognitive development. At the same time, we have the increasingly agentic functions of new tools, the vibe-coding potential, the creating-things potential and everything else that is emerging.

It’s  particularly complex for those people who have still refused to really engage with or touch these technologies in any meaningful way, or who abandoned them early on because they thought, ‘Yeah, this isn’t for me’, or ‘this is killing the planet’, or for any of the other perfectly legitimate reasons people have for resisting. Meanwhile, people in roles like mine are trying to work out how we can best get people up to speed, whether they be staff or students, whilst continuing to acknowledge the very many valid reasons why people might push back, resist or reject. And, increasingly, I think all of this has to connect to the bigger questions about higher education. It’s still all about:

What is education for?

What is assessment for?

What is it that we are actually trying to judge?

What is it that we need to teach?

What is it that we want to evaluate?

And how well do our current structures, pedagogies and assessment designs actually enable us to do those things? Which is perhaps the funny thing about looking back at what I was writing in May. The technology has moved on an the tools are faster and more capable. The rhetoric has become louder and with that the polarisation is greater. The arguments about dependency, cognition, authenticity, assessment and human capability have become more urgent. But the questions underneath it all haven’t really changed. What are we trying to achieve through education, and which bits of that should technology help us with?

And perhaps, somewhere between the evangelists telling us that everything has changed and the uber sceptics telling us that everything is being destroyed, there is still some space for the rest of us to work that out, without shouting at each other like hecklers and snarly, stick wielding pensioners at a reform rally…. sorry, conference.