I have a fairly long experience of both teaching and grading, even though I stopped a couple of years ago. After a couple of years of trial-and-error, I think I found the "right" way:
* Alternate between theory and practice in the course of a single 2-hours slot. Theory is necessary, but the attention span of people is very limited.
* Grade either a homework assignment in the form of a real project with specifications. You give the assignment half-semester, so that students who start early can ask questions. Otherwise, grade a in-session exam, with every resource available, including the course and the whole Internet. However, you don't assess the knowledge, but if the student is able to apply their knowledge to the different tasks at hand.
These approaches are now completely moot in the age of AI.
I was surprised the first time a colleague asked me to do an oral exam. I thought it was a burden on people with poor social interaction skills. It went surprisingly well: you could see in a couple of minutes if the student had integrated the concepts of the course, even if they were shy. Now I wonder if it's the way to go. There's one caveat, though: it doesn't scale.
I was thinking that one of the great ironies of artificial intelligence is that it will make good teaching even more labour intensive. Those kids that can afford the one to one tuition required to push past our natural inclination to be lazy will learn things, while everyone else will succumb to outsourcing all their thinking.
There’s no way I’d have learnt everything I know now about software in today’s environment. So much of software development - debugging techniques, architectural decisions, structuring data etc is learned through trial and error. As the models get better is becoming easier and easier to just not look too closely at their output. I suspect it’s human nature.
Oral exams don't scale at the beginning with hundreds of students per course, agreed. So make them in-person written exams, German universities managed that just fine a few decades ago. Oral exams were reserved to advance from Grundstudium (basic) to Hauptstudium (main).
In my very limited experience (I have German major degrees in Informatik and Philosophy, I was a TA for Logic in Philosophy) the more advanced the topic the thinner the attendance anyway.
The country of Argentina, despite its crippling debt, is able to provide free university education where the majority of classes for most majors are graded with an oral exam.
They've shown that it is entirely possible to scale a system like that, as long as the society sincerely values the role of an educator.
The oral exams scale better than one might expect, at least if you're not just doing multiple choice exams and you have to actually grade the exam. The time spent on grading the exam could have also been spent on an oral exam... For bachelor level it's usually 15-20min (usually quicker for the well prepared students), and for master courses it is 30min.
Usually you know after a few minutes if it's going to be a fail, and then otherwise you only need to figure out where on the passing scale the student ends up.
Maybe the Germans were right? One in-person test, preferably an interview, at the end of the semester should be enough to motivate the smarter part of the bunch to make sure they understand the issues at hand. The rest is chaff anyway. Don't optimize for chaff.
I had that a variation of that experience. This was mostly in the pre-AI days.
Homework assignments were mandatory but only counted towards exam admission not the final grade.
Much better for motivation.
When I was a student, nobody cared if you did your home work. The teacher would go through the home work in the next class, and if you didn't do it, you would just lose the opportunity.
Then there was the written exam at the end of each course, which constituted 100% of the grade. The only reason that got abolished was because it made too many students fail, and the board hates that. So that became a race to the bottom, and apparently, we can still go lower.
It's reaching levels where most uni diplomas don't matter any longer. The only thing that will then matter is your network. That's the end of social mobility.
This makes me think of exams — I always did well on open-book exams, but I didn't do well on exams where I couldn't look at the book. Whether it's because of ADHD or just a lack of memorization ability.
Similarly, I did well in interviews where I could use the internet, but I didn't do well in interviews where I couldn't use the internet.
Everyone has a testing format they're good at. I'm curious about how this kind of thing is evaluated, and whether there are any papers related to this.
I now shifted to in-person interactions with a TA. After every assignment, each student needs to schedule a 15-minute meeting with a TA to answer a couple of questions in a live conversation
Isn't the answer the 'flipped classroom' - students are assigned reading/studying to do before a class and the class becomes more interactive, answering questions/discussing topics/solving problems etc depending on subject. Of course, this is more expensive than a hall with 400 people listening to a lecture with a couple of tests/essays.
This only works if you either have very intrinsically motivated students or you verify that really everyone has actually read the assigned materials before class. Otherwise, some people will turn up unprepared, you will have to explain stuff they should have gotten from the reading assignment, other students will be annoyed, students will stop preparing for class, vicious cycle.
And then you fail them. I suspect an institute genuinely empowered to do so, where the expectations were declared up front and followed through on consistently, would fare well, after the dust settled.
As far as I can tell (based on university and high school level instruction forums) institutions increasingly are not backing the teachers that want to let kids that have earned that F fail the students.
I've heard of places where the minimum points for an assignment is 50% even if you never turn it in which is insanity.
I've been in uni classes which used this concept (about 10 years ago though). It also requires that content is given in the right size and complexity each week, otherwise there will also be people coming to class who did do their homework but didn't understand it properly and just like the people who didn't read it need full explanations.
I also felt having multiple explanations (in a regular class where you could also read the book in advance) from regular lecturers helped me in gaining understanding because having something explained in multiple ways made sure there would be at least one way, or a combination, which made it 'click'.
I always liked the regular lectures where you would be required to do some pre class reading and a smaller set of exercises the most, as you would be engaged in the class itself and would be getting repetition with different explanations.
We still have developers because we can't outsource responsibility to AI. If shit breaks, we still need someone to fix it. Maybe you think that can also be automated. I have a counter-example. AI will rarely use the existing state when solving a bug. It will tack on it's own event queue, it will make up it's own messaging system, it is a cancerous growth. And this observation is based on both opus5.5 and gpt5.6-astra. The results are way past magical at this point, but you still have to weigh in pretty heavily for orderly code growth where solutions are succint, focused and preserve the existing structure. They are still noobs at fixing/adding stuff.
I am teaching some math courses, and I see how LLMs disrupted all standard approaches, and I don't know what to do.
I rely on written exams that are open book, but forbid any use of computers and smartphones in class.
The university insists that homework can't be optional, but it lost its meaning. I've tried to explain to my students that it's in their interest to think about home assignments on their own, but the majority will obviously use gen AI, and it's such a waste of time to check and grade LLM output.
I just had an experience where all students were given "sample problems" to try at home and prepare for the written test. Many of them just dumped the document into an LLM, asked it to produce some "exam guide", and showed up with that thing printed out, asking me during the test to explain what LLM output meant.
It seems like the university also has many people pushing for AI use for everything, but I teach basic stuff where the goal is to make students think on their own and digest some fundamental ideas, LLMs can produce perfect solutions, but relying on them is pointless.
UK (and the wider Anglo-American system) is different from the continent, where in many countries the MA is the basic degree -- a mere BA will get you nowhere -- and this easily takes 5-6 years to achieve.
There is indeed no point in grading home assignments anymore, but why not randomly ask a student at the start of each class to present their solution on the board?
And do a final exam on oen and paper and nothing else, and an exam that relies more on thinking than rote memorization of formulas (or print the relevant parts of the course on the exam document itself).
there probably is no time for 100 students to each do that in a university mathematics lecture, but there are probably ways to make sure students are still learning by restructuring the university into mandatory recitation and labs or something
I'm taking a year-long sabbatical to study a language at university right now and they're changing their assignment processes to curtail AI use. We have two versions of written assignments - the first to be written by hand in class with just a Swedish-Swedish word book available, and the second version to be typed at home implementing the professor's corrections from v1.
At least in this course students seem to be taking their strict 'no AI' policy quite seriously. Then again this is also the kind of course you do when you want to actually learn the material. It provides a qualification for further higher education in Swedish, so if we don't legitimately learn the stuff there's just no point as we won't manage in future courses.
Recently it has been reported that many Chinese students in Moscow are no longer learning Russian during their years at the university (to the dismay of Russian officials who hoped this would be a form of soft power): even for in-person lectures their Bluetooth earbuds will just translate the lecturer's Russian into Chinese in real time. So, perhaps people around the world will start feeling that, thanks to technology, exams testing language skills are a mere formality and can be safely cheated on.
Just don't make it open note then? Or you provide the only notes or cheat sheet available to them at the start of each test. That's what some math tests I've had had done before and I believe the SAT also does the same thing with formulas in the back of the test to reference.
I've always wanted to find out if this would work:
If I knew most of my students were using LLMs I would encourage the rest of them to use them too. I would then change the marking criteria such that if they get a single point wrong that's 20% off the entire grade. 5 mistakes gets you a big fat zero.
LLMs are great but they make mistakes. At this point your students would have had to spend so long checking, re-checking and triple checking the LLM's output that they will have accidentally learn what they need to. They will likely need to cross-reference multiple LLMs and at least have read their output which is likely an upgrade on today.
Or some variation of the above, I'd be interested to hear your thoughts.
This sounds like it would come at the cost of the honest students who don't want to use LLM's and train their own understanding.
Reviewing is just not the same as working through a problem yourself. You often can take shortcuts when checking answers for correctness, which generation (with mind or LLM) of the solution can't take. This isn't limited to Math but equally true for programming and many other skills.
I'd push back slightly on this: I spend more of my time reading / reviewing code than writing it yet I still find reading / reviewing code much more mentally demanding. I don't think I'm alone in that.
The people that do their own work would be at a huge advantage as the LLM(s) can work as reviewers instead of doing the work so there is more chance of catching a mistake before you lose 20%.
On top of that there will be some issues where the LLM is just plain wrong. Being able to work out when that is the case is a hugely useful skill in 2026. It is also something the people who blindly feed their work into an LLM and print its answer without reading it will be unable to discover. Much to their own detriment.
I'm not trying to suggest that reviewing (code) can't feel (more) mentally demanding. The thing with at least reviews is that if you have worked through the problem yourself, you will also hit certain walls where stuff didn't work or wasn't as beautiful as you wanted (ie trial and error), and you have to activate your brain to think of other solutions, ie really think a problem through multiple times with active feedback. As a reviewer you don't get the active feedback usually.
Now this i think applies in lesser degree to the comment I was replying to. With maths, and some comp sci problems, you often can check an answer much faster than actually solving it, eg by using simple substitution for simple math problems. That is a useful skill, but a different one from solving the problem yourself. It employs different types of skills.
I think you might be letting perfect be the enemy of good here. Right now some students are using an LLM to do all of their work. I don't suggest for a moment my method is perfect but I do claim it could be better than the status quo (or at least warrants further investigation in my opinion).
What you would grade with that approach is how well your students can operate a LLM as well as the failure rate of whatever LLM infrastructure the individual student is using. This is usually not the goal of your educational format and hence not what you would grade for - apart from some (meta-)skill courses offered by uni libraries and the like. Those are mostly ungraded, though.
Anyone that knows the subject matter well enough will spot the LLMs mistakes and correct them. If you don't know the subject matter well enough to spot a mistake then you don't know the subject matter well enough.
Put a mixture of questions in that are impossible to answer but don't tell them, and as usual give them marks for explaining their working. Tell them you've devised the questions so some are "AI resistant" but don't tell them how. They'll never know if the AI has tripped up and hallucinated or given them the right answer. Only those that really understand what's going on will set out some working that shows they were on the right route and then got stuck where they were supposed to.
That will either lead to students becoming frustrated after spending inordinate amounts of time on the unsolvable problem or students giving up too early on the solvable problems.
That sounds pretty frustrating. I wonder if it wouldn't be a good approach to, say:
1. Make the exam the only thing that determines the grade (or technically, proves that the student learned enough);
2. Give out test assignments students can do, but don't have to;
3. Offer students for the professor to grade their work if they want to get some feedback on how they're doing.
Would that lead to less time wasted by tutors grading AI generated answers? From where I stand, students should be allowed to prepare for exams any way they like, with or without their professor's help. If they managed to learn, they pass.
When I was in engineering school ~20 years ago, some lecturers openly said that they gave the assignments just enough credit to be worth doing. Something like 10-15% of your final grade would be from assignments, the rest from the exam. So we didn't blow off the assignments to go drinking.
Of course, this a lot easier for some subjects than others. Subjects like Film Studies relied on exams much less, and assignments much more.
According to the article, #1 leads to "failure and drop-out rates of 50 to 80%". #2 and #3 might work for organized and motivated students, but graded assignments from other classes will be prioritized by most.
such a waste of time to check and grade LLM output
Small nitpick: you teach math, why don't you use an automated grader?
I agree homework should be optional and your college is wrong. Even without AI, for some students it's easy to cheat and maybe for some students it's a waste of time. Maybe you can assign it a very low score, like 5% of the final grade, and/or award the grade based on attempting rather than getting the right answer.
Many of them just dumped the document into an LLM, asked it to produce some "exam guide", and showed up with that thing printed out, asking me during the test to explain what LLM output meant.
That student presumably performed poorly on the test. Your approach is working as intended here, isn't it?
Also, is the AI-generated exam guide all that different to a human-written guide they could have downloaded a decade ago?
It must be tiresome to be a student these days… with each teacher having different pet theories about AI use and sometimes trying to “trick” students who use AI, or implement some draconian rules about the way homework or exams must be done. It’s tiresome for you, but just imagine getting varying lectures on AI use as you go from class to class as a student. Must be hell
It's funny how people talk about "teaching" getting more difficult with LLMs.
Clearly teaching got a lot better, since we now have a new teaching tool at our use, the LLMs. It's evaluating that got harder, because students can use the same tool to cheat in traditional forms of evaluation.
If "teaching the students" means the same as "evaluating the students" to you, maybe you shouldn't teach in the first place.
It’s not just evaluating. Students used to learn skills while doing an assignment. LLMing through assignments teaches copy-paste at best.
It’s like copy-paste Wikipedia presentations 2 decades ago. People used to learn things doing research for a project. Then suddenly it became copying off Wikipedia and people not even pre-reading what they copied into PPT.
What you’re calling evaluating is a part of the learning process for the student, not just about confirming that the student knows the thing. LLMs are actively undermining learning in this way. There was useful friction which has been removed.
Teaching and evaluation are of course not the same thing, but evaluation is an important part of teaching.
There is a mismatch between students’ overall desire to learn things and the moment-to-moment experience that learning is hard.
Even devoid of all the economic and social pressures that grades impose students would compromise their learning experience.
Ever regretted looking up the solution to a puzzle in a video game or sneaked a peak at the crossword puzzle solutions?
One of the hardest part of teaching is keeping students from self-sabotage.
Evaluation helps with that.
Beyond that evaluation is of course also useful just as feedback.
Teaching is imo mostly about techniques that battle laziness. Teaching intrinsically motivated intelligent students has always been trivial. All progress in education is about scaling it to work for people who only reluctantly engage with the material.
That outlook is way too bleak. As a person who has taught at universities for fifteen years and has run lots of open educational formats and infrastructure with other public institutions and citizen groups, my experience is the opposite. If anything, the desire to learn and curiosity are among the strongest common traits humans possess. Grades are just a tool and a proxy that most people working in formalized educational settings have to work with at specific points.
Seems the core problem is that many/most students aren't there to learn the material, just to get grades. The class is a complex mechanism to force them to learn material against their will in a validation arms race. The teacher's job becomes something like a military commander. Seems depressing.
This is because we shield children too much from reality. If they had a fire under their ass that made them realize if they don't study well, they will forever be locked into flipping burgers or being homeless, most of them would move their assess.
Don't do anything productive for the society but still get to live a decent life on benefits, job security etc combined with addictive social media have ruined most kids will to study.
And the schools are even worse. When I was in school, teachers encouraged learning things; whether for school or outside, with teacher or self help. Everything was fair game and supported. Debate was normal.
Nowadays the teachers in schools are like a cult. Kids are tuned to reject anything from outside; anything that contradicts teacher is a no go zone. Anything that's "too advanced" compared to what teacher is doing is forbidden. Kids are scared of learning by themselves because the teachers punish it. What do you think these kids will do when they grow up?
The least paragraph... has always been the case. My parents told me stories about teachers denying facts not to lose authority. I distinctly recall my teacher pretending that peripheral vision doesn't exist not to lose authority.
If they had a fire under their ass that made them realize if they don't study well, they will forever be locked into flipping burgers or being homeless, most of them would move their assess.
That strikes me as backward. The higher the expected monetary value of good grades, the more students will prioritise grades over deep learning.
20 years ago my department head was saying homework was pretty pointless because everyone was copying each other. They assigned a token percentage for homework so at least the students were submitting homework.
School homework is one of those things which, like the OnlyFans economy, I hope AI utterly destroys, despite my general detestation for the corrosive effects of AI on society.
I teach mostly smaller courses but my approach to grading has been the same for years: A combination of a semester-long group project (ending in a paper, a prototype and a group presentation) with individual oral exams about aspects of the project and of the lecture. The lecture is flanked by optional TA sessions as well as access to labs which, for most hours of the day, have some experienced/teaching people being around (this results in a lot of ad-hoc learning situations and a general diffusion of knowledge and practices).
This gives me/us a lot of angles to understand what progress has been made throughout the semester as well as to consider and react to individual differences between students. In my book, it’s also resistant to faking it convincingly throughout all modalities. However, scaling that approach to big entry level courses would require more money for teaching than any research group in Europe I’ve ever seen has available.
I dunno, if I had to teach students now I think I’d just force them to stay in the lecture hall without internet for the hour I need them to work on their assignments that week. I just can’t see any other reasonable way to do this. Everyone leaves their phones at the door in some designated locker or something.
I was at EduLearn (a major edtech conference) this year and talked to many professors exactly about this. A lot of them are pulling work back into the classroom, but keep running into the same problems:
- some skills only develop through the actual writing (similar to how you can't learn to code just by watching YouTube videos)
- in-class assignments are not always scalable (and students end up using AI anyway)
- students aren't motivated enough to participate actively
I think looking at the writing process is a more realistic approach, so I built Turingo (https://www.turingo.net). It replays how a Google Doc was written, highlights unusual edits (large pastes) and tracks how that text changed afterwards. Surprisingly, several professors reported that students often underestimate how much LLM-generated text is left in their final drafts, so just seeing the replay changes the conversation.
It's not meant to be a verdict machine like the existing AI detectors (another big problem in education), more of a starting point for the kind of conversations the author's TAs are having. In fact, we're also beta testing a feature that suggests a few questions for the professor to ask based on the replay.
And yes, it's not perfect and autotypers do exist, but they're dumber than you'd expect and nowhere close to mimicking a real human (I hope to address that soon).
It sad that it has come to this, but your approach is correct - it is the act of writing (code or prose) that is important, not the finished output itself.
Interestingly, this applies not just for learning, but also to a great extent in production/work contexts too.
I cannot shake the feeling that we are chasing the wrong goose.
The end goal of education is to help you understand how things work, with problem-solving being a means of developing and applying that understanding. LLMs shortcut the former and turn the latter into a pointless arms race that we cannot win.
I loved the example of increasing the complexity of learning projects, but I still feel like one question is missing: are people really trying to learn the same thing we were trying to teach them in the first place?
In my experience a significant proportion of "evidence-based best pedagogy practices" is hogwash. I'm pretty anti-AI, but going against education-school dogma isn't a reason to avoid AI.
Just grade the process, not the result. It's literally that simple.
AI can talk the talk, but it can't walk the walk.
A simple "explain why this is the best solution" will make it shit the bed 10 out 10 times.
Thanks for sharing. Interesting to see how teachers are adapting to LLMs. I agree that trying to ban AI use is futile.
I have a fairly long experience of both teaching and grading, even though I stopped a couple of years ago. After a couple of years of trial-and-error, I think I found the "right" way:
* Alternate between theory and practice in the course of a single 2-hours slot. Theory is necessary, but the attention span of people is very limited. * Grade either a homework assignment in the form of a real project with specifications. You give the assignment half-semester, so that students who start early can ask questions. Otherwise, grade a in-session exam, with every resource available, including the course and the whole Internet. However, you don't assess the knowledge, but if the student is able to apply their knowledge to the different tasks at hand.
These approaches are now completely moot in the age of AI.
I was surprised the first time a colleague asked me to do an oral exam. I thought it was a burden on people with poor social interaction skills. It went surprisingly well: you could see in a couple of minutes if the student had integrated the concepts of the course, even if they were shy. Now I wonder if it's the way to go. There's one caveat, though: it doesn't scale.
I was thinking that one of the great ironies of artificial intelligence is that it will make good teaching even more labour intensive. Those kids that can afford the one to one tuition required to push past our natural inclination to be lazy will learn things, while everyone else will succumb to outsourcing all their thinking.
There’s no way I’d have learnt everything I know now about software in today’s environment. So much of software development - debugging techniques, architectural decisions, structuring data etc is learned through trial and error. As the models get better is becoming easier and easier to just not look too closely at their output. I suspect it’s human nature.
Oral exams don't scale at the beginning with hundreds of students per course, agreed. So make them in-person written exams, German universities managed that just fine a few decades ago. Oral exams were reserved to advance from Grundstudium (basic) to Hauptstudium (main).
In my very limited experience (I have German major degrees in Informatik and Philosophy, I was a TA for Logic in Philosophy) the more advanced the topic the thinner the attendance anyway.
They do scale if the examiner is an LLM.
The country of Argentina, despite its crippling debt, is able to provide free university education where the majority of classes for most majors are graded with an oral exam.
They've shown that it is entirely possible to scale a system like that, as long as the society sincerely values the role of an educator.
It scales just fine. Use an AI to do oral exam (recorded). Flag the problematic orals and reevaluate the scripts/recorded session.
Finally, if a student feels their result is unjustified, reevaluate the recorded session.
Use more AI to solve the problem of AI?
The oral exams scale better than one might expect, at least if you're not just doing multiple choice exams and you have to actually grade the exam. The time spent on grading the exam could have also been spent on an oral exam... For bachelor level it's usually 15-20min (usually quicker for the well prepared students), and for master courses it is 30min.
Usually you know after a few minutes if it's going to be a fail, and then otherwise you only need to figure out where on the passing scale the student ends up.
Maybe the Germans were right? One in-person test, preferably an interview, at the end of the semester should be enough to motivate the smarter part of the bunch to make sure they understand the issues at hand. The rest is chaff anyway. Don't optimize for chaff.
I had that a variation of that experience. This was mostly in the pre-AI days. Homework assignments were mandatory but only counted towards exam admission not the final grade. Much better for motivation.
When I was a student, nobody cared if you did your home work. The teacher would go through the home work in the next class, and if you didn't do it, you would just lose the opportunity.
Then there was the written exam at the end of each course, which constituted 100% of the grade. The only reason that got abolished was because it made too many students fail, and the board hates that. So that became a race to the bottom, and apparently, we can still go lower.
It's reaching levels where most uni diplomas don't matter any longer. The only thing that will then matter is your network. That's the end of social mobility.
One test is too stressful and may be affected by bad luck. I'd recommend many small quizzes. You can use LLMs to grade.
It's terrifying that some people actually do think this. Not you obviously, but I had enough conversations with other people.
on the off chance you’re being sincere: the article mentions that they allow free retakes of the oral exam
I agree.
This makes me think of exams — I always did well on open-book exams, but I didn't do well on exams where I couldn't look at the book. Whether it's because of ADHD or just a lack of memorization ability.
Similarly, I did well in interviews where I could use the internet, but I didn't do well in interviews where I couldn't use the internet.
Everyone has a testing format they're good at. I'm curious about how this kind of thing is evaluated, and whether there are any papers related to this.
Isn't the answer the 'flipped classroom' - students are assigned reading/studying to do before a class and the class becomes more interactive, answering questions/discussing topics/solving problems etc depending on subject. Of course, this is more expensive than a hall with 400 people listening to a lecture with a couple of tests/essays.
https://fltmag.com/the-flipped-classroom/
https://en.wikipedia.org/wiki/Flipped_classroom
This only works if you either have very intrinsically motivated students or you verify that really everyone has actually read the assigned materials before class. Otherwise, some people will turn up unprepared, you will have to explain stuff they should have gotten from the reading assignment, other students will be annoyed, students will stop preparing for class, vicious cycle.
And then you fail them. I suspect an institute genuinely empowered to do so, where the expectations were declared up front and followed through on consistently, would fare well, after the dust settled.
As far as I can tell (based on university and high school level instruction forums) institutions increasingly are not backing the teachers that want to let kids that have earned that F fail the students.
I've heard of places where the minimum points for an assignment is 50% even if you never turn it in which is insanity.
I've been in uni classes which used this concept (about 10 years ago though). It also requires that content is given in the right size and complexity each week, otherwise there will also be people coming to class who did do their homework but didn't understand it properly and just like the people who didn't read it need full explanations.
I also felt having multiple explanations (in a regular class where you could also read the book in advance) from regular lecturers helped me in gaining understanding because having something explained in multiple ways made sure there would be at least one way, or a combination, which made it 'click'.
I always liked the regular lectures where you would be required to do some pre class reading and a smaller set of exercises the most, as you would be engaged in the class itself and would be getting repetition with different explanations.
We still have developers because we can't outsource responsibility to AI. If shit breaks, we still need someone to fix it. Maybe you think that can also be automated. I have a counter-example. AI will rarely use the existing state when solving a bug. It will tack on it's own event queue, it will make up it's own messaging system, it is a cancerous growth. And this observation is based on both opus5.5 and gpt5.6-astra. The results are way past magical at this point, but you still have to weigh in pretty heavily for orderly code growth where solutions are succint, focused and preserve the existing structure. They are still noobs at fixing/adding stuff.
I am teaching some math courses, and I see how LLMs disrupted all standard approaches, and I don't know what to do.
I rely on written exams that are open book, but forbid any use of computers and smartphones in class.
The university insists that homework can't be optional, but it lost its meaning. I've tried to explain to my students that it's in their interest to think about home assignments on their own, but the majority will obviously use gen AI, and it's such a waste of time to check and grade LLM output.
I just had an experience where all students were given "sample problems" to try at home and prepare for the written test. Many of them just dumped the document into an LLM, asked it to produce some "exam guide", and showed up with that thing printed out, asking me during the test to explain what LLM output meant.
It seems like the university also has many people pushing for AI use for everything, but I teach basic stuff where the goal is to make students think on their own and digest some fundamental ideas, LLMs can produce perfect solutions, but relying on them is pointless.
Perhaps some people do not belong at university? It is just a waste of everyone time, to insist 50% of population studies until 25 years old.
It's like we're putting everybody on crack and then say they're too weak if they become addicted.
If you do not want to learn, and take every shortcut available to avoid it, university might not be the path for you.
25? At least in the UK, most degrees are 3 years long and most people start when they are 18.
UK (and the wider Anglo-American system) is different from the continent, where in many countries the MA is the basic degree -- a mere BA will get you nowhere -- and this easily takes 5-6 years to achieve.
There is indeed no point in grading home assignments anymore, but why not randomly ask a student at the start of each class to present their solution on the board?
And do a final exam on oen and paper and nothing else, and an exam that relies more on thinking than rote memorization of formulas (or print the relevant parts of the course on the exam document itself).
there probably is no time for 100 students to each do that in a university mathematics lecture, but there are probably ways to make sure students are still learning by restructuring the university into mandatory recitation and labs or something
Having two random people picked each time might cause all of them to at least pay attention.
I'm taking a year-long sabbatical to study a language at university right now and they're changing their assignment processes to curtail AI use. We have two versions of written assignments - the first to be written by hand in class with just a Swedish-Swedish word book available, and the second version to be typed at home implementing the professor's corrections from v1.
At least in this course students seem to be taking their strict 'no AI' policy quite seriously. Then again this is also the kind of course you do when you want to actually learn the material. It provides a qualification for further higher education in Swedish, so if we don't legitimately learn the stuff there's just no point as we won't manage in future courses.
Recently it has been reported that many Chinese students in Moscow are no longer learning Russian during their years at the university (to the dismay of Russian officials who hoped this would be a form of soft power): even for in-person lectures their Bluetooth earbuds will just translate the lecturer's Russian into Chinese in real time. So, perhaps people around the world will start feeling that, thanks to technology, exams testing language skills are a mere formality and can be safely cheated on.
Just don't make it open note then? Or you provide the only notes or cheat sheet available to them at the start of each test. That's what some math tests I've had had done before and I believe the SAT also does the same thing with formulas in the back of the test to reference.
I've always wanted to find out if this would work:
If I knew most of my students were using LLMs I would encourage the rest of them to use them too. I would then change the marking criteria such that if they get a single point wrong that's 20% off the entire grade. 5 mistakes gets you a big fat zero.
LLMs are great but they make mistakes. At this point your students would have had to spend so long checking, re-checking and triple checking the LLM's output that they will have accidentally learn what they need to. They will likely need to cross-reference multiple LLMs and at least have read their output which is likely an upgrade on today.
Or some variation of the above, I'd be interested to hear your thoughts.
This sounds like it would come at the cost of the honest students who don't want to use LLM's and train their own understanding.
Reviewing is just not the same as working through a problem yourself. You often can take shortcuts when checking answers for correctness, which generation (with mind or LLM) of the solution can't take. This isn't limited to Math but equally true for programming and many other skills.
I'd push back slightly on this: I spend more of my time reading / reviewing code than writing it yet I still find reading / reviewing code much more mentally demanding. I don't think I'm alone in that.
The people that do their own work would be at a huge advantage as the LLM(s) can work as reviewers instead of doing the work so there is more chance of catching a mistake before you lose 20%.
On top of that there will be some issues where the LLM is just plain wrong. Being able to work out when that is the case is a hugely useful skill in 2026. It is also something the people who blindly feed their work into an LLM and print its answer without reading it will be unable to discover. Much to their own detriment.
I'm not trying to suggest that reviewing (code) can't feel (more) mentally demanding. The thing with at least reviews is that if you have worked through the problem yourself, you will also hit certain walls where stuff didn't work or wasn't as beautiful as you wanted (ie trial and error), and you have to activate your brain to think of other solutions, ie really think a problem through multiple times with active feedback. As a reviewer you don't get the active feedback usually.
Now this i think applies in lesser degree to the comment I was replying to. With maths, and some comp sci problems, you often can check an answer much faster than actually solving it, eg by using simple substitution for simple math problems. That is a useful skill, but a different one from solving the problem yourself. It employs different types of skills.
I think you might be letting perfect be the enemy of good here. Right now some students are using an LLM to do all of their work. I don't suggest for a moment my method is perfect but I do claim it could be better than the status quo (or at least warrants further investigation in my opinion).
What you would grade with that approach is how well your students can operate a LLM as well as the failure rate of whatever LLM infrastructure the individual student is using. This is usually not the goal of your educational format and hence not what you would grade for - apart from some (meta-)skill courses offered by uni libraries and the like. Those are mostly ungraded, though.
Anyone that knows the subject matter well enough will spot the LLMs mistakes and correct them. If you don't know the subject matter well enough to spot a mistake then you don't know the subject matter well enough.
Put a mixture of questions in that are impossible to answer but don't tell them, and as usual give them marks for explaining their working. Tell them you've devised the questions so some are "AI resistant" but don't tell them how. They'll never know if the AI has tripped up and hallucinated or given them the right answer. Only those that really understand what's going on will set out some working that shows they were on the right route and then got stuck where they were supposed to.
That will either lead to students becoming frustrated after spending inordinate amounts of time on the unsolvable problem or students giving up too early on the solvable problems.
That sounds pretty frustrating. I wonder if it wouldn't be a good approach to, say:
1. Make the exam the only thing that determines the grade (or technically, proves that the student learned enough);
2. Give out test assignments students can do, but don't have to;
3. Offer students for the professor to grade their work if they want to get some feedback on how they're doing.
Would that lead to less time wasted by tutors grading AI generated answers? From where I stand, students should be allowed to prepare for exams any way they like, with or without their professor's help. If they managed to learn, they pass.
When I was in engineering school ~20 years ago, some lecturers openly said that they gave the assignments just enough credit to be worth doing. Something like 10-15% of your final grade would be from assignments, the rest from the exam. So we didn't blow off the assignments to go drinking.
Of course, this a lot easier for some subjects than others. Subjects like Film Studies relied on exams much less, and assignments much more.
According to the article, #1 leads to "failure and drop-out rates of 50 to 80%". #2 and #3 might work for organized and motivated students, but graded assignments from other classes will be prioritized by most.
Small nitpick: you teach math, why don't you use an automated grader?
I agree homework should be optional and your college is wrong. Even without AI, for some students it's easy to cheat and maybe for some students it's a waste of time. Maybe you can assign it a very low score, like 5% of the final grade, and/or award the grade based on attempting rather than getting the right answer.
LLM verifies output of another LLM. The circle is complete, and nobody has any doubts that this is just a theatre.
With manual grading, that one student who genuinely tries still has a chance to meaningfully engage.
That student presumably performed poorly on the test. Your approach is working as intended here, isn't it?
Also, is the AI-generated exam guide all that different to a human-written guide they could have downloaded a decade ago?
It must be tiresome to be a student these days… with each teacher having different pet theories about AI use and sometimes trying to “trick” students who use AI, or implement some draconian rules about the way homework or exams must be done. It’s tiresome for you, but just imagine getting varying lectures on AI use as you go from class to class as a student. Must be hell
Thinking too small. The interactive testing needs to be done by AI. Then the transcript is reviewed by the teacher.
It's funny how people talk about "teaching" getting more difficult with LLMs.
Clearly teaching got a lot better, since we now have a new teaching tool at our use, the LLMs. It's evaluating that got harder, because students can use the same tool to cheat in traditional forms of evaluation.
If "teaching the students" means the same as "evaluating the students" to you, maybe you shouldn't teach in the first place.
It’s not just evaluating. Students used to learn skills while doing an assignment. LLMing through assignments teaches copy-paste at best.
It’s like copy-paste Wikipedia presentations 2 decades ago. People used to learn things doing research for a project. Then suddenly it became copying off Wikipedia and people not even pre-reading what they copied into PPT.
What you’re calling evaluating is a part of the learning process for the student, not just about confirming that the student knows the thing. LLMs are actively undermining learning in this way. There was useful friction which has been removed.
Teaching and evaluation are of course not the same thing, but evaluation is an important part of teaching. There is a mismatch between students’ overall desire to learn things and the moment-to-moment experience that learning is hard.
Even devoid of all the economic and social pressures that grades impose students would compromise their learning experience. Ever regretted looking up the solution to a puzzle in a video game or sneaked a peak at the crossword puzzle solutions?
One of the hardest part of teaching is keeping students from self-sabotage. Evaluation helps with that.
Beyond that evaluation is of course also useful just as feedback.
Teaching is imo mostly about techniques that battle laziness. Teaching intrinsically motivated intelligent students has always been trivial. All progress in education is about scaling it to work for people who only reluctantly engage with the material.
Without grading, students don't learn. It's as simple as that.
Pretty strong opinion that overlooks all those students that study or attend courses specifically to learn.
That outlook is way too bleak. As a person who has taught at universities for fifteen years and has run lots of open educational formats and infrastructure with other public institutions and citizen groups, my experience is the opposite. If anything, the desire to learn and curiosity are among the strongest common traits humans possess. Grades are just a tool and a proxy that most people working in formalized educational settings have to work with at specific points.
Seems the core problem is that many/most students aren't there to learn the material, just to get grades. The class is a complex mechanism to force them to learn material against their will in a validation arms race. The teacher's job becomes something like a military commander. Seems depressing.
This is because we shield children too much from reality. If they had a fire under their ass that made them realize if they don't study well, they will forever be locked into flipping burgers or being homeless, most of them would move their assess.
Don't do anything productive for the society but still get to live a decent life on benefits, job security etc combined with addictive social media have ruined most kids will to study.
And the schools are even worse. When I was in school, teachers encouraged learning things; whether for school or outside, with teacher or self help. Everything was fair game and supported. Debate was normal.
Nowadays the teachers in schools are like a cult. Kids are tuned to reject anything from outside; anything that contradicts teacher is a no go zone. Anything that's "too advanced" compared to what teacher is doing is forbidden. Kids are scared of learning by themselves because the teachers punish it. What do you think these kids will do when they grow up?
The least paragraph... has always been the case. My parents told me stories about teachers denying facts not to lose authority. I distinctly recall my teacher pretending that peripheral vision doesn't exist not to lose authority.
That strikes me as backward. The higher the expected monetary value of good grades, the more students will prioritise grades over deep learning.
20 years ago my department head was saying homework was pretty pointless because everyone was copying each other. They assigned a token percentage for homework so at least the students were submitting homework.
The most interesting part here is that AI isn't replacing the learning goals, it's replacing the evidence instructors used to measure them
School homework is one of those things which, like the OnlyFans economy, I hope AI utterly destroys, despite my general detestation for the corrosive effects of AI on society.
You're not learning anything if you don't put in your hours. Maybe you're fine with that but make sure you realize what you're getting into.
I teach mostly smaller courses but my approach to grading has been the same for years: A combination of a semester-long group project (ending in a paper, a prototype and a group presentation) with individual oral exams about aspects of the project and of the lecture. The lecture is flanked by optional TA sessions as well as access to labs which, for most hours of the day, have some experienced/teaching people being around (this results in a lot of ad-hoc learning situations and a general diffusion of knowledge and practices).
This gives me/us a lot of angles to understand what progress has been made throughout the semester as well as to consider and react to individual differences between students. In my book, it’s also resistant to faking it convincingly throughout all modalities. However, scaling that approach to big entry level courses would require more money for teaching than any research group in Europe I’ve ever seen has available.
I dunno, if I had to teach students now I think I’d just force them to stay in the lecture hall without internet for the hour I need them to work on their assignments that week. I just can’t see any other reasonable way to do this. Everyone leaves their phones at the door in some designated locker or something.
I was at EduLearn (a major edtech conference) this year and talked to many professors exactly about this. A lot of them are pulling work back into the classroom, but keep running into the same problems:
- some skills only develop through the actual writing (similar to how you can't learn to code just by watching YouTube videos)
- in-class assignments are not always scalable (and students end up using AI anyway)
- students aren't motivated enough to participate actively
I think looking at the writing process is a more realistic approach, so I built Turingo (https://www.turingo.net). It replays how a Google Doc was written, highlights unusual edits (large pastes) and tracks how that text changed afterwards. Surprisingly, several professors reported that students often underestimate how much LLM-generated text is left in their final drafts, so just seeing the replay changes the conversation.
It's not meant to be a verdict machine like the existing AI detectors (another big problem in education), more of a starting point for the kind of conversations the author's TAs are having. In fact, we're also beta testing a feature that suggests a few questions for the professor to ask based on the replay.
And yes, it's not perfect and autotypers do exist, but they're dumber than you'd expect and nowhere close to mimicking a real human (I hope to address that soon).
It sad that it has come to this, but your approach is correct - it is the act of writing (code or prose) that is important, not the finished output itself.
Interestingly, this applies not just for learning, but also to a great extent in production/work contexts too.
I cannot shake the feeling that we are chasing the wrong goose.
The end goal of education is to help you understand how things work, with problem-solving being a means of developing and applying that understanding. LLMs shortcut the former and turn the latter into a pointless arms race that we cannot win.
I loved the example of increasing the complexity of learning projects, but I still feel like one question is missing: are people really trying to learn the same thing we were trying to teach them in the first place?
very simple. All exams to be paper and pen.
In my experience a significant proportion of "evidence-based best pedagogy practices" is hogwash. I'm pretty anti-AI, but going against education-school dogma isn't a reason to avoid AI.
Why teach someone who doesn't want to learn?
I managed to push AI to make me a math solver: https://nonconfirmed.com/mathsolver-n/