Please stop using that MIT Cognitive Debt paper as evidence of something
or is it cognitive atrophy? wait no, cognitive cost? wait, no
At least once a week, I encounter something like this:
And every single time, the ‘emerging research’ comes back to this MIT Media Lab study from mid last year:
Mostly this stuff gets consumed through blogs, articles, and podcasts that act as mediums translating the research itself. They trust that you trust that the academics understand this more than you, so you don’t have to go the primary source.
Well, I have read the actual paper.
It tells me nothing about the effect of AI use on my cognitive abilities. And it actually worries me to see the path of limited research design to actual national policy discussion and debate play out without any resistance at all. There’s a topic related meta joke in there somewhere but there’s no time.
Anyway, media and fake intellectuals exaggerating a paper’s findings is par for the course. Fine. But here’s what the authors claim their paper shows…
This study focuses on finding out the cognitive cost of using an LLM in the educational context of writing an essay.
….
As the educational impact of LLM use only begins to settle with the general population, in this study we demonstrate the pressing matter of a likely decrease in learning skills based on the results of our study. The use of LLM had a measurable impact on participants, and while the benefits were initially apparent, as we demonstrated over the course of 4 months, the LLM group’s participants performed worse than their counterparts in the Brain-only group at all levels: neural, linguistic, scoring.
…
We hope this study serves as a preliminary guide to understanding the cognitive and practical impacts of AI on learning environments.
Students that use LLMs show a decrease in learning skills. There is a real, measurable ‘cognitive cost’.
How did they reach that conclusion? Well, they start with their research questions, emerging at the end of the introduction:
We attempt to respond to the following questions in our study:
1. Do participants write significantly different essays when using LLMs, search engine and their brain-only?
2. How do participants’ brain activity differ when using LLMs, search or their brain-only?
3. How does using LLM impact participants’ memory?
4. Does LLM usage impact ownership of the essays?
Immediate note: these are the actual questions they thought they could answer with their data. Are any of these the same as ‘Do LLMs negatively impact learning skills?’, or will this always require a level of conjecture to turn this into a policy implication?
So before we even get into the measurement, some things to initially consider about these questions:
Obviously yes they will. That’s a descriptive finding. Is that a good or bad thing? Will these researchers say if it’s good or bad? Will they unsubtly imply it? Even if they don’t, will the media attach some normativity to this?
Bonus Q: Do you even need to run research to answer this question? Or perhaps, do we live in a system where something has to be said by a respected university for it to count as a fact, even though everyone believed it already? Are you maybe seeing how funding works now?
Of course it will differ, but fair play, here they are asking how. The basic guess would be the brain pivots towards planning and executive functions. The higher level guiding the LLM into helping you produce the output. Again, is that good or bad? Can the researchers tell you that?
Incredibly vague, but let’s assume they mean memory of what the student produced. A student who wrote every word themselves will almost certainly remember more. How much more will be interesting too.
This is the first time the word ‘ownership’ comes up so I literally don’t know what it, or the question, means yet.
Point being, there is a lot of affect being smuggled in here. You are being guided to feel a certain way about what you are about to see.
Which, by the way, is the whole point of everything you encounter in a paper before you reach the methodology. If they think you are especially dumb, which they do, they’ll even add an extra section for statements such as:
‘According to Turner and Rainie [21],“81 percent of Americans rely on information from the Internet ‘a lot’ when making important decisions,“ many of which involve learning activities [22].
And 19% of Americans are essentially/literally dead or experiencing power outages at the moment?
The fact that a statement like ‘Americans rely on the internet for information’ requires an appeal to authority goes a long to show that conceptions of what ‘good’ looks like in Western education has wayyy more to do with conformity than it does brain power.
Speaking of conformity, it is finally time to get into the experiment.
Here is who we are dealing with:
Originally, 60 adults were recruited to participate in our study, but due to scheduling difficulties, 55 completed the experiment in full (attending a minimum of three sessions, defined later). To ensure data distribution, we are here only reporting data from 54 participants (as participants were assigned in three groups, see details below). These 54 participants were between the ages of 18 to 39 years old (age M = 22.9, SD = 1.69) and all recruited from the following 5 universities in greater Boston area: MIT (14F, 5M), Wellesley (18F), Harvard (1N/A, 7M, 2 Non-Binary), Tufts (5M), and Northeastern (2M) (Figure 3). 35 participants reported pursuing undergraduate studies and 14 postgraduate studies. 6 participants either finished their studies with MSc or PhD degrees, and were currently working at the universities as post-docs (2), research scientists (2), software engineers (2) (Figure 2). 32 participants indicated their gender as female, 19 - male, 2 - non-binary and 1 participant preferred not to provide this information.
I’ve left that unedited under the assumption that you would glaze over it. Which is probably what the authors wanted you to do too. I don’t mean that as a personal jab at this research team. It is just important to note that this is how this stuff always is. And it’s not good for anyone.
Here is what you should pick up on:
Originally, 60 adults were recruited to participate in our study, but due to scheduling difficulties, 55 completed the experiment in full (attending a minimum of three sessions, defined later). To ensure data distribution, we are here only reporting data from 54 participants (as participants were assigned in three groups, see details below). These 54 participants were between the ages of 18 to 39 years old (age M = 22.9, SD = 1.69) and all recruited from the following 5 universities in greater Boston area: MIT (14F, 5M), Wellesley (18F), Harvard (1N/A, 7M, 2 Non-Binary), Tufts (5M), and Northeastern (2M) (Figure 3). 35 participants reported pursuing undergraduate studies and 14 postgraduate studies. 6 participants either finished their studies with MSc or PhD degrees, and were currently working at the universities as post-docs (2), research scientists (2), software engineers (2) (Figure 2). 32 participants indicated their gender as female, 19 - male, 2 - non-binary and 1 participant preferred not to provide this information.
And a note they drop in on the next page:
Each participant received a $100 check as a thank-you for their time, conditional on attending all three sessions, with additional $50 payment if they attended session 4.
I’m not overly worried about the sample size, you can do a lot with that with some good research design. I’m not worried about the WEIRD-ness of the sample. The results of this stuff are mainly going to be relevant for educated students anyway.
I’m very worried about the mindset of the participant. These are clearly college students that could do with $100 and will show up to sit in a room three times to get it. Plus the extra $50 which 18 of them decided to go for too. I used to be that guy too. Good times.
But annoyingly important thing: if you are paying them just to show up, what reason do they actually have to care about what they are doing? There is no incentive to do well, or to take it seriously.
“Well isn’t that true for all lab studies, but given the randomised design, you can still find meaningful differences across groups?”
Yes, but the meaning of the difference has to be analysed through that lens.
Ok, first of all the groups, bolding is mine:
Participants were randomly assigned across the three following groups:
LLM Group (Group 1): Participants in this group were restricted to using OpenAI’s GPT-4o as their sole resource of information for an essay writing task. No other browsers or other apps were allowed;
Search Engine Group (Group 2): Participants in this group could use any website to help them with their essay writing task, but ChatGPT or any other LLM was explicitly prohibited; all participants used Google as a browser of choice. Google search and other search engines had “-ai” added on any queries, so no AI enhanced answers were used by the Search Engine group.
Brain-only Group (Group 3): Participants in this group were forbidden from using both LLM and any online websites for consultation.
And eventually, the actual task:
Once the EEG headset was calibrated, they were introduced to their task: essay writing.
For each of three sessions, a choice of 3 topic prompts were offered to a participant to select from, totalling 9 unique prompts for the duration of the whole study (3 sessions). All the topics were taken from SAT tests.
…
The participants were instructed to pick a topic among the proposed prompts, and then to produce an essay based on the topic’s assignment within a 20 minutes time limit.
<participants receive instructions on what tools they’re allowed use>
All participants were then reassured that though 20 minutes might be a rather short time to write an essay, they were encouraged to do their best.
Firstly, 20 minutes isn’t a short time to write an essay. It is no time.
So when you give a college student this (real) prompt:
“1. Many people believe that loyalty whether to an individual, an organization, or a nation means unconditional and unquestioning support no matter what. To these people, the withdrawal of support is by definition a betrayal of loyalty. But doesn’t true loyalty sometimes require us to be critical of those we are loyal to? If we see that they are doing something that we believe is wrong, doesn’t true loyalty require us to speak up, even if we must be critical?
Assignment: Does true loyalty require unconditional support?”
And give them 20 minutes to write a response, what the hell are we studying? Dwell on that for a moment if you don’t have an answer.
I’ve read that prompt three times just in writing this, and I can’t remember any of the words. It is the exact type of English language reserved only for something like the SAT, where even as a school student you understand this is more game than skill, and will produce whatever bullshit you need to in order for the grader to tell you ‘well played’ and take your B+.
And if you want to pass your SATs and get into college, of course you may as well try bullshit as effectively as possible, but that isn’t what’s happening here. Here, you have already got into college, and now some psychology researcher is giving you $100 to write this stuff again which you already know is nonsense.
Point being, if you were in the ChatGPT group, and you had a choice between totally offloading it to GPT-4o, or painstakingly making up a bullshit answer yourself, you are so obviously going to have it produce some slop for you.
And not only because you see no point in expending the effort, but because these questions are the human version of how AI writing acts deep but has nothing of meaning underneath the surface. This is exactly what LLMs are built on and trained for. The human written answer would be just as meaningless.
So not only do I think participants relying on ChatGPT in this condition is inevitable (i.e. having it write the whole thing vs checking some facts/phrasing), but if a participant didn’t do that, and worked way harder in writing it themselves, then I would have less respect for that participant.
That would be a person who has no respect for their time and efforts. Or, more commonly with Western educated people, a person who will blindly follow any old authority figure for no reason. And a society filled with people like that would not be good. That’s the ideal setting for dystopian tyranny.
Or forget the character aspect, who is smarter? The one who tries or the one who doesn’t? Or are there, perhaps, different ways to think about what smart means?
Yet, I think we already know where this inevitable experimental result is going to take us by the end of this paper.
We then see a classic play of getting you in the right mental zone to take some results for granted.
To ensure you doubly understand that what is being presented to you is hard scientific stuff, we move onto ‘Quantitative Statistical Findings’, which leads with this awesome graph:
Before getting more complicated:
Literally what the fuck is going on here. The results section of this paper feels like traversing those moving staircases in Hogwarts, except now you forget which House you’re even in and how you ended up there, and are just begging to get on the train home again. Oh shit now you’ve taken another wrong turn and there’s 10 heatmaps (?) charging at you:
Is this meta-psychology? Are we behind the curtain? Is this the afterlife? What the fuck is this next thing?
Has an LLM been fed an instruction to write a psychology paper that perfectly illustrates everything wrong with academic research to the point of satire? Maybe the answer to that question can be found here:
HUMAN. 85. ART. 2 (IN YELLOW). 4. RIGHT AT THE INTERSECTION OF ‘DECISION AND LIFE CHOICE’ AND …. ‘XEROX’? (AND THE WORD ‘FROM’ HOVERING VAGUELY NEAR IT)
ARE YOU WITNESSING THIS SCIENCING RIGHT NOW? WE NEED TO END AI STAT, NO ROBOT COULD EVER COME UP WITH THIS MANY ONTOLOGY PAIRS PER TOPIC.
(but if the AI could never do that, why worry about shutting it dow-NOT THE POINT)
I am not exaggerating when I say I am now forgetting what this paper is meant to be about. How’s that for cognitive atrophy?
This at least mentions the word teachers which reminds me it’s something to do with education:
In fact, why do you even need me here at all?
That graph, quite simply, speaks for itself.
The findings aren’t just quant btw. They did qual interviews with the participants, too.
Here, check out the results:
“Cluster 15 shows popularity of using ChatGPT to generate the intro, and mostly in sessions 1, 2, 3, and almost not in session 4”
That circle marked 15 shows that. Somehow. I’m eyeballing that at -5.5 PaCMAPs on the Y axis. Nice.
There is then a section of EEG results which I will skip because I can’t pretend to understand them, although my trust that anyone anywhere may understand them has been severely damaged.
I was expecting at some point to find the main dependent variables (out of a seemingly infinite number), which correspond to those 4 research questions at the start, and to have them clearly marked with the different means per group. If that happened, I didn’t see it. And I ended up at the discussion, which they are hoping you just skipped to at the offset (or never even read, the stuff they want you to see is in the abstract anyway).
And back we go to some words we can understand:
The results of our study offer several intriguing insights into the differences in cognitive and performance outcomes in essay writing tasks for 54 participants, who used LLMs such as ChatGPT, traditional web search, or were tools-free over a span of 4 sessions per participant over a period of 4 months.
‘Several intriguing insights’ is a weird set of words. Let’s see the first one:
We found that the Brain-only group exhibited strong variability in how participants approached essay writing across most topics. In contrast, the LLM group produced statistically homogeneous essays within each topic, showing significantly less deviation compared to the other groups.
i.e. the ChatGPT group used ChatGPT. As they should. The task was utterly pointless and any sensible person would have done the same.
This is added as a key piece of information, also:
Interestingly, in the Brain-only group the social media influence found its way around, here is a quote from one of the essays “So why we are not talking about it on Instagram, for example?”
Do I not speak English as well as I thought I could or something? What do those words mean?
No time to work that out, let’s continue with the findings:
Participants in the LLM and Search Engine groups were more inclined to focus on the output of the tools they were using because of the added pressure of limited time (20 minutes). Most of them focused on reusing the tools’ output, therefore staying focused on copying and pasting content, rather than incorporating their own original thoughts and editing those with their own perspectives and their own experiences.
Why. Would. They. Bother. Incorporating. Their. Original. Thoughts. They. Don’t. Care. About. Your. Bullshit. Essay.
And here’s an example of one of the EEG findings, bolding is theirs:
Interestingly, the Search Engine group exhibited increased activity in the occipital and visual cortices, particularly in alpha and high alpha sub-bands. This pattern most likely reflects the group’s engagement with visually acquired information during the research and content-gathering phase during the use of the web browser. These occipital-to-frontal flows (e.g. Oz→Fp2, PO4→AF3) support the interpretation that participants were actively scanning, selecting, and evaluating information presented on the screen to construct their essays, a cognitively demanding integration of visual, attentional, and executive resources.
In contrast, despite also using a digital interface, the LLM group did not exhibit comparable levels of visual cortical activation. While participants interacted with the LLM via a screen, the purpose of this interaction was distinct: LLM use reduced the need for prolonged visual search and semantic filtering, shifting cognitive load toward procedural integration and motor coordination (e.g. FC6→CP5, Fp1→Pz), as supported by dominant beta band activity in fronto-parietal networks. This suggests a more automated, scaffolded cognitive mode, with reduced reliance on endogenous semantic construction or visual content evaluation.
Meanwhile, the Brain-only group showed the strongest activations outside of the visual cortex, particularly in left parietal, right temporal, and anterior frontal areas (e.g. P7→T8, T7→AF3). These regions are involved in semantic integration, creative ideation, and executive self-monitoring. The elevated delta and theta coherence into AF3, a known site for cognitive control, underscored the high internal demand for content generation, planning, and revision in the absence of external aids.
You are using your brain differently depending on whether you had the help or not.
Which is a measurable difference sure, but so what?
Put on an EEG helmet and do a complicated calculation in your head, or with a calculator, and you will get a similar difference. First of all, it’s the answer that’s important, but whether your effort is also important depends entirely on the task at hand. In this task, your effort does not matter. You are getting $100 anyway.
These distinctions carry significant implications for cognitive load theory, the extended mind hypothesis [102], and educational practice.
Why?
Notice the jump to talking about educational practice here. That is vital.
This paper is now turning to talking about learning again. And this is the angle it will allow itself to be interpreted through in the media hereafter.
But this isn’t a learning task. It’s a simulation of a writing sprint that students do as part of their education. But even as the no-AI group wrote their 20 minute essays, what are they supposed to be learning? Everything they are writing they are producing themselves. There is literally nothing for them to learn. Only the other groups actually have that opportunity, ironically enough.
There is also findings about the LLM group not being able to quote from their own essays. And here they bury one of the craziest pieces of analysis of all, bolding mine:
Search Engine and Brain-only participants did not display such impairments. By Session 2, both groups achieved near-perfect quoting ability, and by Session 3, 100% of both groups’ participants reported the ability to quote their essays, with only minor deviations in quoting accuracy. This behavioral preservation correlates with stronger parietal-frontal and temporal-frontal connectivity in alpha and theta bands, observed especially in the Brain-only group, and to a lesser degree in the Search Engine group. In the Brain-only group, the P7→T8 and Pz→T8 connections suggest deep semantic processing, while Oz→Fz and FC6→AF3 reflect sustained executive monitoring, both of which support stronger integration of content into memory systems.
They are trying to tie quoting to the strengthening of your memory systems. Meaning, they are setting up an implication around LLM use and reduced memory.
This is fucking crazy.
The participants logically had ChatGPT had AI write their pointless essay. Their failure to produce a quote from it, literally any quote, suggests not that they un-practiced brain is losing it’s ability to remember basic stuff. It tells you they never even bothered to read it, because why would they bother to read it? There was no point.
And sure enough, here it is wrapped up and ready to be quoted in the media:
Correct quoting ability, which goes beyond simple recall to reflect semantic precision, showed the same hierarchical pattern: Brain-only group > Search Engine group > LLM group. The complete absence of correct quoting in the LLM group during Session 1, and persistent impairments in later sessions, suggested that not only was memory encoding shallow, but the semantic content itself may not have been fully internalized.
Ok I can’t do this anymore, let’s go to their conclusions section.
Brace yourself:
As we stand at this technological crossroads, it becomes crucial to understand the full spectrum of cognitive consequences associated with LLM integration in educational and informational contexts. While these tools offer unprecedented opportunities for enhancing learning and information access, their potential impact on cognitive development, critical thinking, and intellectual independence demands a very careful consideration and continued research.
The LLM undeniably reduced the friction involved in answering participants’ questions compared to the Search Engine. However, this convenience came at a cognitive cost, diminishing users’ inclination to critically evaluate the LLM’s output or ”opinions” (probabilistic answers based on the training datasets).
‘Cognitive cost’. What does that mean, I hear you cry?
I have no idea. No one does. What we have is a huge set of confusing research questions, methods, outcome measurements, and analyses. And we need some term to cap it off with to make it understandable to anyone beyond our lab.
But the words cognitive cost mean nothing. The words cognitive atrophy mean nothing. The words cognitive debt mean nothing. The fact that they all mean nothing is further reinforced by how they are used interchangeably in the discourse around this stuff and within the actual paper itself.
In fact, the term cognitive debt doesn’t feature in the conclusions at all yet the title of the paper is ‘Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task’
What the hell is this debt and where in the paper did we start accumulating it?
Ctrl-f for ‘debt’ and you find this note buried in the results:
Perhaps one of the more concerning findings is that participants in the LLM-to-Brain group repeatedly focused on a narrower set of ideas, as evidenced by n-gram analysis and supported by interview responses. This repetition suggests that many participants may not have engaged deeply with the topics or critically examined the material provided by the LLM.
When individuals fail to critically engage with a subject, their writing might become biased and superficial. This pattern reflects the accumulation of cognitive debt, a condition in which repeated reliance on external systems like LLMs replaces the effortful cognitive processes required for independent thinking.
Cognitive debt defers mental effort in the short term but results in long-term costs, such as diminished critical inquiry, increased vulnerability to manipulation, decreased creativity. When participants reproduce suggestions without evaluating their accuracy or relevance, they not only forfeit ownership of the ideas but also risk internalizing shallow or biased perspectives.
This is where things get more interesting. Because the ‘cognitive debt’ happens when the ChatGPT group then get to the task brain only (last session), and still show reduced performance and sticking to the style of writing the LLM came up with.
I’ve written before about students using ChatGPT to cheat in university, and still think people make the wrong conclusions about this. They are not robbing themselves of education, they are just showing you they were just pretending to learn anyway.
Think again about the experimental task. SAT essay writing. Has there ever been a clearer example of a ‘creative task’ that in fact encourages no creativity whatsoever, and leads to thousands of highly varied students somehow producing boring homogenous work? Everyone in every condition was pretending to write an original essay, the ChatGPT group just had a tool that let them do that without having to think as much. And so they let the LLM-based arbitrary style stick because the one they could have come up with themselves would have been just as arbitrary.
But this, my friends, is the key thing to remember if you want to understand what’s going on with this paper existing.
The authors have just, out of nowhere, decided what they want to be true. They know what they want the paper title to be, so they introduce this note where they get to define this thing you’ve never heard of, based off the data they’re claiming means things it doesn’t even nearly mean.
And the reason they do this is (1) this is the opinion they already held before the research began, they just needed a convoluted way to give it some authority and (2) cognitive debt is a catchy name.
And catchy names are where the money is at.
But this is where I am really really serious. It is disastrous and heartbreaking to see social science turn into this.
Because all the people writing articles and blogs, youtube videos and podcasts, on stage doing speeches about this stuff, all the policy-workers desperately trying to work out where education should go, they are going to take this as face value because we should live in a world where research abstracts should be able to be taken literally.
But we don’t live in that world, because researchers are wilfully blind about what their experiments mean.
And now, millions of people have had their minds made up on AI in education based on this paper, despite none of them reading it.
Psychology and social science needed to be ready to help the world adapt to AI. They aren’t. Makes me sad.
I wish I had a lighter note to end this on but I don’t.
Cheers.











This is so funny 😂😭and ridiculously sad