Language as a cognitive scaffold: inner speech and chain-of-thought
A term essay: inner speech in humans and chain-of-thought prompting in LLMs do the same job, using language to break hard problems into steps. The differences a well-known comparison found are mostly methodology.
Written for a COG250Y1 (Introduction to Cognitive Science) essay assignment, submitted April 7, 2025. Rewritten here from the APA manuscript: the citations became inline links and the register is looser; the argument is unchanged. Recompiled as submitted, the manuscript and its topic proposal are available as PDFs.
Beyond carrying messages between people, language structures cognition itself. The Sapir-Whorf hypothesis holds that linguistic categories shape thought; the Language of Thought Hypothesis replies that cognition runs on something deeper than any natural language. Large language models have quietly recontextualized this old debate. Chain-of-thought prompting makes a model reason step by step in words, and the technique looks suspiciously like inner speech, the silent verbalization humans use to manage complex reasoning. I argue the resemblance is structural: in both humans and LLMs, language works as a cognitive scaffold that breaks complex problems into sequential, manageable steps, and the differences between the two are mostly about who erects the scaffold.
Shaper and medium of thought
The early Sapir-Whorf formulations said language determines how you perceive and categorize the world. The strongest evidence is more modest. Boroditsky showed that English speakers reach for horizontal metaphors when thinking about time (“ahead”, “behind”) while Mandarin speakers use vertical ones more often: language biases attention and processing, and the effect is modulatory. Later cross-linguistic work sharpened this into a neo-Whorfian position: lexical categories measurably shift memory and attention, yet speakers of different languages share one perceptual framework underneath. Language organizes cognitive material.
The Language of Thought Hypothesis pulls the other way: cognition operates over an internal symbolic system, “Mentalese”, independent of any spoken language. Recent reviews find mental representations behaving compositionally, the way linguistic syntax does, and current versions of the hypothesis allow a dual system: innate representational machinery, refined and shaped by natural-language input. Both camps end up somewhere compatible: language biases and organizes thought as a flexible tool rather than a rigid determinant, which is exactly the property a scaffold has.
What the brain shows
People with aphasia, whose language production is severely impaired, keep mathematical reasoning and spatial navigation largely intact: core cognition survives without language. Neuroimaging agrees, up to a point. Classical language areas engage selectively for linguistic tasks while abstract reasoning recruits frontal and parietal circuits, but during complex tasks the language regions co-activate with the circuits for executive control and working memory. Language-based rehearsal helps memory retrieval, verbal labels improve discrimination and categorization, and disrupting language areas with TMS impairs verbal reasoning while sparing nonverbal problem-solving. Thought runs without language; language organizes and amplifies it. That is what a scaffold does.
Inner speech
Alderson-Day and Fernyhough distinguish condensed inner speech, brief directive cues, from expanded inner speech, a full narrative dialogue with yourself. Both do real work. In Baddeley’s model of working memory the phonological loop rehearses verbal information through inner speech, and articulatory suppression, forcing someone to repeat an irrelevant word, sharply degrades memory performance. Jorba and Vicente describe how adults monitor and adjust their own problem-solving through internal dialogue, talking themselves through the steps of a task. Even the phrasing matters: third-person self-talk measurably improves emotion regulation without recruiting extra cognitive control. Inner speech is not a byproduct of thinking. It is one of the mechanisms.
Chain-of-thought, the artificial analogue
Wei and colleagues showed that a few exemplars of step-by-step reasoning dramatically improve LLM performance on arithmetic and symbolic tasks. The prompt makes the model think aloud, and the parallel to human practice is explicit in the prompting literature: humans also perform better when they verbalize their process, and explaining a problem, even to yourself, promotes the kind of generalization that shows up reliably in education research. The follow-up techniques double down on process over answer: sampling multiple chains and voting, searching over trees of intermediate thoughts. And the chain is legible, which is worth something on its own: in medicine, reasoning prompts make the model’s diagnostic path auditable, the machine equivalent of a student required to show their work.
The reported divergence, examined
The strongest counterargument comes from Yax, Anlló, and Palminteri, who gave new variants of classic cognitive tests to both humans and LLMs. Chain-of-thought prompting markedly improved the models; analogous prompts did not much improve the humans. They read this asymmetry as a difference between human and machine reasoning. I think the study shows something else.
First, the human sample was psychology students, a population already known to be unrepresentative of human cognition at large; participants trained in formal logic might have responded to structured prompts quite differently. Second, the test items were rewritten to avoid training-data contamination but never validated for difficulty, so nobody knows whether the human and model versions of the task were even comparable. Third, and most importantly, the asymmetry has a simpler explanation: humans already run inner speech. An external prompt telling a person to think step by step is redundant scaffolding on top of scaffolding they generate themselves, while a model, which has no spontaneous inner dialogue, gets its entire scaffold from the prompt. Read this way, the finding is the parallel showing up in the data.
The positive case is that both systems lean on language for the same operation: decomposition. Dove calls language a neuroenhancement, Fernyhough and Borghi describe inner speech as a cognitive tool, and Lupyan and Bergen argue language effectively programs the mind. On the machine side, Goldstein and colleagues found shared computational principles between human language processing and deep language models. The mechanisms differ, spontaneous in one case and externally cued in the other, but the function is the same.
What follows
If language is the scaffold, some things follow. For education, explicit verbal strategy, thinking aloud and structured self-talk, is the mechanism working as designed. For AI, reasoning-in-language is also a window into the system, and architectures that borrow the developmental role of inner speech are a live research direction, with language treated as a cognitive and social tool for machines too. The parallel even predicts failures: chain-of-thought hurts model performance on exactly the kinds of tasks where deliberation hurts humans, which is a strange coincidence unless the scaffold really is shared. And if it is shared, then the debate about whether these systems understand anything has to reckon with the fact that human understanding leans on the same linguistic machinery it wants to reserve for itself.
The differences the critics measure are real, and they are differences in how the scaffold gets built: humans grow their own, and models need it handed to them in the prompt. The scaffold itself is shared.