Transcribbit frog mascot beside a speech bubble and an open notebook, representing spoken versus written memory
Research
9 min read

Spoken vs Written Memory: What the Evidence Actually Shows

By Kyle RProduct

Spoken vs written memory, tested: the famous reading and hearing percentages are fabricated. Real studies on the modality effect show what you actually retain.

You've probably heard the claim that spoken vs written memory splits into a tidy pyramid: we supposedly remember 10% of what we read, 20% of what we hear, and 90% of what we do. It's one of the most repeated statistics in corporate training decks, and a 2014 investigation by four researchers found no study behind any of those numbers [1].

The real research is messier and more useful. For the last few items in a list you just heard, spoken wins by a small margin, for a few seconds. For almost everything else, especially anything you need to review, search, or act on later, a written record beats a recording you can only play back from the start.

This piece works through the actual cognitive science: the modality effect, the production effect, and why a transcript beats a re-listen.


Spoken vs written memory at a glance

QuestionShort answer
Is the "10% read, 20% hear" pyramid real?No traceable study supports it [1]
Do spoken items have any real memory edge?Yes, briefly, for the last few items in a list (recency) [2][3]
Does that edge last?No, it fades within seconds once attention shifts [4]
What wins for delayed recall?Written text, in controlled testing [5]
Why does text hold up better long-term?You can re-read, search, and quiz yourself; speech vanishes once spoken
Does saying things aloud help too?Yes, modestly, within a single study session [7]

The pyramid stat you've heard is fake

The "10/20/90" chart, often shown as a cone or pyramid, is usually credited to Edgar Dale's 1946 "Cone of Experience." Dale never attached percentages to his cone. The numbers were bolted on decades later by unknown hands, and every version disagrees slightly on the exact figures, which is itself a tell: real experimental data doesn't round so neatly to multiples of five and ten.

Dean Subramony, Michael Molenda, Anthony Betrus, and Will Thalheimer traced the chart's history in a 2014 special issue of *Educational Technology* and concluded there is no body of research behind any version of it [1]. So we won't use it here. Instead, this article sticks to named studies you can look up.


What the modality effect actually says

The modality effect is the real, tested finding hiding behind the myth: whether you hear or read a list of items changes how well you recall it, and the effect is strongest for the last few items.

Catherine Penney reviewed decades of this work in *Memory & Cognition* in 1989 and proposed that spoken and written words are processed through two separate streams, each with different properties [2]. In her account, a spoken word leaves a brief acoustic trace that lingers for a moment after you hear it, which is why the last few items in a spoken list often come back to you more easily than the last few items you silently read [2].

That is the entire basis for saying speech "wins." It's real, and it's narrow.


Why the last few things you hear stick, briefly

Bennet Murdock mapped this out in 1962 with the classic serial position curve: people recall the first few and the last few items in a list well, and the middle worse, whatever the list length [3]. Items near the start get rehearsed into longer-term memory. Items near the end are still sitting in short-term storage when recall starts, which is why they come back so easily.

Murray Glanzer and Anita Cunitz showed just how fragile that end-of-list advantage is. In 1966, they had people count backward for a few seconds before recalling a list, and that alone wiped out the recency boost while leaving the earlier items intact [4]. The last things you heard weren't gone from memory in general, they were gone from a short-term store that a few seconds of distraction can empty.

That's the catch with voice notes and meetings. The moment your attention shifts, the part of the message you'd otherwise remember best disappears.


Where written text pulls ahead

Once you look past the first few seconds after hearing something, the advantage flips. In a separate 1989 paper in the *Quarterly Journal of Experimental Psychology*, Penney tested delayed recall and recognition rather than immediate recall, and visual presentation beat auditory presentation across her conditions. The paper's title says it plainly: visual is better than auditory [5].

A written record persists. You can slow down, jump back a paragraph, and reread the one line you missed, none of which a live stream of speech allows. That gap is well documented in research comparing reading and listening for retention and comprehension, which finds reading has a real edge once the material gets long or complex.

Reviewing only helps if you do something with the material, though. Henry Roediger and Jeffrey Karpicke showed that testing yourself on material produces stronger long-term retention than simply rereading it again [6]. A written record is what makes that kind of active recall possible in the first place. A voice note you've already listened to once gives you nothing to quiz yourself against unless you replay the whole thing.


Why long voice notes are so hard to retain

This is where spoken vs written memory stops being an academic question and starts costing you time.

A five-minute voice note takes five minutes to get through, whatever your reading speed, and you cannot skim it, search it, or jump to the part that matters. If you miss a name or a date halfway through, your only option is to start again from wherever you paused. That's a direct, practical consequence of the transient nature of speech: unless it's recorded and transcribed, spoken information exists only for the moment it's spoken.

It's also why long, unstructured voice notes tend to leave the person receiving them worse off, not better: the sender got to organize their thoughts once, out loud, and the listener has to reconstruct that structure from a single, unskimmable pass. A written transcript turns the same message into something you can scan, search for a keyword, and return to a week later.

If you regularly get long voice notes you need to act on, that's the decision point: keep re-listening from the start every time, or convert the note to text once and read it in a fraction of the time. Transcribbit does the second one for free, up to 50 transcriptions a month. Get Started for Free.


The production effect: saying it helps too

There's a genuine, if narrower, upside to speaking that's worth keeping alongside the case for written records.

Colin MacLeod and colleagues named the production effect in a 2010 paper: words read aloud during study are remembered better than the same words read silently, most likely because saying something aloud creates an extra, distinctive trace at the moment you encode it [7]. It shows up reliably within a single study list, where the spoken items stand out against a background of silent ones.

The mechanism is still debated. A 2014 analysis by Daniel Algom and colleagues argued that "distinctiveness" is actually hiding at least two separate effects that don't always move together [8]. So treat the production effect as real but not fully explained.

None of this contradicts the written-text advantage above. The strongest setup is usually a written record you also engage with actively, not a silent skim and not a recording you only ever hear once.


Frequently asked questions

Do you remember what you read or hear better?

It depends what you're measuring. For the last few items in a list recalled immediately, spoken input has a small, real edge (the modality effect) [2][3]. For anything recalled after a delay, written presentation did better in controlled testing [5]. The honest summary: speech wins a narrow, short-lived contest, and text wins the one that matters for actually keeping information.

Is the "we remember 10% of what we read" statistic true?

No. It comes from a chart with no traceable research behind any of its numbers, as documented by a 2014 investigation into the chart's origins [1]. Treat any specific retention percentage attached to reading, hearing, or "doing" as unsupported unless it cites a named study.

What is the modality effect in memory?

It's the finding that whether information arrives by ear or by eye changes how well you recall it, particularly for the most recent items in a list, which are recalled better when heard than when read [2]. It fades within seconds once your attention moves elsewhere [4].

Why are voice notes so hard to remember accurately?

Because speech is transient. Once a spoken sentence is over, it's gone unless it's recorded, so you cannot re-scan it the way you can reread a sentence. A long voice note also removes the option to skim ahead or search for a name or date, which written text allows by default.

Does reading something aloud help you remember it?

Often, yes, within limits. The production effect shows that words spoken aloud during study are recalled better than the same words read silently, an effect named by MacLeod and colleagues [7]. It's a real but modest boost, and researchers still disagree on exactly why it happens [8].

Is rereading the best way to remember something you've read?

Not the best, no. Roediger and Karpicke found that testing yourself on material beats simply rereading it for long-term retention [6]. A written record is most valuable as something you can quiz yourself against, not just something you look at twice.


Final Thoughts

Spoken vs written memory doesn't have a single winner, but it isn't a coin flip either. Speech has one real advantage: the last few things you hear are, for a few seconds, unusually easy to recall. That's it. Once any delay enters the picture, once you need to review something, search it, or act on it later, a written record consistently does more for you than a recording ever will.

That's the case for turning spoken content into text. Billions of voice messages get sent every day now, 7 billion a day on WhatsApp alone by the company's own count, and most of them are heard once and then forgotten in the way the research above predicts. A transcript is what turns a one-listen message into something you can actually keep.

Transcribbit turns forwarded WhatsApp voice notes into clean, searchable text in over 50 languages, free for up to 50 transcriptions a month, with no audio kept afterward. Get Started for Free.


References

  1. The Mythical Retention Chart and the Corruption of Dale's Cone of Experience (Subramony, Molenda, Betrus & Thalheimer, 2014, Educational Technology, 54(6), 6-16) - worklearning.com
  2. Modality effects and the structure of short-term verbal memory (Penney, 1989, Memory & Cognition) - link.springer.com
  3. The serial position effect of free recall (Murdock, 1962) - semanticscholar.org
  4. Two storage mechanisms in free recall (Glanzer & Cunitz, 1966) - semanticscholar.org
  5. Modality Effects in Delayed Free Recall and Recognition: Visual is Better than Auditory (Penney, 1989, Quarterly Journal of Experimental Psychology) - journals.sagepub.com
  6. The Power of Testing Memory (Roediger & Karpicke, 2006) - psychnet.wustl.edu
  7. The Production Effect: Delineation of a Phenomenon (MacLeod, Gopie, Hourihan, Neary & Ozubko, 2010, JEP: LMC) - uwaterloo.ca
  8. The production effect in memory: multiple species of distinctiveness (Algom et al., 2014, Frontiers in Psychology) - pmc.ncbi.nlm.nih.gov

More research and writing from Transcribbit: https://transcribbit.io/blog