Neurako Learn
Learning Science

Dual Coding

Paivio's dual coding theory and Mayer's multimedia learning research: why combining words and images produces stronger memory than either alone — and how to use it on flashcards.

Last updated 2026-05-23

~7 min read~15 min to applyLearning Science

Dual coding theory, developed by Canadian psychologist Allan Paivio (1971, 1986), proposes that the mind processes information through two linked but independent systems — one verbal, one visual. Information encoded in both systems has two retrieval paths, producing stronger memory than either system alone. Richard Mayer's later "Cognitive Theory of Multimedia Learning" refined this with practical principles for when image-plus-text combinations help and when they hurt.

Key Takeaways

  • The brain has two semi-independent processing channels: verbal (words, language) and nonverbal (images, sounds, spatial) - Information stored in both channels has two retrieval paths, producing better recall than single-channel storage - Paivio (1971, 1986) showed pictures are remembered better than words and concrete words better than abstract words — both consistent with dual coding
  • Mayer's research (2009) found multimedia lessons combining narration and visuals produced large gains on transfer tests, but only when the two channels were aligned and not redundant - For flashcards: a visual on the back of a card (mental image, photo, diagram) can strengthen memory — but only if the image is relevant, not decorative

The basic theory

Allan Paivio proposed in his 1971 book Imagery and Verbal Processes that human cognition includes two systems for representing information:

  • The verbal system handles language — spoken words, written text, abstract concepts.
  • The nonverbal system handles imagery — visual mental images, sounds, smells, spatial layouts.

Each system has its own internal representations and processing rules. Most concepts can be encoded in either or both. The word "dog" can be stored as a verbal label, as a visual mental image, or — most strongly — as both, linked together.

Paivio's experimental work demonstrated three relevant patterns:

  1. The picture superiority effect. People remember pictures better than words. In free-recall studies, pictures of objects are recalled at roughly twice the rate of the corresponding words.

  2. The concreteness effect. Concrete words (like "apple" or "table") are remembered better than abstract words (like "justice" or "intention"). Concrete words can be encoded in both systems; abstract words primarily exist in the verbal system.

  3. Bilingual independence. When the same concept is learned in two languages, the two verbal labels are stored more independently than a unified theory would predict — but both connect through the shared nonverbal system. This is one reason translation-based vocabulary study works.

Mayer's multimedia extension

Richard Mayer, building on Paivio (and on John Sweller's cognitive load theory), developed the Cognitive Theory of Multimedia Learning (CTML) in the late 1990s and early 2000s. Mayer's research focused specifically on combining narration with diagrams, animations, or other visual content.

His major findings, distilled into design principles:

The multimedia principle. People learn better from words and relevant pictures than from words alone — when the two are aligned. Mayer's lab experiments consistently showed substantial gains (often 50–100% on transfer tests) when text was paired with a relevant diagram.

The modality principle. Spoken narration plus visual diagrams outperforms on-screen text plus the same diagrams. Reason: the eye is already loaded processing the diagram; adding text forces the visual channel to do extra work, while spoken narration uses an otherwise idle auditory channel.

The redundancy principle. Adding on-screen text that duplicates spoken narration hurts learning. The redundant text creates split attention with no informational gain.

The coherence principle. Adding interesting-but-irrelevant material ("seductive details") to a lesson reduces learning. Background music, decorative images, fun anecdotes — they sound like they make lessons more engaging, but they consume working memory without contributing to the topic.

The signaling principle. Drawing attention to important elements (arrows, highlighting, labels) improves learning. Without signaling, learners may not allocate attention to the right parts of a complex image.

The spatial contiguity principle. Labels should be placed near the parts they describe, not in a legend at the bottom of the diagram. Forcing eye-movement between label and referent imposes extraneous cognitive load.

These principles compose. The strongest multimedia instruction is coherent (no decoration), signaled (attention directed), contiguous (labels near referents), and modality-appropriate (narration over text when visuals are present).

Why dual coding works at the brain level

Functional MRI studies have largely confirmed the two-system architecture, though the modern view is more nuanced than strict separation. Words and images activate partially overlapping but distinct networks. The verbal system involves Broca's and Wernicke's areas, the angular gyrus, and parts of the temporal lobe; visual imagery activates parts of the occipital cortex (especially when generating mental images) and the parahippocampal region.

A 2019 review by Madan and Singhal found consistent fMRI evidence for dual-pathway encoding when participants studied concrete words. The hippocampus integrates the two pathways, which is part of why memory palaces (which deliberately combine verbal content with spatial-visual context) are so effective.

Where the theory has been challenged

Dual coding is not without critics. Propositional theorists argue that all knowledge is ultimately stored in an abstract, amodal format, and that the "image" experience is reconstructed at retrieval rather than stored as such. This debate is decades old and unresolved at the theoretical level — but the practical finding (words + relevant images beat words alone) holds regardless of which underlying account is correct.

A 2025 piece by educational researcher David Didau pushed back on what he called "the dual coding delusion" — the tendency to interpret dual coding as "add any image to any lesson." His point: dual coding only helps when the image is conceptually coherent with the verbal content, when the two are processed together in working memory, and when the image is not merely decorative. Random clipart does not produce dual coding benefits.

This caveat is important. The Mayer principles are not "always add pictures." They are "add the right pictures in the right way."

How to use dual coding on flashcards

Several practical applications follow from dual coding research:

Add relevant images to cards where they aid encoding. Anatomy cards benefit enormously from labeled images. Chemistry mechanisms benefit from structural diagrams. Geography facts benefit from maps. Historical events benefit from related photos or paintings.

Mental imagery counts. You don't need to attach a photograph to a card to use dual coding. When you encounter a new vocabulary word, deliberately imagine a vivid scene linking the word to its meaning. This is the same principle behind the keyword method and memory palaces.

Avoid decorative images. A flashcard about French verb conjugations does not benefit from a stock photo of the Eiffel Tower in the background. That image consumes attention without contributing to learning — exactly Mayer's coherence-principle warning.

Use labels close to what they describe. If you're studying a labeled anatomical diagram, the label should be inside or adjacent to the structure, not in a separate legend below. Cards with split-attention layouts impose extraneous load.

Don't duplicate. Diagrams plus minimal text labels are better than diagrams plus long descriptive paragraphs. The text repeats what the diagram already shows; the duplication wastes working memory.

In Neurako

Neurako supports image cards through both manual upload and AI image capture. When you photograph a textbook page or whiteboard, the system extracts both the visual content and any associated text. This dual extraction matches the dual coding principle: you study with both representations, not just transcribed text. See Capture from Images for the workflow.

A worked example

Suppose you're studying the brachial plexus — the network of nerves running from the spinal cord through the shoulder.

Words-only card (weak): "The brachial plexus is formed from spinal nerves C5, C6, C7, C8, and T1, which combine into trunks (upper, middle, lower), then divisions, then cords (lateral, posterior, medial), then terminal branches including the musculocutaneous, axillary, median, radial, and ulnar nerves."

This is a single multi-fact card that exceeds working-memory capacity for a novice.

Diagram card (stronger): A labeled anatomical diagram showing the C5–T1 roots, the trunk-division-cord-branch structure, and each terminal nerve. The learner sees the spatial relationship — which roots feed which trunk, which cord gives rise to which terminal branch — directly.

Diagram + selective text (strongest): The same diagram, with key relationships emphasized by signaling (arrows, color coding for terminal branches by origin cord), and brief text providing a memorable phrase for the levels ("Real Texans Drink Cold Beer" for Roots, Trunks, Divisions, Cords, Branches).

The text alone, the diagram alone, and the text+diagram combinations are all available study material. The combination wins because two retrieval paths exist, and the verbal phrase provides a structural skeleton the diagram's details can hang on.

Sources

  1. Paivio, A. (1971). Imagery and Verbal Processes. Holt, Rinehart, and Winston.

  2. Paivio, A. (1986). Mental Representations: A Dual Coding Approach. Oxford University Press.

  3. Clark, J. M., & Paivio, A. (1991). Dual coding theory and education. Educational Psychology Review, 3(3), 149–210. https://doi.org/10.1007/BF01320076

  4. Mayer, R. E. (2009). Multimedia Learning (2nd ed.). Cambridge University Press.

  5. Mayer, R. E., & Moreno, R. (2003). Nine ways to reduce cognitive load in multimedia learning. Educational Psychologist, 38(1), 43–52. https://doi.org/10.1207/S15326985EP3801_6

Ready to turn this into a study ritual?

Start studying with Neurako

Related reading

On this page