Skip to content

What Paper Still Earns in a Digital-First Study Stack: Three Places It Works, and One It Doesn't

Aug 6, 2026 1 min
TL;DR Once studying went fully digital, paper still holds three places with real evidence: reading (paper over screens at g ≈ −0.21 across 171,055 participants, widening to 0.35–0.48 when scrolling is required), writing while you answer (on screen, harder questions draw less scratch work, not more), and drawing (45% recall against 20% for writing). The one claim most people lean on, that handwritten notes stick better, spans −0.008 to +0.248 across four meta-analyses with no consensus.

🌏 中文版

Learning something today has a default path, and it is digital end to end. Online courses, PDFs, a notes app, a tablet and stylus, a spaced repetition scheduler. You can run the whole chain without touching paper.

So what is left for the physical side? I assumed the answer would be mostly sentiment. It isn't. Paper still does measurable work, and the effect sizes are larger than I expected, just not in the place most people point to.

The next three sections cover where physical holds up, strongest evidence first. The fourth covers the one place it doesn't, which happens to be the argument people reach for most often.

Where the physical advantage is clearest: reading

Meta-analysisScaleEffect size
Delgado et al. 201854 studies / 171,055 participantsg = −0.21 (paper over screen)
Clinton 2019g = −0.25
Salmerón et al. 202449 studies (handheld devices specifically)g = −0.113 / −0.103
2025 network meta-analysis56 studies / 4 device typesRanking: paper > tablets > e-readers > computers > smartphones

A 171,055-participant sample and four meta-analyses pointing the same way. This is the least contested block in the whole field.

It also comes with boundaries, and knowing them is what tells you when printing is worth it.

Text type first. In Delgado's data informational text gives g = −0.27 while narrative-only text gives g = .01, which is no difference at all, and Clinton found the same shape. Novels on a Kindle are fine. Textbooks and papers are where it bites.

Then time pressure. Time-constrained reading gives g = −0.26 against −0.09 for self-paced. Exams and deadline reading lose the most.

The last boundary may be the important one, and it loosens the whole paper-versus-digital frame. The 2025 network meta-analysis split scrolling out: when scrolling is required, paper's advantage is g = 0.35–0.48; when it is not, the range drops to 0.03–0.12 with no reliable difference. What wins may not be paper so much as pagination. A paginated PDF or E-ink page behaves close to paper, an infinitely scrolling web page does not. Worth flagging that Delgado 2018 found scrolling was not a significant moderator, so the two conflict, and the 2025 result has not been replicated.

On mechanism, the most persuasive account is not eye strain but broken metacognitive calibration. Screen readers systematically overestimate how well they understood and therefore stop investing effort early. Clinton said exactly this in interviews. It also explains why time pressure amplifies the effect, since overconfident readers under time pressure are the least likely to go back and reread.

One counterintuitive detail. Delgado regressed effect size on publication year from 2000 to 2017 and found paper's advantage growing over time. If unfamiliarity with digital tools were the cause, it should have shrunk.

Second place: whether you write anything while answering

Prisacari & Danielson (2017) compared computer- and paper-based quizzes in undergraduate general chemistry. Scores and subjective cognitive load showed no difference, but students used scratch paper significantly more on the paper version, and the gap was larger on harder questions.

Pengelley, Whipp & Rovis-Hermann (2023) replicated the pattern with a repeated-measures design and N = 263 Western Australian year-9 students. The result comes in three layers, and any citation needs all three.

First, paper produced significantly higher scores on difficult questions, plus higher cognitive load and scratch paper use across all paper questions. Second, once working memory capacity was controlled, the main effects of mode on score and on both cognitive load measures were no longer significant, and only the mode-by-difficulty interaction survived. Third, the scratch paper pattern: on paper, harder questions drew more written work, while on computer the trend reversed and harder questions drew less.

The third layer is the sharpest. The authors conclude that "these results contradict previous findings that computer-based testing can be implemented without consequence for all learners".

Move a problem onto a screen and students stop writing exactly when writing matters most. Scores do not necessarily drop right away; the problem-solving behaviour changes. So paper's contribution here is not that writing is better, it is that paper makes writing cheap enough that you don't skip it.

Third place: drawing

Wammes et al. (2016) found across seven experiments that drawn words were recalled about 45% of the time against about 20% for written words. The effect held up well enough that the authors called it reliable and robust.

A tablet can do this too, but the friction differs. On paper, an arrow, a box, or a line joining two concepts costs nothing; in most notes apps you first switch tools, pick a colour, set a width. The gap is small, and small gaps decide whether you actually draw.

The one that doesn't hold: "handwritten notes stick better"

This is the argument people reach for most often to defend paper, and it is the only one of the four with no consensus behind it.

Meta-analysisScaleEffect sizeRelationship to digital distraction
Allen et al. 202014 studies / 3,075 participantsr = −.142 (d ≈ 0.25)Includes self-report and survey studies; does not separate distraction
Flanigan et al. 202424 studies / 3,005 participantsg = 0.248, 95% CI [0.181, 0.315]Experimental and quasi-experimental only, but distraction is not an inclusion criterion
Lau 202233 reports / 88 effect sizesg = 0.144, 95% CI [0.023, 0.265]Experimental and quasi-experimental, secondary plus college
Voyer et al. 202236 articles / 77 effect sizesg = −0.008, 95% CI [−0.18, 0.16]States its purpose as unconfounded by distractions

The ordering looks tidy. The more a design separates writing medium from digital distraction, the smaller the effect, until it hits zero at Voyer. Voyer's team read it the same way in their abstract:

an apparent advantage of longhand notetaking reported in some previous studies can be explained at least partially by distractions from notetaking by other applications that are present only with digital devices.

They also wrote a line you rarely see in a journal article: "We began this meta-analysis fully expecting to find evidence favoring the use of longhand notetaking… two of us who are instructors often suggested to students at the beginning of each semester that longhand…" They expected to confirm handwriting and came out with zero.

That ordering is consistent with the distraction hypothesis rather than proof of it. The four differ simultaneously in population, design, outcome measure, and search era, so nothing is held constant except the question.

The founding study, Mueller & Oppenheimer (2014), is in worse shape than most retellings suggest. Urry et al. (2021) ran a preregistered replication with N = 142 against the original's 65. The headline conceptual-application effect went from g = 0.34 to g = −0.13, flipping direction, and differing significantly from the original, t(139.03) = −2.78, p = .003. In Morehead et al. (2019), even the group that took no notes at all did not fall significantly behind.

One finding does hold up every time: typists write more words and overlap more verbatim with the speaker. In Urry's replication verbatim overlap came out at g = −0.85 against the original's −0.93, not significantly different (p = .326). That part replicated cleanly.

None of this is bad news for using paper. It changes the reason. The live question is probably not the hand movement but the fact that typing is smooth enough that you drift into transcribing without noticing, and paper works by throttling input bandwidth until you have to choose what matters. Defending paper with "handwriting burns it into your brain" means standing on the weakest ground available.

That brain-connectivity study cannot carry the argument

The 2024 high-density EEG study by Van der Weel & Van der Meer found greater theta- and alpha-band connectivity during handwriting than typing, and got relayed everywhere as proof that handwriting is better for learning.

The methods section says three things. The handwriting condition used a digital pen on a touchscreen, not paper. The typing condition used the right index finger only. Each trial ran 25 seconds with just the first 5 seconds recorded, and the paper contains no memory or learning test at all.

What it supports is that pen movements are more complex than key presses. It cannot support handwriting improving test scores, and it certainly cannot support paper beating screens, since its handwriting condition was itself on a screen.

Six hybrid modes

Han et al. (2021) gave the formal HCI definition at DIS: a hybrid paper-digital interface is "any interface embedding digital or electronic functionality in physical paper to enable its use as an input or output device". The six modes below are my own practical sorting rather than a citation of theirs.

ModeHow it worksRepresentative toolsMain friction
1 Paper in, digital outWrite on paper → scan → OCR → notes systemRocketbook, MathpixSync tax; OCR depends on handwriting
2 Digital prompt, paper answerApp schedules and prompts, paper carries the workingAnki plus paper, Skritter prompts with paper dictationApp never sees what you got wrong
3 Digital source, paper processingVideo or PDF on screen, drawing and notes on paperCornell notes, sketchnotesAlmost none
4 Paper as interfacePaper is a controller, not a notebookPlickersClassroom only
5 Overlay and live syncCapture happens as you writeAnoto dot paper, SpARklingPaperHardware friction is fatal
6 Paper-like digitalKeep the writing, drop the physical paperreMarkable / Supernote / BOOXDevice cost

Mode 2 is the most underrated of the six, and it has direct evidence behind it. What it does is put back the writing behaviour the screen quietly removed, and it needs no scanning, so there is no sync cost.

Mode 3 runs on the drawing effect covered earlier, and paper's contribution is that the tool friction rounds to zero.

Mode 4 shines where resources are scarce. Plickers gives each student a printed QR card held with their chosen answer facing up, and the teacher scans the whole room with one phone, compressing a device per student into a sheet of paper per student.

Mode 5 is the prettiest in the lab and the most likely to fail in practice. In a two-year longitudinal classroom study from Stanford HCI, 8 of 18 students abandoned the Anoto digital pen because the pen was bulky and needed charging every day. Hardware friction erased every software benefit.

Devices and off-the-shelf programmes

E-ink writing tablets deserve to move up a rung in this analysis. They buy both the absence of an app store and page turns instead of scrolling, and those are the two evidence-backed sources of paper's advantage. Official pricing as of 2026-08-05:

DevicePriceScreenPen
reMarkable Paper Purefrom $39910.3" mono, 226 PPI, 360 gIncluded
reMarkable Paper Pro Movefrom $4497.3" colour, 264 PPI, 230 gIncluded
reMarkable Paper Profrom $62911.8" colour, 229 PPI, 525 gIncluded
Supernote Nomadfrom $3297.8"Separate, from $65
Supernote Manta$50510.7"Separate, from $65
BOOX Go Color 7 (Gen II)$289.997" colourSeparate, from $45.99
BOOX Note Air5 C$529.9910.3" Kaleido 3, Android 15Separate, from $45.99

The pen is the easiest line item to miss. All three reMarkable models include the Marker, while Supernote and BOOX add $46 to 100. reMarkable Connect runs $3.99/month after a 50-day trial; several comparison sites list $2.99, so use the official page.

In Taiwan the ready-made digital half is Adaptive Learning Platform (因材網) from the Ministry of Education, which does knowledge-structure diagnosis, and Junyi Academy with 39,000+ videos and 91,000+ exercises. Junyi's own case library includes a teacher interview titled "Using differentiation and a paper-plus-digital combination to help every child succeed", so classroom practice was hybrid all along.

Japan's two largest correspondence-education programmes are a neat commercial contrast. Benesse's Shinken Zemi elementary course lets parents choose between Challenge Touch (tablet-led, though it still ships some paper material) and Challenge (paper-led), with human "red pen teacher" marking shared by both. Annual grade-4 pricing is 68,400 yen and 70,400 yen respectively, so the paper option is slightly the more expensive one. It also ships a paper workbook twice a year whose stated purpose is "practice writing with a pencil the way you do at school", which is mode 2 in commercial form.

Smile Zemi takes the opposite route with tablet only and no paper, from 3,630 yen per month plus 10,978 yen for the device. One of its selling points is that "there are no apps or games, so there is no temptation and children can concentrate". A purely digital product marketing itself on the absence of distraction converges on Voyer's hypothesis from the commercial side.

Benesse, meanwhile, names something paper cannot do: "only digital can judge stroke order". Real-time evaluation of stroke order and stroke endings genuinely is a digital-only capability.

Where to start

The least demanding combination is mode 2 plus mode 3, with no scanning. The app handles scheduling and prompting, paper handles working and drawing. No sync cost, no device cost, and it keeps most of what has evidence behind it, namely retrieval practice, spaced repetition, drawing, and reopening the scratch paper the screen closed, while avoiding the known failure modes.

If you only change one thing, it should not be your pen. Move the informational long-form reading you actually need to absorb off scrolling screens and onto paper or a paginated E-ink page. That is the largest-sample, largest-effect, steadiest result in the whole review.

If you are willing to spend money, buy an E-ink tablet with no app store.

When designing your own setup, three decisions determine whether it survives. Whether syncing is manual or automatic, since manual means one more daily chore and usually means quitting within three weeks. Whether paper is an asset or a consumable, which decides whether you need OCR. And which side is the single source of truth, because when both are, you can never find anything.

Limits of this piece

Several things have no settled answer, and saying so beats hiding it.

Whether review amplifies or erases the handwriting advantage is unresolved, and both sides are too fragile to build on. Lau's multivariate model shows review compressing the advantage from g ≈ +0.47 to +0.05 (review coefficient −0.414, p = .009), but the same variable is not significant univariately (p = .091) and the model uses only 14 of 33 reports. Lau writes that "it is possible that it is simply an artifact". Flanigan concludes the opposite, with review g = 0.421 against 0.208 without, but that cell holds only 9 effect sizes.

Whether the handwriting advantage concentrates in conceptual items is also unsettled. Flanigan's conceptual effect is the smallest of three at 0.199, Urry's is the largest at 0.14, the directions are opposite, and neither is significant.

The scrolling result carries the most weight in my recommendations while resting on a single unreplicated study, and it conflicts with Delgado 2018's moderator analysis.

The deepest problem comes from Lau's methodological critique, which he calls the fundamental problem of modality research. Randomly assigning participants to a writing medium also indirectly assigns them a writing style, and the two cannot be separated. Of the 33 reports, only 2 manipulated writing style as an experimental factor, and transcription capacity "as far as I can tell has never been controlled for in a note-taking study". This whole literature may believe it is measuring the medium while actually measuring the behaviour the medium induces.

His verdict on the literature's external validity deserves quoting directly:

while the effect size observed in this systematic review and meta-analysis is probably true, we do not know much about how these findings, based on a set of narrowly designed studies, will translate into a more practical context.

Most studies look like this: a single 10 to 15 minute video lecture on unfamiliar content, an author-written short quiz, administered immediately or at most a week later. Nothing follows anyone using a hybrid system for an entire semester.

So every recommendation here is a combination inferred from short-term experiments, not a scheme anyone has tested directly.

References