Skip to content
Language through
Literature
The reading journalRussian · Language through Literature

Language through Literature

The short story is not the easy one


24 texts by Chekhov, Gogol, Turgenev, Dostoevsky and Tolstoy · measured each against its own vocabulary

Every teacher, every forum answer and every reading list gives the same first instruction: start with a short story. It is good advice for reasons that have nothing to do with vocabulary — a short text can be finished, and finishing is most of what keeps a reader reading. But it is usually offered as though the short text were also the easier one, word for word. Measured against the texts themselves, that part is false, and it is false in a direction worth knowing about.

The measurement

For each of 24 public-domain texts I asked one question: how many of that text's own dictionary words must a reader hold to reach 95% of its running words — the coverage figure at which Hu & Nation put comfortable unassisted reading? The words are taken in the only order a learner could actually follow, commonest-in-that-text first. Proper names are counted as covered rather than as vocabulary; nobody looks up Очумелов.

The answer is not a constant, and it does not move the way one would expect.

TextAuthor Running words Distinct words Appear once Needed for 95% Share of its own vocabulary
Толстый и тонкий Chekhov 511 268 70.5% 243 90.7%
Смерть чиновника Chekhov 677 317 64.4% 284 89.6%
Хамелеон Chekhov 901 403 67.5% 358 88.8%
Лошадиная фамилия Chekhov 904 380 69.5% 335 88.2%
Злоумышленник Chekhov 1,063 421 63.9% 368 87.4%
Студент Chekhov 1,124 498 69.9% 442 88.8%
Ванька Chekhov 1,153 556 71.9% 499 89.7%
Спать хочется Chekhov 1,570 611 64.6% 533 87.2%
Мальчики Chekhov 1,791 696 66.2% 607 87.2%
Скрипка Ротшильда Chekhov 3,017 1,014 62.4% 864 85.2%
О любви Chekhov 3,056 963 64.3% 811 84.2%
После бала Tolstoy 3,075 966 64.5% 813 84.2%
Крыжовник Chekhov 3,241 1,118 64.8% 956 85.5%
Медведь Chekhov 3,647 1,025 59.5% 843 82.2%
Душечка Chekhov 3,850 1,201 62.5% 1,009 84.0%
Анна на шее Chekhov 3,935 1,294 64.7% 1,098 84.9%
Человек в футляре Chekhov 4,005 1,225 65.0% 1,025 83.7%
Дама с собачкой Chekhov 5,109 1,405 61.7% 1,150 81.9%
Дом с мезонином Chekhov 5,600 1,542 60.6% 1,262 81.8%
Нос Gogol 7,645 1,893 60.2% 1,511 79.8%
Шинель Gogol 10,007 2,313 60.5% 1,813 78.4%
Первая любовь Turgenev 15,781 3,215 56.6% 2,426 75.5%
Палата № 6 Chekhov 16,043 3,339 55.3% 2,537 76.0%
Белые ночи Dostoevsky 16,845 2,513 55.0% 1,671 66.5%

Read the last column from the top down. Толстый и тонкий is 511 words long and asks for 90.7% of its own vocabulary. Белые ночи is 16,845 words long and asks for 66.5%. The shortest text on the shelf is the least forgiving one on it, by 24.2 percentage points.

The trend is not perfectly monotonic — a few short texts are denser than their length predicts — but the shape is not in doubt. All 9 texts under two thousand words sit in a narrow band between 87.2% and 90.7%. The fall only begins after that, and it keeps going as long as the texts do.

Why length is forgiving

Coverage is bought with repetition, and repetition needs room. In a long text a few hundred words — the prepositions, the pronouns, the handful of nouns the story is actually about — recur so often that they carry an enormous share of every page, and a reader who holds only those is already most of the way there. A short story has nowhere to put that repetition. 70.5% of the vocabulary of Толстый и тонкий occurs exactly once inside it. There is no frequency core to stand on: the weight is spread almost flat across the whole vocabulary, so nearly every word you do not know costs you something.

The same effect separates two texts of almost identical length near the bottom of the table. Белые ночи runs 6.7 words for every dictionary word it uses; Палата № 6 runs 4.8. One narrator circles the same small vocabulary for pages at a time and the other does not, and the difference in what they demand of a reader is 66.5% against 76.0% — a wider gap than the one between the shortest text here and the longest of the short ones. Length is a proxy for repetition. It is not the thing itself.

The objection, which is correct

There is an honest rebuttal to all of this and it should be made in full: in absolute terms the short story really is easier. Five per cent of Толстый и тонкий is about 26 unknown words in total; five per cent of Белые ночи is about 842. And the list you would have to learn is 243 words against 1,671 — the difference between an afternoon and a season.

Both things are true, because they are answers to two different questions. How much work is this text has the answer everyone assumes: the short one, easily. How complete must my knowledge be before this text stops interrupting me has the opposite answer. A learner who picks a short story expecting the second kind of ease has been told something true in a way that will mislead them, and will conclude from being stopped every third line that they are not ready to read — when what has actually happened is that they chose the format with the least slack in it.

What follows

The last tenth is the expensive tenth, and it cannot be pre-learned. Going from 95% to 98% on Толстый и тонкий means 258 of its 268 words rather than 243 — 96.3% of everything in it. Those last words are the ones that occur once, which is to say the ones no frequency list will ever hand you in time, in this text or in any other.

So the choice is not which text to postpone. If the words that stop you are mostly words you would meet once anyway, the useful move is not to go and learn them first — it is to read with them already attached, and spend the attention on the half of the vocabulary that does recur, in the order the books themselves impose.

And pick short texts for the right reason. Not because they are gentle, which they are not, but because the whole ask is small and visible: 243 words is a list you can see the end of, and at the end of it is a finished story by Chekhov rather than a percentage.

Method

Public-domain texts with lemmatisation corrected through dictionary checks and reader feedback. How the reader is made explains the process and its limits. Tokenisation and lemmatisation are the ones the reader itself runs, so these figures and the ones on the word census are computed the same way and can be held against each other. Proper names are excluded from the vocabulary counts and counted as covered running text. Every number on this page is computed from the shipped corpus when the page is served rather than typed in, so it cannot drift away from the texts it describes. The underlying word lists are published, the full ledger — every book's counts in one table, with a file — is at Russian classics, counted, and the corpus itself is available to anyone who wants to check this — write and ask.

On the 98% figure. The 95–98% coverage band this page measures against is Paul Nation's. I wrote to him with the methodological question behind these numbers, he answered it, and he has given permission to be named here. The band is his; the measurements are mine, and so is any error in them.

The texts, with every word carrying its stress mark and its meaning where it stands, are at languagethroughliterature.com. The seven Chekhov stories in the table above are free to read with no account — Толстый и тонкий is here.

Some complete works are free without an account in the interactive reader; other books offer free openings. Free lesson samples are also available. A free account keeps your saved words and progress across devices. Free reading and membership.