Language through Literature

The short story is not the easy one


Fifteen texts by Chekhov, Gogol, Turgenev, Dostoevsky and Tolstoy · measured each against its own vocabulary

Every teacher, every forum answer and every reading list gives the same first instruction: start with a short story. It is good advice for reasons that have nothing to do with vocabulary — a short text can be finished, and finishing is most of what keeps a reader reading. But it is usually offered as though the short text were also the easier one, word for word. Measured against the texts themselves, that part is false, and it is false in a direction worth knowing about.

The measurement

For each of fifteen public-domain texts I asked one question: how many of that text's own dictionary words must a reader hold to reach 95% of its running words — the coverage figure at which Hu & Nation put comfortable unassisted reading? The words are taken in the only order a learner could actually follow, commonest-in-that-text first. Proper names are counted as covered rather than as vocabulary; nobody looks up Очумелов.

The answer is not a constant, and it does not move the way one would expect.

TextAuthor Running words Distinct words Appear once Needed for 95% Share of its own vocabulary
Толстый и тонкий Chekhov 511 269 70.6% 244 90.7%
Смерть чиновника Chekhov 677 317 64.4% 284 89.6%
Хамелеон Chekhov 901 403 67.5% 358 88.8%
Лошадиная фамилия Chekhov 904 380 69.5% 335 88.2%
Злоумышленник Chekhov 1,063 421 63.9% 368 87.4%
Студент Chekhov 1,124 498 69.9% 442 88.8%
Ванька Chekhov 1,153 557 71.8% 500 89.8%
Спать хочется Chekhov 1,570 611 64.6% 533 87.2%
Мальчики Chekhov 1,791 696 66.2% 607 87.2%
После бала Tolstoy 3,075 966 64.5% 813 84.2%
Дама с собачкой Chekhov 5,109 1,405 61.7% 1,150 81.9%
Нос Gogol 7,645 1,893 60.2% 1,511 79.8%
Шинель Gogol 10,007 2,314 60.5% 1,814 78.4%
Первая любовь Turgenev 15,781 3,216 56.7% 2,427 75.5%
Белые ночи Dostoevsky 16,861 2,521 55.1% 1,678 66.6%

Read the last column from the top down. Толстый и тонкий is 511 words long and asks for 90.7% of its own vocabulary. Белые ночи is 16,861 words long and asks for 66.6%. The shortest text on the shelf is the least forgiving one on it, by 24.1 percentage points.

The trend is not perfectly monotonic — a few short texts are denser than their length predicts — but the shape is not in doubt. All 9 texts under two thousand words sit in a narrow band between 87.2% and 90.7%. The fall only begins after that, and it keeps going as long as the texts do.

Why length is forgiving

Coverage is bought with repetition, and repetition needs room. In a long text a few hundred words — the prepositions, the pronouns, the handful of nouns the story is actually about — recur so often that they carry an enormous share of every page, and a reader who holds only those is already most of the way there. A short story has nowhere to put that repetition. 70.6% of the vocabulary of Толстый и тонкий occurs exactly once inside it. There is no frequency core to stand on: the weight is spread almost flat across the whole vocabulary, so nearly every word you do not know costs you something.

The same effect separates two texts of almost identical length near the bottom of the table. Белые ночи runs 6.7 words for every dictionary word it uses; Первая любовь runs 4.9. One narrator circles the same small vocabulary for pages at a time and the other does not, and the difference in what they demand of a reader is 66.6% against 75.5% — a wider gap than the one between the shortest text here and the longest of the short ones. Length is a proxy for repetition. It is not the thing itself.

The objection, which is correct

There is an honest rebuttal to all of this and it should be made in full: in absolute terms the short story really is easier. Five per cent of Толстый и тонкий is about 26 unknown words in total; five per cent of Белые ночи is about 843. And the list you would have to learn is 244 words against 1,678 — the difference between an afternoon and a season.

Both things are true, because they are answers to two different questions. How much work is this text has the answer everyone assumes: the short one, easily. How complete must my knowledge be before this text stops interrupting me has the opposite answer. A learner who picks a short story expecting the second kind of ease has been told something true in a way that will mislead them, and will conclude from being stopped every third line that they are not ready to read — when what has actually happened is that they chose the format with the least slack in it.

What follows

The last tenth is the expensive tenth, and it cannot be pre-learned. Going from 95% to 98% on Толстый и тонкий means 259 of its 269 words rather than 244 — 96.3% of everything in it. Those last words are the ones that occur once, which is to say the ones no frequency list will ever hand you in time, in this text or in any other.

So the choice is not which text to postpone. If the words that stop you are mostly words you would meet once anyway, the useful move is not to go and learn them first — it is to read with them already attached, and spend the attention on the half of the vocabulary that does recur, in the order the books themselves impose.

And pick short texts for the right reason. Not because they are gentle, which they are not, but because the whole ask is small and visible: 244 words is a list you can see the end of, and at the end of it is a finished story by Chekhov rather than a percentage.

Method

Fifteen public-domain texts, lemmatised by hand and checked against dictionaries. Tokenisation and lemmatisation are the ones the reader itself runs, so these figures and the ones on the word census are computed the same way and can be held against each other. Proper names are excluded from the vocabulary counts and counted as covered running text. Every number on this page is computed from the shipped corpus when the page is served rather than typed in, so it cannot drift away from the texts it describes. The underlying word lists are published, the full ledger — every book's counts in one table, with a file — is at Russian classics, counted, and the corpus itself is available to anyone who wants to check this — write and ask.

The texts, with every word carrying its stress mark and its meaning where it stands, are at languagethroughliterature.com. The seven Chekhov stories in the table above are free to read with no account — Толстый и тонкий is here.