Russian classics, counted
This page is the ledger: the numbers this house is asked for, measured on the books themselves and printed where anyone can check them. Nothing here is estimated, averaged from elsewhere, or typed in by hand — every figure is computed from the same data files the reader is served, and moves when the corpus moves. The same table is a file, below.
The corpus, whole
68,172 running words across 5,075 counted sentences of fifteen texts by Chekhov, Gogol, Turgenev, Dostoevsky and Tolstoy, containing 7,746 distinct lemmas — of which 3,824, 49.4%, occur exactly once. The vocabulary is ordered into a ladder of 7,000 rungs, each rung the word that opens the most whole sentences not yet readable; here is what each depth of it buys:
| Words known | Coverage of running text | Sentences fully readable | Share of all sentences |
|---|---|---|---|
| 100 | 39.4% | 405 | 8.0% |
| 250 | 55.7% | 751 | 14.8% |
| 500 | 67.3% | 1,195 | 23.5% |
| 1000 | 77.1% | 1,802 | 35.5% |
| 1500 | 80.2% | 2,309 | 45.5% |
| 2000 | 82.8% | 2,778 | 54.7% |
| 3000 | 89.8% | 3,398 | 67.0% |
| 4000 | 93.4% | 3,841 | 75.7% |
| 5000 | 95.9% | 4,221 | 83.2% |
| 6000 | 97.4% | 4,521 | 89.1% |
| 7000 | 98.9% | 4,851 | 95.6% |
The books, one by one
Two ways of asking what a book costs. Ladder to 95 / 98 is how many words of the ladder, first rung first, before that book's running text is 95% and 98% covered — the thresholds reading research treats as workable and comfortable. Readable at 1,000 is the share of its sentences in which the first thousand words leave no stranger at all. An em-dash means the whole ladder never gets there: that book's tail is vocabulary the rest of the shelf never uses.
| Book | Author | Words | Lemmas | Appear once | Ladder to 95 | Ladder to 98 | Readable at 1,000 |
|---|---|---|---|---|---|---|---|
| Толстый и тонкий free | А. П. Чехов | 511 | 269 | 70.6% | 4,994 | 5,009 | 39.5% |
| Смерть чиновника free | А. П. Чехов | 677 | 317 | 64.4% | 3,822 | 4,951 | 50.5% |
| Хамелеон free | А. П. Чехов | 901 | 403 | 67.5% | 3,919 | 5,489 | 44.0% |
| Лошадиная фамилия free | А. П. Чехов | 904 | 380 | 69.5% | 3,850 | 4,973 | 52.2% |
| Злоумышленник free | А. П. Чехов | 1,063 | 421 | 63.9% | 6,472 | 6,504 | 46.5% |
| Студент free | А. П. Чехов | 1,124 | 498 | 69.9% | 6,932 | 6,966 | 20.6% |
| Ванька free | А. П. Чехов | 1,153 | 557 | 71.8% | — | — | 8.2% |
| Спать хочется | А. П. Чехов | 1,570 | 611 | 64.6% | 4,682 | 6,549 | 28.4% |
| Мальчики | А. П. Чехов | 1,791 | 696 | 66.2% | 4,367 | 5,737 | 33.5% |
| После бала | Л. Н. Толстой | 3,075 | 966 | 64.5% | 4,945 | — | 26.2% |
| Дама с собачкой | А. П. Чехов | 5,109 | 1,405 | 61.7% | 4,306 | 5,614 | 33.9% |
| Нос | Н. В. Гоголь | 7,645 | 1,893 | 60.2% | 4,785 | 6,769 | 28.4% |
| Шинель | Н. В. Гоголь | 10,007 | 2,314 | 60.5% | — | — | 12.7% |
| Первая любовь | И. С. Тургенев | 15,781 | 3,216 | 56.7% | 4,576 | 6,150 | 33.4% |
| Белые ночи | Ф. М. Достоевский | 16,861 | 2,521 | 55.1% | 3,465 | 5,155 | 49.0% |
The same rows, said in sentences
«Толстый и тонкий» asks 4,994 lemmas of you before it reads at 95 in a hundred, and 5,009 before 98.
«Смерть чиновника» asks 3,822 lemmas of you before it reads at 95 in a hundred, and 4,951 before 98.
«Хамелеон» asks 3,919 lemmas of you before it reads at 95 in a hundred, and 5,489 before 98.
«Лошадиная фамилия» asks 3,850 lemmas of you before it reads at 95 in a hundred, and 4,973 before 98.
«Злоумышленник» asks 6,472 lemmas of you before it reads at 95 in a hundred, and 6,504 before 98.
«Студент» asks 6,932 lemmas of you before it reads at 95 in a hundred, and 6,966 before 98.
«Ванька» is not read down by the ladder alone: all seven thousand rungs leave it at 93.0 in a hundred, and the rest of its vocabulary is its own.
«Спать хочется» asks 4,682 lemmas of you before it reads at 95 in a hundred, and 6,549 before 98.
«Мальчики» asks 4,367 lemmas of you before it reads at 95 in a hundred, and 5,737 before 98.
«После бала» asks 4,945 lemmas of you to read at 95 in a hundred; 98 lies beyond the ladder, in words the rest of the shelf never uses.
«Дама с собачкой» asks 4,306 lemmas of you before it reads at 95 in a hundred, and 5,614 before 98.
«Нос» asks 4,785 lemmas of you before it reads at 95 in a hundred, and 6,769 before 98.
«Шинель» is not read down by the ladder alone: all seven thousand rungs leave it at 94.8 in a hundred, and the rest of its vocabulary is its own.
«Первая любовь» asks 4,576 lemmas of you before it reads at 95 in a hundred, and 6,150 before 98.
«Белые ночи» asks 3,465 lemmas of you before it reads at 95 in a hundred, and 5,155 before 98.
Can you read one of them yet?
The table says what each book asks; whether you hold it is a different question, and it has its own page. Mark forty words you know and the corpus answers with its own arithmetic — can you read this book yet?
The file
russian-classics-counted.csv — one row per book, the columns above, openable in any spreadsheet. The word lists that pair with it are published separately, with stress marks and meanings.
Method
Mined, never made. The corpus is fifteen public-domain texts, entire and unabridged — no invented sentences, no simplified retellings, and if an adapted edition ever joins the shelf it will be labelled a retelling and excluded from every row of this page and its file. Lemmatisation and glossing were done by hand and checked against dictionaries. Tokenisation is the build's own, the same one the reader runs. Proper names are counted as covered running text, never as vocabulary to be learned. A sentence enters the sentence counts when it contains at least one dictionary word; a line of bare names or punctuation measures nothing. A sentence is fully readable at a given depth when every dictionary word in it stands within that many rungs of the ladder — the ordering in which each next word opens the most whole sentences, published as the word lists. The 95/98 thresholds are the working heuristic of the reading research (Hu & Nation); the argument for them is set out in the guide, and the two findings drawn from this table stand at half the words appear once and the short story is not the easy one.
The German room keeps its own ledger, in its own terms — Grimm, counted.