Word tokens
Words characteristic of wiki_linguistics
Word frequency data from Wikipedia Level 5 language articles
Ranked by Zipf delta: how much more common a word is here than its
average across the 14 other corpora
(19th_books, 20th_books, wiki_math, wiki_geography, wiki_biology, wiki_modern_life, wiki_arts, wiki_society, wiki_physical_science, wiki_history, early_modern_science, religious_translated, cooking, legal_scotus).
+1.00 means ten times as common here. The ranks beside it say the same thing
in readable terms — they are not used for sorting, because the corpora are
different sizes and rank 500 does not mean the same thing in each.
Most skewed toward wiki_linguistics
100 shown
| Word | Zipf Δ | Rank here | Mean rank elsewhere | Rank ratio | Corpora |
|---|---|---|---|---|---|
| languages | +2.00 | 16 | 2371 | 148.2× | 12 |
| tense | +1.88 | 179 | 7046 | 39.4× | 3 |
| pronunciation | +1.80 | 143 | 6410 | 44.8× | 3 |
| dialect | +1.80 | 68 | 4650 | 68.4× | 4 |
| script | +1.73 | 65 | 4080 | 62.8× | 7 |
| language | +1.71 | 10 | 822 | 82.2× | 14 |
| verb | +1.70 | 70 | 4149 | 59.3× | 3 |
| alphabet | +1.69 | 83 | 4589 | 55.3× | 4 |
| vocabulary | +1.68 | 163 | 5622 | 34.5× | 3 |
| syllable | +1.64 | 130 | 4623 | 35.6× | 3 |
| spelling | +1.63 | 202 | 6105 | 30.2× | 5 |
| grammar | +1.61 | 131 | 4426 | 33.8× | 5 |
| noun | +1.60 | 105 | 3522 | 33.5× | 3 |
| speakers | +1.58 | 63 | 2703 | 42.9× | 4 |
| dialects | +1.55 | 57 | 2215 | 38.9× | 2 |
| linguistics | +1.54 | 168 | 4460 | 26.6× | 3 |
| spoken | +1.54 | 43 | 1959 | 45.6× | 10 |
| pronounced | +1.51 | 125 | 3669 | 29.3× | 11 |
| linguistic | +1.42 | 135 | 3624 | 26.8× | 6 |
| masculine | +1.41 | 507 | 6229 | 12.3× | 3 |
| slang | +1.38 | 797 | 8044 | 10.1× | 3 |
| voiced | +1.36 | 359 | 5713 | 15.9× | 3 |
| plural | +1.36 | 161 | 3869 | 24.0× | 4 |
| Language | +1.33 | 162 | 3564 | 22.0× | 5 |
| indicative | +1.31 | 1122 | 8995 | 8.0× | 2 |
| auxiliary | +1.29 | 665 | 6470 | 9.7× | 3 |
| accents | +1.26 | 712 | 6317 | 8.9× | 4 |
| singular | +1.25 | 204 | 3340 | 16.4× | 8 |
| imperative | +1.25 | 880 | 7396 | 8.4× | 3 |
| varieties | +1.23 | 145 | 3391 | 23.4× | 11 |
| letters | +1.22 | 81 | 1728 | 21.3× | 11 |
| accent | +1.22 | 361 | 3652 | 10.1× | 3 |
| sounds | +1.21 | 160 | 2724 | 17.0× | 9 |
| stops | +1.20 | 504 | 5603 | 11.1× | 8 |
| borrowed | +1.20 | 495 | 5128 | 10.4× | 7 |
| words | +1.20 | 33 | 784 | 23.7× | 14 |
| scripts | +1.19 | 336 | 4880 | 14.5× | 2 |
| Standard | +1.17 | 245 | 3566 | 14.6× | 6 |
| speaker | +1.17 | 430 | 4557 | 10.6× | 7 |
| ch | +1.17 | 714 | 6332 | 8.9× | 5 |
| suffix | +1.16 | 285 | 3945 | 13.8× | 2 |
| phrases | +1.15 | 368 | 3985 | 10.8× | 6 |
| writing | +1.15 | 84 | 1636 | 19.5× | 13 |
| characters | +1.14 | 129 | 2137 | 16.6× | 11 |
| dictionary | +1.13 | 509 | 5456 | 10.7× | 4 |
| emphatic | +1.12 | 1630 | 8877 | 5.4× | 2 |
| Modern | +1.11 | 290 | 3825 | 13.2× | 9 |
| ending | +1.10 | 374 | 4065 | 10.9× | 11 |
| borrowing | +1.10 | 1384 | 7909 | 5.7× | 2 |
| syntactic | +1.10 | 547 | 3742 | 6.8× | 2 |
| native | +1.10 | 121 | 1916 | 15.8× | 13 |
| syllables | +1.09 | 215 | 2812 | 13.1× | 2 |
| comparative | +1.09 | 638 | 5360 | 8.4× | 6 |
| feminine | +1.08 | 503 | 4384 | 8.7× | 5 |
| speech | +1.08 | 101 | 1707 | 16.9× | 11 |
| adjective | +1.08 | 436 | 3087 | 7.1× | 2 |
| verbal | +1.08 | 516 | 5077 | 9.8× | 4 |
| linguistically | +1.07 | 1922 | 9191 | 4.8× | 2 |
| standard | +1.07 | 89 | 1919 | 21.6× | 14 |
| Proto | +1.06 | 343 | 3976 | 11.6× | 3 |
| syntax | +1.06 | 467 | 3156 | 6.8× | 2 |
| sentence | +1.05 | 173 | 2233 | 12.9× | 9 |
| word | +1.05 | 35 | 534 | 15.3× | 15 |
| intelligible | +1.04 | 777 | 5627 | 7.2× | 5 |
| stress | +1.04 | 299 | 3486 | 11.7× | 10 |
| written | +1.04 | 55 | 837 | 15.2× | 14 |
| phrase | +1.03 | 259 | 2734 | 10.6× | 10 |
| official | +1.02 | 118 | 1701 | 14.4× | 12 |
| nd | +1.02 | 2164 | 9378 | 4.3× | 2 |
| tones | +1.01 | 410 | 3224 | 7.9× | 5 |
| stressed | +1.00 | 556 | 4655 | 8.4× | 4 |
| sentences | +1.00 | 423 | 3446 | 8.1× | 6 |
| Old | +1.00 | 224 | 2654 | 11.8× | 13 |
| literary | +0.99 | 229 | 2786 | 12.2× | 8 |
| indicate | +0.99 | 262 | 2835 | 10.8× | 14 |
| initials | +0.98 | 2059 | 8691 | 4.2× | 2 |
| register | +0.97 | 902 | 5747 | 6.4× | 6 |
| intonation | +0.97 | 1058 | 6859 | 6.5× | 3 |
| dictionaries | +0.97 | 719 | 5396 | 7.5× | 3 |
| te | +0.97 | 633 | 4177 | 6.6× | 5 |
| indefinite | +0.96 | 887 | 5120 | 5.8× | 6 |
| cf | +0.96 | 1270 | 6062 | 4.8× | 3 |
| mark | +0.94 | 268 | 2188 | 8.2× | 14 |
| proverb | +0.94 | 2385 | 9463 | 4.0× | 2 |
| marks | +0.94 | 351 | 2906 | 8.3× | 13 |
| marking | +0.93 | 884 | 5700 | 6.4× | 6 |
| merged | +0.93 | 893 | 5724 | 6.4× | 5 |
| ni | +0.93 | 1237 | 6468 | 5.2× | 3 |
| o | +0.93 | 208 | 2131 | 10.2× | 15 |
| hi | +0.92 | 2608 | 9600 | 3.7× | 2 |
| speaking | +0.91 | 185 | 1672 | 9.0× | 15 |
| semantic | +0.91 | 526 | 4132 | 7.9× | 2 |
| conjugation | +0.91 | 936 | 3938 | 4.2× | 2 |
| mutually | +0.90 | 709 | 4504 | 6.4× | 4 |
| letter | +0.89 | 139 | 1444 | 10.4× | 13 |
| i | +0.89 | 78 | 950 | 12.2× | 15 |
| example | +0.89 | 46 | 692 | 15.1× | 15 |
| clauses | +0.89 | 440 | 2856 | 6.5× | 2 |
| texts | +0.88 | 241 | 2672 | 11.1× | 9 |
| sh | +0.88 | 1349 | 6851 | 5.1× | 3 |