Word Tokens

A word token is one spelling in one language — “top”, “will”, “ice cream” — regardless of which sense means it. Frequency is measured here, then split across the senses that share the string.

Top unlinked tokens

Frequent spellings that no lemma claims yet — the queue of vocabulary the corpora measured but the database has no sense for.

Most common unlinked

Corpus vocabulary

Words that lean toward one corpus more than the rest — the cooking words, the science words.


Register-neutral words

The complement: words that sit at the same frequency in every corpus.

Steadiest across corpora