Benchmarks
Available benchmark suites
| Name | Questions | Runs | Avg Score | Last Run | Actions |
|---|---|---|---|---|---|
|
Algebra
A benchmark to evaluate a model's ability to solve linear and quadratic equations with integer solutions. Includes single-variable linear equations (ax + b = c) and quadratic equations with one or two integer roots. |
1120 | 1120 | 62.0/100 | 2026-07-27 17:19 | View Run |
|
Antonym Identification
A benchmark to evaluate a model's ability to identify the antonym of a word. |
1120 | 1120 | 89.5/100 | 2026-07-27 17:13 | View Run |
|
Book Author Match
A benchmark to evaluate matching famous books to their correct authors. |
432 | 432 | 77.0/100 | 2026-07-27 17:36 | View Run |
|
Definitions
A benchmark to evaluate a model's ability to identify the correct definition of words. |
920 | 920 | 83.7/100 | 2026-07-27 17:33 | View Run |
|
English Plural Generation
A benchmark to evaluate a model's ability to produce the correct plural form of English nouns, covering regular, -es, -ies, -ves, irregular, invariant, and Latin/Greek pluralization rules. |
1000 | 1000 | 92.3/100 | 2026-07-27 17:34 | View Run |
|
Food Category Classification
A benchmark to evaluate classification of food items by category. |
500 | 500 | 78.4/100 | 2026-07-27 17:36 | View Run |
|
Fractions and Percentages
A benchmark to evaluate a model's ability to calculate percentages and fractions, including percent-of, fraction-of, and percent change problems. |
1120 | 1120 | 87.6/100 | 2026-07-27 17:18 | View Run |
|
Geography Knowledge
A benchmark to evaluate a model's knowledge of world geography through multiple-choice questions about countries, capitals, physical features, and other geographical information. |
1000 | 1000 | 84.5/100 | 2026-07-27 17:35 | View Run |
|
Geometry
A benchmark to evaluate a model's ability to calculate area, perimeter, and volume for standard shapes: rectangles, triangles, rectangular boxes, and circles (using π ≈ 3.14159). |
1120 | 1120 | 76.5/100 | 2026-07-27 17:20 | View Run |
|
Historical Event Year
A benchmark to evaluate selecting the correct year for major historical events. |
450 | 450 | 71.7/100 | 2026-07-27 17:37 | View Run |
|
Lemma Identification
A benchmark to evaluate a model's ability to identify the lemma (base form) of a given word. The lemma is the dictionary form: - For nouns: the singular form (e.g., "cats" → "cat") - For verbs: the infinitive form without "to" (e.g., "running" → "run") - For adjectives: the positive form (e.g., "better" → "good") No questions — generate first |
0 | 0 | - | Never | View Generate |
|
Letter Count
A benchmark to evaluate a model's ability to count how many times a specific letter appears in a word. |
1120 | 1120 | 39.8/100 | 2026-07-27 17:09 | View Run |
|
Math Word Problems
A benchmark to evaluate a model's ability to read math word problems and extract the relevant numbers to compute the correct answer. Approximately one third of questions contain distractor/unused information. |
1120 | 1120 | 86.1/100 | 2026-07-27 17:17 | View Run |
|
Multilingual Synonym Generation
A benchmark to evaluate a model's ability to generate noun synonyms in multiple languages. |
1456 | 1456 | 76.9/100 | 2026-07-27 17:14 | View Run |
|
Part of Speech
A benchmark to evaluate a model's ability to identify the part of speech of a specific word in a sentence. |
1000 | 1000 | 90.3/100 | 2026-07-27 17:33 | View Run |
|
Pinyin Letter Count
A benchmark to evaluate a model's ability to count how many times a specific letter appears in the Pinyin representation of a Chinese sentence. |
560 | 560 | 23.0/100 | 2026-07-27 17:15 | View Run |
|
Python GCD With Validation
Write a Python 3.12 function for GCD with invalid-input exceptions. |
24 | 24 | 81.7/100 | 2026-07-27 17:37 | View Run |
|
Python Hello World Function
Write a Python 3.12 function that prints Hello world. |
24 | 24 | 91.7/100 | 2026-07-27 17:37 | View Run |
|
Python Letter Count in String
Count occurrences of a target letter in a string. |
23 | 23 | 81.7/100 | 2026-07-27 17:37 | View Run |
|
Python Minimum Coin Change
Compute minimum number of coins to make a target amount. |
23 | 23 | 74.4/100 | 2026-07-27 17:37 | View Run |
|
Python Prime Factorization
Return the prime factorization of a positive integer. |
23 | 23 | 72.2/100 | 2026-07-27 17:37 | View Run |
|
Sentence Decomposition
A benchmark to evaluate a model's ability to produce multilingual token-level sentence decomposition with grammatical metadata. No questions — generate first |
0 | 0 | - | Never | View Generate |
|
Simple Arithmetic
A benchmark to evaluate a model's ability to perform basic arithmetic: addition, subtraction, multiplication, and division. |
1120 | 1120 | 94.2/100 | 2026-07-27 17:15 | View Run |
|
Spell Check
A benchmark to evaluate a model's ability to identify misspelled words in a sentence and provide their correct spelling. |
1120 | 1120 | 83.2/100 | 2026-07-27 17:12 | View Run |
|
Syllable Count
Tests ability to count syllables in words across Latin-alphabet languages. |
1120 | 1120 | 54.5/100 | 2026-07-27 17:11 | View Run |
|
Syllogism Validity
A benchmark to evaluate whether a model can determine if short categorical syllogisms are logically valid. |
384 | 384 | 68.7/100 | 2026-07-27 17:35 | View Run |
|
Time Arithmetic
A benchmark to evaluate a model's ability to add and subtract durations from clock times in 24-hour HH:MM format. |
1120 | 1120 | 60.1/100 | 2026-07-27 17:19 | View Run |
|
Translation en_de
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 80.8/100 | 2026-07-27 17:23 | View Run |
|
Translation en_es
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 81.9/100 | 2026-07-27 17:22 | View Run |
|
Translation en_fr
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 79.9/100 | 2026-07-27 17:21 | View Run |
|
Translation en_ja
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 77.7/100 | 2026-07-27 17:25 | View Run |
|
Translation en_zh
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 81.7/100 | 2026-07-27 17:24 | View Run |
|
Translation fr_es
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 74.0/100 | 2026-07-27 17:24 | View Run |
|
Translation fr_ko
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 74.6/100 | 2026-07-27 17:26 | View Run |
|
Translation it_lt
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 64.4/100 | 2026-07-27 17:27 | View Run |
|
Translation ja_lt
A benchmark to evaluate a model's ability to translate words from one language to another. |
1120 | 1120 | 71.1/100 | 2026-07-27 17:27 | View Run |
|
Unit Conversion
A benchmark to evaluate a model's ability to accurately convert between different units of measurement. |
1120 | 1120 | 73.1/100 | 2026-07-27 17:16 | View Run |
|
Validate Bulk IPA/Phonetic (bebras)
A regression benchmark for Bebras bulk pronunciation verification. Tests whether the model returns only words with wrong IPA/phonetic mappings from 20-word lists with English + Chinese disambiguation. |
5 | 5 | 20.0/100 | 2026-04-13 20:24 | View Run |
|
Validate Definition (lokys)
A regression benchmark for the lokys agent's validate_definition() function. Tests whether the LLM correctly identifies well-formed vs. problematic word definitions (e.g. circular definitions, translations used as definitions). No questions — generate first |
0 | 0 | - | Never | View Generate |
|
Validate Lemma Form (lokys)
A regression benchmark for the lokys agent's validate_lemma_form() function. Tests whether the LLM correctly identifies if a word is in its base/lemma form and suggests the correct form when it is not. No questions — generate first |
0 | 0 | - | Never | View Generate |
|
Validate Translation (voras)
A regression benchmark for the voras agent's validate_all_translations_for_word() function. Tests whether the LLM correctly identifies semantically incorrect or non-lemma translations across multiple target languages. No questions — generate first |
0 | 0 | - | Never | View Generate |
|
Verb Forms
A benchmark to evaluate a model's ability to generate full verb-form paradigms across persons and tenses in multiple languages. No questions — generate first |
0 | 0 | - | Never | View Generate |
|
Vowel Count
Tests ability to count vowels (a, e, i, o, u and accented forms) in a word across Latin-alphabet languages. |
1120 | 1120 | 44.8/100 | 2026-07-27 17:10 | View Run |
|
Word Length
A benchmark to evaluate a model's ability to count the total number of letters in a given word. |
1120 | 1120 | 64.8/100 | 2026-07-27 17:08 | View Run |
|
Word to IPA
A benchmark to evaluate a model's ability to convert words from multiple languages to their IPA (International Phonetic Alphabet) pronunciation. |
80 | 80 | 78.5/100 | 2026-03-24 20:55 | View Run |