The items
Version 0.1.0 has 268 items. There are three kinds.
- Sealed sentences. One sentence in two images. The first is in the professor's own hand. The second is the same text rewritten in a period-style hand by a hieratic specialist, closer to what a scribe would have produced. Both were commissioned in 2022. The script is public. What the sentence says has never been published.
- The public identification set. Photographs and facsimiles of real Egyptian documents from Wikimedia Commons, The Metropolitan Museum of Art, Chester Beatty in Dublin and the Yale Peabody Museum. Most are hieratic. The rest are controls in demotic and hieroglyphs, so a model that answers hieratic for anything Egyptian loses points.
- Single signs. Hieratic signs from AKU-PAL, the Mainz Academy's palaeography database, each labelled with its Gardiner code by the database itself. A label is kept only when published palaeographies agree on it. Signs span the Old Kingdom to the Roman period.
Every public image is public domain, CC0, CC BY or CC BY-SA, and its record keeps the source page, license and attribution. Captions and labels that name the script are cropped out. Items from widely reproduced documents are flagged as famous, because a model may have seen that exact photograph in training.
The prompts
One image and one user message. No system prompt, no tools, no examples. Identify is asked cold. Every other rung tells the model the script, so failing rung one doesn't sink the rest. Each prompt asks for a final answer line, which is the only part that gets scored.
Identify
Can you identify and translate this script? Finish your answer with one final line in exactly this format: SCRIPT: <the name of the writing system, or UNKNOWN>
Signs, for a sentence
This image shows Egyptian hieratic, the cursive script of ancient Egypt, written with a reed brush or pen. Transcribe it sign by sign into hieroglyphs, in reading order, using codes from Gardiner's sign list (category letters followed by a number, no spaces inside a code). Finish your answer with one final line in exactly this format: SIGNS: <Gardiner codes separated by spaces>
Signs, for a single sign
This image shows Egyptian hieratic, the cursive script of ancient Egypt, written with a reed brush or pen. It contains a single sign. Which hieroglyph does this hieratic sign correspond to? Give its code in Gardiner's sign list (category letters followed by a number, no spaces inside a code). Finish your answer with one final line in exactly this format: SIGNS: <one Gardiner code>
Transliterate
This image shows Egyptian hieratic, the cursive script of ancient Egypt, written with a reed brush or pen. Transliterate the text into standard Egyptological transliteration, in Unicode (ꜣ, ꜥ, ḥ, ḫ, ẖ, š, ḳ, ṯ, ḏ and so on). Finish your answer with one final line in exactly this format: TRANSLITERATION: <your transliteration>
Translate
This image shows Egyptian hieratic, the cursive script of ancient Egypt, written with a reed brush or pen. Translate the text into English. Finish your answer with one final line in exactly this format: TRANSLATION: <your English translation>
Scoring
- Identify. 1 for the right script, 0.5 for a different Egyptian script, 0 for anything else, including UNKNOWN or a missing answer line. Abnormal hieratic counts as hieratic and cursive hieroglyphs count as hieroglyphic.
- Signs. A single sign scores 1 when the first code given matches the source's label. A sentence scores one minus the sign error rate, the edit distance between the predicted and reference code sequences divided by the reference length, floored at zero.
- Transliterate and translate. On the sentence these are not scored by machine, because no answer key is stored. The harness still ships character error rate and chrF scoring for future public items with published readings.
- Averages. Samples are averaged per item, then items are averaged per rung.
A refusal scores zero. A failed API call is left out and re-run. Server-side model fallbacks are switched off, so an answer is always credited to the model that wrote it.
Keeping the answer sealed
- There is no answer key. Not in the repository, not on a server, not with the maintainer. It has never been published and never will be.
- When anyone runs a sealed rung, the answers are written to a local inbox that is never committed either. A correct reading would give the answer away, so these stay private.
- The identify question on the sealed sentences also asks for a translation, so a model that can read it would write the answer there. For those items only the script the model named and its score are published. The full text stays private.
Settings
Models that take a reasoning effort run at high. Nothing else is tuned, no temperature is set, and the provider's defaults apply. Leaderboard runs ask each sealed item three times per rung, because there are only two of them, and each public item once, because there are many. The same model at two effort levels is two rows.
Run it yourself
git clone https://github.com/alymoursy/hieraticbench cd hieraticbench && npm install cp .env.example .env # add one provider key npm run bench -- run --models claude-opus-5-5 --samples 3 npm run bench -- leaderboard
Any model can run as provider:model-id, for example openrouter:qwen/qwen3.8-flash. Claude runs through Anthropic's API, other labs' models through OpenRouter. Scored results land in results/runs. Answers about the sentence land in results/inbox. Keep that file private. It is never committed. Questions go to hieraticbench@veeza.ai.
Limits of version 0.1.0
- The sealed set is one sentence in two hands. That is a story, not yet a statistic. More sentences from more hands is the first priority.
- The two sealed images are crops of screenshots. They will be replaced with the original scans.
- Famous public documents may be in training data. The famous flag lets you score with or without them.
- To keep costs down, models from other labs answered only the identify question. GPT-6 Astra and both Gemini models were asked about the sentence alone, and Qwen3.8 Max stopped partway through the documents. The leaderboard marks every gap.
- Chat-app transcripts on the home page are context. They use each app's own system prompt and are never counted on the leaderboard.
What comes next
- New sealed sentences from Egyptologists, in different hands and periods.
- A tool-assisted track where the model can crop and zoom, since the newest models read images better with a crop tool.
- Line-by-line transcriptions of public papyri, annotated sign by sign by specialists.