Sber Releases MERA Text 2.0 Benchmark for Language Models
Sber introduced MERA Text 2.0 on August 31, restructuring its text-language-model benchmark around 12 tests in four skill groups: Human-Centric, Culture-Specific, Agentic and Reasoning. The previous version’s public leaderboard has been frozen and will no longer receive updates.
The MERA team said the redesign responds to saturation in older tasks. Leading models now approach the human baseline on the previous benchmark and exceed it on some tests, compressing differences near the top of the ranking. Rather than adding a few datasets, the team rebuilt the evaluation around abilities it considers difficult and informative for current models.
Human-Centric tasks assess riddles, mechanisms of humor and whether answers fit a described character. Culture-Specific testing covers Russian regional vocabulary, correction of real Russian-language text and cultural references spanning folklore, sayings, memes, songs and films. The team reports that this cultural group has the greatest room for improvement and correlates less strongly with overall model performance than the other groups.
The Agentic group does not attempt to evaluate a complete autonomous agent. IFHardBench requires models to satisfy three to six formal constraints simultaneously, while SOBHard tests document extraction and transformation. GorillaHard asks a model to select tools and supply appropriate call arguments, but measures planning rather than actual tool execution.
Reasoning tests cover words whose meanings reverse with context, short problems containing logical traps, and linguistic ambiguity. All MERA Text 2.0 test sets are private, and each test uses at least five instruction variants. The organizers also fix task protocols, retain run parameters and logs, and avoid adding agentic, reasoning or tool harnesses that could solve parts of the tasks for the evaluated model.
Scores are aggregated in three stages: metrics within each test, tests within each group, and then the four group scores. Every stage uses a geometric mean. A zero is replaced with 0.01 so that one failed metric does not erase the entire score, and valid comparisons require all models to be recalculated on the same test-set version.
Practical context: The grouped design can reveal whether a model’s weakness lies in cultural context, strict output constraints or tool selection instead of reducing every result to an undifferentiated average. Its scope remains narrower than full agent evaluation: long-running scenarios, environment interaction and complex end-to-end workflows are deliberately excluded from this benchmark line.
A full run has been reduced from 23 datasets and about 37,000 requests to 12 datasets and roughly 8,700 requests—a 4.3-fold decrease by request count. Previous submissions and results remain available in user accounts, while new measurements follow the existing submission process. At announcement time, the technical report was still forthcoming; the team directs readers to the live benchmark for current results rather than a static ranking.
| Group | Tests | What they assess |
|---|---|---|
| Human-Centric | Riddles, Humor, Characters | Metaphors, humor, character and conversational context |
| Culture-Specific | RussianRegions, SAGE, RUBIN | Regional vocabulary, text correction and Russian cultural knowledge |
| Agentic | IFHardBench, SOBHard, GorillaHard | Instruction following, structured documents and tool selection |
| Reasoning | Enantiosemy, NewReasoning, LIMUR | Context-dependent meanings, logic problems and linguistic ambiguity |
Sources
Event date: 2026-08-31. Primary source date: 2026-08-31.