Skip to content
-
  • Facebook
  • X
  • Telegram
  • Instagram
  • YouTube
AI Feed AI Feed AI Feed

AI news, tools, comparisons and practical guides

Subscribe
AI Feed AI Feed AI Feed

AI news, tools, comparisons and practical guides

  • AI News
  • Radar
  • AI Comparisons
  • About
  • API Prices
  • Local AI
  • Jobs

Sections

  • AI Comparisons
  • AI Features
  • AI Guides
  • AI News
  • Jobs
  • Uncategorized

Latest stories

  • Google Details How Edy’s Grocer Uses Gemini for Catering
  • VK Adds Canary-Gated Data Snapshots to Its Ad Selection Service
  • Google Showcases Four Projects Built with Gemini 3.8 Flash
  • Google Unifies Vertical Video Buying in Display & Video 360
  • What Does an AI API Really Cost? Chat, RAG and Agent Budgets
  • AI News
  • Radar
  • AI Comparisons
  • About
  • API Prices
  • Local AI
  • Jobs
Subscribe
Close

Search

Home/AI Features/Sber Releases MERA Text 2.0 Benchmark for Language Models
Иллюстрация к новости: Сбер представил бенчмарк MERA Text 2.0 для оценки языковых моделей
AI FeaturesAI News

Sber Releases MERA Text 2.0 Benchmark for Language Models

Alex
By Alex
26.09.2026 2 Min Read
◉2unique readers

Sber introduced MERA Text 2.0 on August 31, restructuring its text-language-model benchmark around 12 tests in four skill groups: Human-Centric, Culture-Specific, Agentic and Reasoning. The previous version’s public leaderboard has been frozen and will no longer receive updates.

The MERA team said the redesign responds to saturation in older tasks. Leading models now approach the human baseline on the previous benchmark and exceed it on some tests, compressing differences near the top of the ranking. Rather than adding a few datasets, the team rebuilt the evaluation around abilities it considers difficult and informative for current models.

Human-Centric tasks assess riddles, mechanisms of humor and whether answers fit a described character. Culture-Specific testing covers Russian regional vocabulary, correction of real Russian-language text and cultural references spanning folklore, sayings, memes, songs and films. The team reports that this cultural group has the greatest room for improvement and correlates less strongly with overall model performance than the other groups.

The Agentic group does not attempt to evaluate a complete autonomous agent. IFHardBench requires models to satisfy three to six formal constraints simultaneously, while SOBHard tests document extraction and transformation. GorillaHard asks a model to select tools and supply appropriate call arguments, but measures planning rather than actual tool execution.

Reasoning tests cover words whose meanings reverse with context, short problems containing logical traps, and linguistic ambiguity. All MERA Text 2.0 test sets are private, and each test uses at least five instruction variants. The organizers also fix task protocols, retain run parameters and logs, and avoid adding agentic, reasoning or tool harnesses that could solve parts of the tasks for the evaluated model.

Scores are aggregated in three stages: metrics within each test, tests within each group, and then the four group scores. Every stage uses a geometric mean. A zero is replaced with 0.01 so that one failed metric does not erase the entire score, and valid comparisons require all models to be recalculated on the same test-set version.

Practical context: The grouped design can reveal whether a model’s weakness lies in cultural context, strict output constraints or tool selection instead of reducing every result to an undifferentiated average. Its scope remains narrower than full agent evaluation: long-running scenarios, environment interaction and complex end-to-end workflows are deliberately excluded from this benchmark line.

A full run has been reduced from 23 datasets and about 37,000 requests to 12 datasets and roughly 8,700 requests—a 4.3-fold decrease by request count. Previous submissions and results remain available in user accounts, while new measurements follow the existing submission process. At announcement time, the technical report was still forthcoming; the team directs readers to the live benchmark for current results rather than a static ranking.

The four MERA Text 2.0 test groups
Group Tests What they assess
Human-Centric Riddles, Humor, Characters Metaphors, humor, character and conversational context
Culture-Specific RussianRegions, SAGE, RUBIN Regional vocabulary, text correction and Russian cultural knowledge
Agentic IFHardBench, SOBHard, GorillaHard Instruction following, structured documents and tool selection
Reasoning Enantiosemy, NewReasoning, LIMUR Context-dependent meanings, logic problems and linguistic ambiguity

Sources

  1. Sber Tech

Event date: 2026-08-31. Primary source date: 2026-08-31.

Follow AI Feed on Telegram

New AI stories, practical guides and tool comparisons — in one concise feed.

Open Telegram→

AI Feed topic

Continue exploring

Related reporting and practical tools selected for this topic.
  1. 29.09.2026Google Showcases Four Projects Built with Gemini 3.8 Flash↗
  2. 24.09.2026Anthropic Says Claude Identified the ART Enzyme System↗
  3. 23.09.2026Google Launches Gemini 3.8 Flash and Flash-Lite TTS↗
  4. 23.09.2026SberTech Reworks Documentation AI Agent Around Vector Search↗
AI model radar→

Tags:

Editor’s Picks
Alex
Author

Alex

Follow Me
Other Articles
Local AI for Mac — модели для 8–128 ГБ памяти
Previous

Best Local AI Models for Mac by RAM: 8GB to 128GB Guide

Иллюстрация к новости: Т-Инвестиции запустили MCP-сервер для инвестиционных ИИ-агентов
Next

T-Investments Opens MCP Server to AI Agents

Recent posts

  • Google Details How Edy’s Grocer Uses Gemini for Catering
  • VK Adds Canary-Gated Data Snapshots to Its Ad Selection Service
  • Google Showcases Four Projects Built with Gemini 3.8 Flash
  • Google Unifies Vertical Video Buying in Display & Video 360
  • What Does an AI API Really Cost? Chat, RAG and Agent Budgets

Recent comments

No comments to show.

Archives

  • September 2026
  • May 2026

Sections

  • AI Comparisons
  • AI Features
  • AI Guides
  • AI News
  • Jobs
  • Uncategorized

    © 2026 AI Feed. All rights reserved.
    RUEN
    AboutEditorial PolicySources & methodologyCorrectionsContactPrivacyAnalytics settings
    AI Feed analytics

    Helps us understand which pages are useful. Advertising tracking is disabled.