Skip to content
-
  • https://www.facebook.com/
  • https://twitter.com/
  • https://t.me/
  • https://www.instagram.com/
  • https://youtube.com/
AI Feed AI Feed AI Feed

AI news, tools, comparisons and practical guides

Subscribe
AI Feed AI Feed AI Feed

AI news, tools, comparisons and practical guides

  • AI News
  • AI Tools Radar
  • AI Comparisons
  • About
  • AI API Prices
  • Local AI
  • AI & Jobs

Sections

  • AI Comparisons
  • AI Features
  • AI Guides
  • AI News
  • Uncategorized

Latest stories

  • VK Reports SESAME Speech-Enhancement Model Results
  • Google Begins Phased Rollout of Gemini Notebook Study Tools
  • Google Releases Gemini 3.8 Live for Real-Time Voice Apps
  • SberTech Describes an MCP Gateway for Enterprise AI Agents
  • T-Bank Publishes Perseus Guide for Four ML Tasks
  • AI News
  • AI Tools Radar
  • AI Comparisons
  • About
  • AI API Prices
  • Local AI
  • AI & Jobs
Subscribe
Close

Search

Home/AI Features/VK Reports SESAME Speech-Enhancement Model Results
Иллюстрация к новости: VK сообщила о модели SESAME для нейросетевого шумоподавления
AI FeaturesAI News

VK Reports SESAME Speech-Enhancement Model Results

Alex
By Alex
16.09.2026 2 Min Read
◉1unique readers

VK reported results for SESAME on September 2, describing an experimental speech-enhancement model developed by students from a joint HSE University and VK workshop. Built on MP-SENet, the system applies a sparse mixture of experts to different acoustic patterns. The project’s code is available on GitHub.

SESAME retains MP-SENet’s encoder, decoders and spectral-processing pipeline while replacing selected transformer feed-forward components with MoE modules. A shared bidirectional GRU supplies temporal context, after which four lightweight experts process features through separate two-layer MLPs. Exactly two experts are activated for each time-frequency token.

The router combines local features with a 64-dimensional recording-level profile derived by averaging the compressed magnitude spectrogram over time and passing it through a linear layer with GELU. According to the developers, this gives routing decisions broader information about the noisy input. The profile is used only by the router and does not enter the experts’ main processing path.

The team rejected Expert Choice, a routing method that can leave tokens unprocessed when expert capacity is limited. In 100-epoch ablations, Expert Choice reduced PESQ from 3.52 to 3.38. The final design instead uses Token Choice, sending every token to exactly two experts without token dropping. MoE modules appear only in the first two of four time-frequency self-attention blocks; placing them in all four also lowered PESQ.

Training used VoiceBank+DEMAND, with 11,572 training pairs and 824 test pairs resampled to 16 kHz. After 400 epochs, the developers reported PESQ of 3.64, CSIG of 4.84, CBAK of 4.04, COVL of 4.37 and SSNR of 10.93. In their comparison, SESAME exceeded PrimeK-Net, base MP-SENet and MP-SENet large on all five metrics.

SESAME has 3.53 million parameters in total, with 2.87 million active for each token. The authors list 3.59 million total and active parameters for MP-SENet large. Their nearly 20% reduction refers specifically to active parameters relative to a similarly sized dense model, rather than measured latency, energy consumption or execution cost.

Practical context: In practical terms, the reported experiments indicate that routing designed for language transformers cannot simply be transferred to audio when it may drop time-frequency tokens: continuity requires every token to be processed. The findings remain limited to one corpus and the developers’ own experiments; the publication provides no independent reproduction, real-device testing or speed measurements.

Final SESAME configuration
Parameter Value
Base architecture MP-SENet
Experts 4
Active experts per token 2 (Token Choice)
MoE-FFN modules 4, in TSB₀ and TSB₁
Recording profile 64-dimensional, router only
Total parameters 3.53 million
Active parameters 2.87 million
Routing variants after 100 epochs
Configuration PESQ Change
SESAME: E=4, k=2, TSB₀–₁, Token Choice 3.52 —
Expert Choice 3.38 −0.14
MoE in all four TS blocks 3.44 −0.08
E=8, k=2, TSB₀–₁ 3.47 −0.05
E=16, k=4, TSB₀–₁ 3.42 −0.10
Limitations of the reported results

All metrics were reported by the project’s developers on the VoiceBank+DEMAND test set. The publication contains no independent reproduction, latency or energy-consumption measurements, or production-environment testing.

Sources

  1. VK Tech

Event date: 2026-09-02. Primary source date: 2026-09-02.

Follow AI Feed on Telegram

New AI stories, practical guides and tool comparisons — in one concise feed.

Open Telegram→

Tags:

Editor’s Picks
Alex
Author

Alex

Follow Me
Other Articles
Иллюстрация к новости: Google добавляет в Gemini Notebook голосовой диалог и учебные обзоры
Previous

Google Begins Phased Rollout of Gemini Notebook Study Tools

Recent posts

  • VK Reports SESAME Speech-Enhancement Model Results
  • Google Begins Phased Rollout of Gemini Notebook Study Tools
  • Google Releases Gemini 3.8 Live for Real-Time Voice Apps
  • SberTech Describes an MCP Gateway for Enterprise AI Agents
  • T-Bank Publishes Perseus Guide for Four ML Tasks

Recent comments

No comments to show.

Archives

  • September 2026
  • May 2026

Sections

  • AI Comparisons
  • AI Features
  • AI Guides
  • AI News
  • Uncategorized

    © 2026 AI Feed. All rights reserved.
    RUEN
    AboutEditorial PolicySources & methodologyCorrectionsContactPrivacyAnalytics settings
    AI Feed analytics

    Helps us understand which pages are useful. Advertising tracking is disabled.