VK Adds Canary-Gated Data Snapshots to Its Ad Selection Service
VK has deployed canary delivery for data snapshots in the production cluster of its advertising selection service. A candidate snapshot is first applied to a limited group handling real traffic; only after technical and product checks pass is it released to the main production group.
In VK’s architecture, a snapshot is a prepared data image that the service loads into memory to process requests. A component called Seeder periodically assembles a consistent dataset from the database, packages it and sends it into the delivery pipeline. VK separates global snapshots containing core service data from snapshots holding advertising campaign data.
Before the canary gate was added, building a snapshot took 30–40 minutes. The file was uploaded to S3, while its metadata entered a snapshot index. The existing orchestrator then notified subscribed instances, coordinated downloads—including peer-to-peer transfers—and instructed them to apply the version. Successful transport, however, did not establish that ad selection still produced the expected result.
VK avoided rebuilding that critical delivery path by routing snapshots through two type_id values. The base type goes to the canary group, while a type carrying the _verified suffix serves the main environment. After approval, the operator publishes another index entry with the same metadata and a reference to the same file, so the large snapshot is neither copied nor rebuilt.
Verification has two consecutive stages. For five minutes, with measurements once per minute, the system checks that most canary instances received the version and remained ready. It then observes impact for 20 minutes, comparing normalized canary and control-group measurements for errors, latency, empty results, selected-ad volume and CPM on key placements.
A Kubernetes CRD-based operator manages the process rather than a periodic script. Configuration, the current phase and measurement results are retained in the resource, allowing work to resume after a controller restart. The design also includes configuration validation, pausing, retries for temporary failures and a manual force-verified option for authorized engineers when automated checks are unavailable.
The gate does not replace recovery measures. VK says a complete canary process requires a representative group, sensitive metrics and the ability to roll back. The team added controls to stop deployment, force a snapshot rebuild and return to the previous version, alongside alerts for the scenario.
Practical context: The broader value is that data quality becomes part of deployment rather than a check performed only after release. The same pattern may suit machine-learning models, search indexes, recommendation features and configurations that change product behavior without changing application code. Its effectiveness still depends on representative traffic, suitable metrics, defensible thresholds and a working rollback route.
| Stage | Analysis window | Success criterion |
|---|---|---|
| Delivery and readiness | 5 minutes, measured once per minute | The version appeared on most canary instances, and the instances remained ready |
| Impact verification | 20 minutes, measured once per minute | Technical and product measurements remained within the permitted deviation from the control group |
Architectural limitations
Two logical streams make configuration more complex, observation windows lengthen delivery, and admission strictness depends on the selected metrics and thresholds. The type_id implementation provides no separate API for managing deployment state. Canary checks also require a viable rollback path; VK says it added the ability to stop deployment, rebuild snapshots and return to the previous version.
Sources
Event date: 2026-09-28. Primary source date: 2026-09-28.