One model instead of a pipeline
Most production recommenders work like a chain: candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex's Sona brings candidate generation and ranking into a single generative AI system.
What the live test showed
Yandex tested the model in a seven-day live A/B experiment on its smart speakers. In the test, more than 15 candidate generators, the pre-ranking stage and the ranking stage were replaced by one served transformer. According to the company, playback can begin without the user first picking an artist, genre or mood.
How the architecture works
Sona reads the listener's history once per request through an encoder. A decoder generates candidates, and a ranking module scores them against the same shared representation. There are no hand-engineered features; inputs are logged event fields and learned Semantic IDs.
- Every track becomes a tuple of 3 discrete codes.
- The encoder attends to 8,192 past events, with deeper processing for the recent 2,048.
- A frozen teacher ranker guides training and is removed at serving time.
- Training stays online, with new weights reaching serving every 10 minutes.
Why it matters
Recommenders are usually built by stitching together separately trained models, each optimizing its own objective. Sona puts candidate generation and ranking around one shared user representation. Yandex says serving runs on NVIDIA Triton Inference Server and reaches 41% model FLOPs utilization.



