Contextual Building Massing Generation (CoMa)

Vision-language models that generate site-aware 3D building massings for architectural design.

Role: R&D Leader, Sber AI Period: Feb 2026 – present

Massing — the first rough 3D volume of a future building — has to fit the site, its urban context and the economic and programmatic constraints of the project. We frame massing generation as a conditional task for vision-language models that produce site-aware 3D volumes.

We introduced CoMa-20K, a dataset that pairs detailed massing geometries with economic and programmatic data and imagery of the development site in its urban context, benchmarked fine-tuned and large zero-shot vision-language models, and proposed a learned contextual relevance metric (Maslov et al., 2026). The dataset, model weights and code are publicly available.

The project is part of our broader work on generative AI for architecture at Sber AI, which also includes floor plan generation and interior layout synthesis.

References

2026

  1. Smart Cities
    CoMa: Contextual Massing Generation with Vision-Language Models
    Evgenii Maslov, Aleksandra Vabnits, Vladimir Vorona, Anastasia Antsiferova, Valentin Khrulkov, Anastasia Volkova, Anton Gusarov, Andrey Kuznetsov, and Ivan Oseledets
    Smart Cities, 2026