Pool the non-English catalogs into one run, never one run per language

Cosine similarities from two different embedding models are not comparable, so the language groups cannot be merged after the fact. But per-language runs would give five score spaces with two to four cities each, and the evidence unit in this analysis is how many independent cities a theme appears in — which two cities cannot support.

Pooled under the multilingual model they share one space and eleven cities. The cost is that the weakest language sets the floor: Portuguese controls land at the 32nd and 44th percentile against 89th for Italian, so the Brazilian portion of the tail is the least trustworthy part of the run, and the page says so.