Every other analysis here runs city → UN: take an SDG indicator, find the municipal dataset that matches it. That can only discover what the framework already asks about. This runs it backwards — take every dataset a city publishes, find its nearest SDG indicator, and look at what is left over.

11,206 datasets from 37 city portals (language en), matched against all 519 enumerated SDG indicators — not the 442 with usable data, because a framework gap is a question about vocabulary rather than coverage.

What a result here means. A dataset far from every indicator means no SDG indicator's text is near this dataset's text. That is evidence about vocabulary, not proof of a conceptual gap — this project has already learned once that a null from the matcher is not evidence of absence. So the unit of evidence below is how many independent cities a theme appears in. A theme in thirty city catalogs is a category of municipal governance; a theme in one is that city's filing habit.

Positive controls

Datasets from the hand-verified NYC crosswalk. These are known to correspond to an SDG indicator, so they must land in the high-affinity region. If they fall in the tail, the tail is measuring retrieval failure and nothing below is trustworthy.

9 of 9 located · 1 fell in the tail.

Expected correspondence City Dataset Affinity Percentile In tail?
Road traffic deaths New York Motor Vehicle Collisions - Crashes 0.627 99 no
Child mortality (deaths) New York Infant Mortality 0.578 96 no
Maternal mortality New York Pregnancy-Associated Mortality 0.506 86 no
Fine particulate matter (PM2.5), annual mean New York Air Quality and Health Impacts 0.47 74 no
Deaths attributable to ambient air pollution New York Air Quality and Health Impacts 0.47 74 no
Proportion of municipal waste recycled New York DSNY Monthly Tonnage Data 0.438 61 no
Municipal waste collected New York DSNY Monthly Tonnage Data 0.438 61 no
Intentional homicide New York NYPD Complaint Data Historic 0.425 55 no
Inadequate housing New York Housing Maintenance Code Violations 0.367 21 YES

What cities publish that the SDGs have no words for

The bottom 25% of each catalog by affinity — 2,786 datasets. Below are the title phrases that recur across that tail, counted by how many independent cities use them. These are not inferred categories; they are what the cities themselves called the data.

Cities Datasets Phrase
7 22 street sweeping
6 10 building permits
6 8 special events
6 7 zip code
6 6 fire stations
5 5 call center
5 5 bike share
5 5 bus stops
5 5 zip codes
4 7 business licenses
4 6 work orders
4 6 bus stop
4 4 street name
4 4 council district

In detail

street sweeping — 22 datasets across 7 cities

building permits — 10 datasets across 6 cities

special events — 8 datasets across 6 cities

zip code — 7 datasets across 6 cities

fire stations — 6 datasets across 6 cities

call center — 5 datasets across 5 cities

bike share — 5 datasets across 5 cities

bus stops — 5 datasets across 5 cities

zip codes — 5 datasets across 5 cities

business licenses — 7 datasets across 4 cities

work orders — 6 datasets across 4 cities

bus stop — 6 datasets across 4 cities

street name — 4 datasets across 4 cities

council district — 4 datasets across 4 cities

The same tail, clustered by embedding

A second view, kept because it groups datasets that share no vocabulary. k-means returns k clusters whether or not k themes exist, so each carries its coherence — the mean cosine of its members to its own centroid. Below 0.62, a cluster is a partition rather than a theme and is marked diffuse; 17 of 40 are. Read those as noise, not as findings.

Coherence Cities Datasets Terms Closest SDG indicator
0.992 3 231 dfs, check, speed, sign Beach litter items per unit of surface area
0.839 5 70 libraries, location, holds Countries with users/communities participati
0.81 6 42 foia, request, log Countries that have legislative, administrat
0.766 2 50 repave, miles, arterials, arterial, lane Number of deaths rate due to road traffic in
0.742 7 44 insight, edmonton, survey, community Participation rate in organized learning (on
0.713 24 159 (no term covers a fifth of this cluster) Land area
0.707 10 57 election, voting, general, results Number of local governments
0.702 23 151 (no term covers a fifth of this cluster) Countries that have national urban policies
0.692 19 92 (no term covers a fifth of this cluster) Number of deaths rate due to road traffic in
0.687 2 32 optimized, corridors, timing, signal Progress toward productive and sustainable a
0.673 19 128 (no term covers a fifth of this cluster) Land area
0.672 13 56 (no term covers a fifth of this cluster) Number of deaths rate due to road traffic in
0.662 3 100 des Coastal Eutrophication: Total Nitrogen (micr
0.661 16 99 school Extent to which global citizenship education
0.658 21 90 (no term covers a fifth of this cluster) Countries that adopt and implement constitut
0.654 17 72 (no term covers a fifth of this cluster) Number of local governments
0.654 6 51 edmonton Countries that have national urban policies
0.654 4 63 stairway, sidewalks Number of deaths rate due to road traffic in
0.64 22 141 (no term covers a fifth of this cluster) Number of deaths rate due to road traffic in
0.64 15 81 (no term covers a fifth of this cluster) Countries that adopt and implement constitut
0.638 20 88 (no term covers a fifth of this cluster) Net inbound official development assistance
0.629 19 78 bike, parking Number of deaths rate due to road traffic in
0.622 6 19 covid- Number of total conflict-related deaths
0.62 (diffuse) 16 42 csb, update, response, time Police reporting rate for robbery in the pre
0.619 (diffuse) 18 67 (no term covers a fifth of this cluster) Countries with users/communities participati
0.619 (diffuse) 13 41 ems, calls Number of deaths due to disaster
0.611 (diffuse) 17 57 neighborhood, street Number of deaths rate due to road traffic in
0.611 (diffuse) 14 43 sweeping, schedule, street Countries with procedures in law or policy f
0.608 (diffuse) 8 25 beudo, engagement, employee, buildings International financial flows to developing
0.605 (diffuse) 20 67 (no term covers a fifth of this cluster) Land area
0.601 (diffuse) 7 55 school Adjusted gender parity index for participati
0.6 (diffuse) 10 21 (no term covers a fifth of this cluster) Beach litter items per unit of surface area
0.6 (diffuse) 7 43 (no term covers a fifth of this cluster) Score of adoption and implementation of nati
0.598 (diffuse) 17 36 analytics Proportion of results indicators drawn from
0.597 (diffuse) 26 88 permits, building Countries with procedures in law or policy f
0.593 (diffuse) 9 40 storm Total inbound official flows for infrastruct
0.58 (diffuse) 14 59 employee Total government revenue, in local currency
0.514 (diffuse) 16 47 (no term covers a fifth of this cluster) Countries with integrated biodiversity value
0.513 (diffuse) 13 35 library, austin Total inbound official flows for infrastruct
0.481 (diffuse) 11 26 (no term covers a fifth of this cluster) Countries that have legislative, administrat

The coherent clusters in detail

dfs, check, speed, sign — 231 datasets across 3 cities (coherence 0.992, mean affinity 0.323)

Most central to the cluster:

One per city, to show the spread:

libraries, location, holds — 70 datasets across 5 cities (coherence 0.839, mean affinity 0.324)

Most central to the cluster:

One per city, to show the spread:

foia, request, log — 42 datasets across 6 cities (coherence 0.81, mean affinity 0.335)

Most central to the cluster:

One per city, to show the spread:

repave, miles, arterials, arterial, lane — 50 datasets across 2 cities (coherence 0.766, mean affinity 0.25)

Most central to the cluster:

One per city, to show the spread:

insight, edmonton, survey, community — 44 datasets across 7 cities (coherence 0.742, mean affinity 0.334)

Most central to the cluster:

One per city, to show the spread:

(no term covers a fifth of this cluster) — 159 datasets across 24 cities (coherence 0.713, mean affinity 0.338)

Most central to the cluster:

One per city, to show the spread:

election, voting, general, results — 57 datasets across 10 cities (coherence 0.707, mean affinity 0.321)

Most central to the cluster:

One per city, to show the spread:

(no term covers a fifth of this cluster) — 151 datasets across 23 cities (coherence 0.702, mean affinity 0.321)

Most central to the cluster:

One per city, to show the spread:

(no term covers a fifth of this cluster) — 92 datasets across 19 cities (coherence 0.692, mean affinity 0.327)

Most central to the cluster:

One per city, to show the spread:

optimized, corridors, timing, signal — 32 datasets across 2 cities (coherence 0.687, mean affinity 0.265)

Most central to the cluster:

One per city, to show the spread:

Reproducing

python3 probe/fetch_municipal.py
python3 probe/inverse.py