Este artículo aún no se ha traducido al Español: estás leyendo el original en English. También disponible en:Deutsch, English, Українська
What Share of a Country Should Sixty Cities Hold? On Choosing a Threshold You Can Defend
The cheapest data check we run is one line of arithmetic: add up the populations of the cities in a locality file, divide by the population of the country, look at the ratio. It once caught a Taiwanese list summing to 106.9% of Taiwan, which is physically impossible and took a second to see.
Then we said the threshold should not be 100% but "a plausible urban share for that specific country," and moved on as though that were an answer. It is not an answer. It is the name of a question.
This article is our attempt to actually answer it across 39 locales, and the finding is not the one we expected. The share of a country held by our city list is mostly determined by decisions we made about the list, not by facts about the country — which means the number is nearly uninterpretable on its own, and a universal threshold cannot exist even in principle. What we can offer instead is a procedure that replaces one unanswerable question with two answerable ones, plus a warning about naming your metrics that we earned the hard way while writing this.
Two numbers, named
Before any figures, a piece of vocabulary that cost us an afternoon of confusion. There are two entirely different percentages you can compute from a weighted city list, both natural, both roughly in the 50–80% range for small countries, and both easy to write down as "the share":
| Name | Definition | Answers |
|---|---|---|
| Coverage | sum of the whole list ÷ national population | How much of the country does this list reach? |
| Head share | sum of the top ten rows ÷ sum of the whole list | How top-heavy is this list internally? |
Georgia is at 86.8% head share and 64.4% coverage. Both are correct; they are not versions of one number. Unlabelled, the two series merge into one and produce a table where nothing means anything. Every percentage below says which one it is.
Coverage is the quantity the classic sum-against-the-whole check produces, so it is the one this article is mainly about. Head share turns up later as the thing you might reach for to rescue the check — and does not.
The actual spread
Our locality set is 64 files. Of those, 44 carry population weights and 20 ship uniform. The coverage table below was computed when 39 were weighted; bn_BD, en_AU, no_NO, tr_TR and zh_TW have been weighted since and are not in it. Summing each weighted file against its country's population gives coverage:
| Locale | Cities | Coverage | Locale | Cities | Coverage |
|---|---|---|---|---|---|
en_NZ | 60 | 70.4% | ro_RO | 60 | 37.0% |
fi_FI | 60 | 70.1% | vi_VN | 60 | 33.0% |
he_IL | 55 | 67.0% | pl_PL | 55 | 32.3% |
fr_CA | 60 | 66.4% | el_GR | 60 | 31.9% |
hy_AM | 41 | 66.0% | pt_BR | 58 | 31.0% |
lv_LV | 58 | 66.0% | de_DE | 57 | 28.5% |
ka_GE | 45 | 64.4% | en_PH | 60 | 27.6% |
bg_BG | 60 | 61.3% | it_IT | 58 | 25.9% |
hr_HR | 60 | 60.4% | zh_CN | 60 | 21.9% |
kk_KZ | 52 | 58.3% | ne_NP | 51 | 21.1% |
en_CA | 60 | 48.5% | id_ID | 60 | 18.6% |
ro_MD | 43 | 47.1% | fr_BE | 60 | 16.7% |
hu_HU | 60 | 46.5% | fr_FR | 56 | 16.2% |
uk_UA | 54 | 44.0% | en_US | 54 | 15.3% |
sv_SE | 58 | 43.6% | en_ZA | 60 | 14.9% |
de_AT | 60 | 42.9% | en_IN | 60 | 9.0% |
(Locales between 40% and 42% — sk_SK, ja_JP, sl_SI, cs_CZ, fa_IR, ru_RU, es_ES — omitted from the table for length; all fall in the middle of the range.)
Coverage runs from 9.0% to 70.4%, a factor of nearly eight. No single number separates the good files from the bad ones, because 9.0% and 70.4% are both correct.
Note also what is not here. The 100% test had, over the whole set, rejected exactly one file: Taiwan at 106.9%. The other locale we pulled weights from, Australia at 84.7%, passed the 100% test comfortably and was rejected by a human noticing that Sydney's row carried the Greater Sydney figure. A test that fires once across 64 files, and misses the second instance of the very defect it targets, is not carrying much weight. (Both files were repaired afterwards and both now ship weighted — which does not rescue the test, since in neither case was it the test that found the defect.)
Why the number means so little
The share is a ratio, and we controlled the numerator. Two decisions we made dominate it, and neither has anything to do with whether the data is right.
First: how deep the list goes. Every list is capped at roughly 60 cities. What that cap means depends entirely on the country. Here is the smallest city in each list — the point where the list stops:
| Locale | Coverage | Smallest city on the list | Median city |
|---|---|---|---|
lv_LV | 66.0% | 753 | 6,410 |
ka_GE | 64.4% | 899 | 10,747 |
hy_AM | 66.0% | 1,010 | 13,133 |
en_NZ | 70.4% | 2,316 | 18,447 |
sl_SI | 41.3% | 1,519 | 6,037 |
en_ZA | 14.9% | 12,790 | 87,348 |
ja_JP | 41.5% | 189,591 | 474,592 |
en_US | 15.3% | 199,723 | 639,111 |
pt_BR | 31.0% | 302,692 | 698,642 |
zh_CN | 21.9% | 547,879 | 3,608,835 |
en_IN | 9.0% | 644,406 | 1,162,472 |
The smallest city in the Latvian file has 753 residents. The smallest city in the Indian file has 644,406. That is a factor of 856 in where the two lists stop.
This single fact explains most of the spread. Sixty Latvian entries reach down into villages, and Latvia does not have many settlements below that, so the list captures most of the country's urban system and lands at 66%. Sixty Indian entries stop at "cities above 644,000," below which India has several thousand more towns, so the list captures a sliver and lands at 9.0%. Both lists are behaving exactly as a 60-row cap forces them to behave.
Second: how big the country is. A 60-city cap in a nation of 1.4 billion and a 60-city cap in a nation of 1.9 million are not the same instrument. India would need tens of thousands of rows to reach the depth Latvia reaches with 58. No threshold applied to both can mean the same thing.
Put together: expected coverage is set by the cap and the country's size, and the true urban structure only modulates it. Asking whether 31.0% coverage is plausible for Brazil without knowing that Brazil's list stops at 302,692 residents is asking about a ratio while ignoring its numerator's definition.
The depth figure also explains head share, which is why it belongs at the centre of this article rather than either percentage. India's head share is low (51.9%) because its list contains sixty large cities of comparable size — an artefact of stopping at 644,406. Latvia's is high (78.2%) because its list runs from Riga down to a village of 753, so the top ten inevitably dominate. Both percentages are shadows of the same decision about where to stop.
Head share does not rescue coverage
The obvious next move is to correct coverage using something computable from the file itself. If a country is dominated by one huge city, you would expect high coverage from few rows; if its cities are evenly sized, you would expect the opposite. Head share measures exactly that. So does it predict coverage?
It does not. The two are close to unrelated, and two locales demonstrate it by failing in opposite directions:
| Locale | Head share | Coverage | Smallest city | Reading |
|---|---|---|---|---|
fi_FI | 55.4% | 70.1% | 7,766 | Flat list, yet reaches most of the country |
de_AT | 77.3% | 42.9% | 7,349 | Top-heavy list, yet reaches under half |
Finland's list is the least top-heavy of the pair and has the highest coverage. Austria's is heavily concentrated in its top ten and covers barely half as much. Whatever correction you build from head share, one of these two comes out backwards.
The full set behaves the same way. Georgia has the highest head share we measured, 86.8%, at 64.4% coverage. Latvia is 78.2% head share at 66.0% coverage. New Zealand is 74.3% at 70.4%. India is 51.9% head share — a genuinely flat, evenly-sized list of large cities — at 9.0% coverage. Israel is 53.0% head share at 67.0% coverage, almost exactly India's head share with seven times the coverage.
Note what that last pair means. en_IN and he_IL have nearly identical internal shape — 51.9% and 53.0% of each list sits in its top ten — and coverages of 9.0% and 67.0%. The internal shape of the list carries essentially no information about how much of the country it reaches.
So the two metrics answer genuinely different questions, and neither can stand in for the other. Head share tells you about the shape of the list; coverage tells you about its reach. We looked for an internal signal that predicts expected coverage and did not find one that survives these cases.
There is a further trap in the vocabulary itself. Both quantities sit in the same numeric range for small countries — Latvia is 78.2% and 66.0%, Armenia 81.8% and 66.0% — so a table that omits the metric name produces two series that look like one noisy series. We spent real time this session reconciling "78.2%" against "66.0%" for Latvia before noticing they were answers to different questions. Neither figure was wrong. The label was missing, which is a data defect in its own right and one that no arithmetic check can catch.
The bound that is genuinely better than 100%
There is one improvement available, and it is worth stating precisely because it is real but limited.
The sum of a country's city populations cannot exceed the country's urban population. That is a physical constraint and it is strictly tighter than 100%, because no country is 100% urban. Applied to Australia's rejected file: 84.7% of Australians in 60 cities, against an Australian urbanisation rate somewhere around 86%, implies those 60 cities held roughly 98% of every urban Australian — which is absurd for a country with hundreds of urban centres, and it is absurd in a way that a 100% threshold cannot express and a human noticed only by accident.
So the improved check is: compare the sum against the urban population, not the total population. It costs one extra external number per country and it would have caught Australia mechanically.
But it must be labelled honestly, because it inherits the exact disease it is meant to diagnose. A national urbanisation rate is itself the output of somebody deciding what counts as a city — the same definitional choice that produced the defect in the first place. Countries classify settlements as urban by wildly different rules: by population threshold, by administrative status, by density, by economic activity. A bound built on that number is indicative, not decisive. It narrows the suspect list; it does not convict.
We ran this bound over our set and it puts three locales close enough to their estimated urban populations to be worth a look — ka_GE, hy_AM and ro_MD, each a small country where our list reaches down to roughly a thousand residents. We are not reporting those as defects. The known unit traps in Georgia and Armenia — Kaspi's three nested figures, Chambarak's քաղաք-versus-համայնք pair — are already correctly resolved in the shipped files, which we checked. What the bound produced is a queue for human review, which is the most any threshold of this kind can produce.
So: is there a formula?
No. We looked for one and we are reporting its absence as the result.
There is no function from the contents of a locality file to a defensible expected coverage, because the dominant term is a cap we imposed, the second term is country size, and the third — actual urban structure — is the only one carrying signal about correctness and it is the smallest of the three. Any threshold that fires on the output number is really firing on our own row limit. And the most promising internal corrective, head share, turns out to measure something orthogonal.
What replaces it is not a formula but a reframing. The unanswerable question "is this coverage plausible?" splits into two answerable ones:
- Is the top of the list the right unit? This is where the real defects live. Taiwan's list mixed municipalities with their own districts; Australia's substituted agglomerations for local government areas. Both are visible by taking the largest three or four entries and checking each against an independently known figure for that specific city. Sydney at 5,219,674 against a City of Sydney of roughly 200,000 is a one-minute check that needs no threshold at all.
- Where does the list stop, and was that on purpose? Record the smallest entry as a stated depth rule — "this file contains settlements above N residents." Once that is written down, low coverage stops being suspicious and becomes arithmetic. India at 9.0% needs no defence once the file says it stops at 644,406.
And a third, free one: label every percentage with its metric. Coverage and head share cost us more confusion this session than any wrong number did.
The coverage check keeps its place, but demoted: it is a detector of impossibility, not of implausibility. Above 100% something is definitely wrong. Below 100% it says nothing you can act on without the two checks above. That is a much smaller claim than the one we were implicitly making, and it is the one the data supports.
If that sounds like an admission that this check requires manual judgement per country, it is. Twenty of our 64 files ship without weights precisely because that judgement came back "we cannot make this consistent," and we would rather ship a documented approximation than an invisible defect. The honest version of a data-quality process includes a column for undecided.
For the defect class this check was built to catch, see Sixty Cities, 106.9% of a Country. For what happens downstream when weights are absent, see Why Address Generators Produce Implausible Cities. For the general problem of checks that cannot fail, see A Test That Cannot Fail.
Data as of 2026-07-18
Every share, city count, smallest-city and median figure in this article was computed on 18 July 2026 directly from the shipped locality files, by summing the fourth column of each .tsv and dividing by the population table used in our regression suite. Nothing here is quoted from working notes.
Sources and notes:
- Locality files — 64 files in
share/name_gen/locality/. 44 carry a fourth population column and 20 are three-column and ship uniform, re-counted on 22 July 2026; the uniform set includesar_SA,en_GB,en_IE,ko_KR,nl_NL,pt_PTandes_AR. When the coverage figures above were computed the split was 39/25;bn_BD,en_AU,no_NO,tr_TRandzh_TWgained weights afterwards. - Population denominators — the table embedded in
ng_regression_tests.php, which carries rounded national populations (Latvia 1,870,000; India 1,430,000,000; and so on). Shares are therefore accurate to roughly the rounding of that table, which is adequate at the resolution this article uses and inadequate for anything finer. - Extremes —
en_IN127,990,788 over 60 cities = 9.0%;en_NZ3,658,551 over 60 cities = 70.4%. Smallest entries:lv_LV753,en_IN644,406, a ratio of 856. - Both metrics, computed side by side on 18 July 2026. Head share is the sum of the ten largest rows over the sum of all rows; coverage is the sum of all rows over the national population.
| Locale | Head share | Coverage | Smallest city |
|---|---|---|---|
ka_GE | 86.8% | 64.4% | 899 |
hy_AM | 81.8% | 66.0% | 1,010 |
lv_LV | 78.2% | 66.0% | 753 |
de_AT | 77.3% | 42.9% | 7,349 |
en_NZ | 74.3% | 70.4% | 2,316 |
ro_MD | 74.1% | 47.1% | 2,655 |
en_ZA | 57.8% | 14.9% | 12,790 |
fi_FI | 55.4% | 70.1% | 7,766 |
he_IL | 53.0% | 67.0% | 22,896 |
en_IN | 51.9% | 9.0% | 644,406 |
en_US | 50.9% | 15.3% | 199,723 |
zh_CN | 45.2% | 21.9% | 547,879 |
fi_FI sums to 3,927,865 and hy_AM to 1,848,547. The fi_FI / de_AT inversion and the en_IN / he_IL pair are the two cases that rule out predicting coverage from head share.
- Australia's 84.7% — a recollection, not a re-derivation.
en_AUships without weights, so there is no fourth column left to sum; what is confirmed on disk is the consequence. Treat the figure as indicative and see What Counts as a City for the same caveat at its source. - Urbanisation rates — external, approximate, and deliberately not tabulated here. They are used in this article only to make a qualitative point about Australia and to explain why the improved bound is soft. Any specific national urbanisation figure should be taken from the relevant statistics office, with attention to how that office defines "urban."
- Georgia and Armenia spot-checks —
ka_GEships Kaspi at 12,911 (the city figure, not the 41,841 municipality or the 13,729 administrative unit) andhy_AMships Chambarak at 5,468 (the քաղաք figure, not the 12,416 համայնք). Both known traps are resolved correctly in the shipped files; this was verified by reading the rows, not assumed.