Este artículo aún no se ha traducido al Español: estás leyendo el original en English. También disponible en:Deutsch, English, Українська
Rooted or Imported? Telling Them Apart With Nothing but the Registry
Spain's national statistics institute, INE, publishes its surnames without accents. GARCIA, MARTINEZ, TRAORE, CALDERON — 86,512 rows, 1,832 of them carrying Ñ, and zero carrying an acute accent or a diaeresis. The counts are exact and the spelling is stripped.
If you want to add surnames from that register to a corpus that spells Spanish properly, you have to put the accents back. Spanish is unusually cooperative here: stress and accent are rule-governed, and one class is close to airtight. A surname ending in -on, -an or -in takes the accent on the final vowel. CALDERON → Calderón. BELTRAN → Beltrán. MARTIN → Martín.
We measured that rule against 140 surnames in an existing, hand-curated Spanish corpus where the correct spelling was already known. It reproduced 136 of 140 — 97.14%.
Then we looked at the four it missed:
Lin · Hussain · Khan · Pan
The rule is not 97% accurate. It is 100% accurate on Spanish surnames and 0% accurate on everything else, and the register does not tell you which is which. Applied to the register's tail, it would have produced Johnsón, Wilsón, Thompsón, Robinsón, Moldován, Constantín, Rahmán and Hassán — a set of misspellings that would have sat in a name corpus permanently, each one wrong in a confident, Spanish-looking way.
What was needed was a test that separates rooted surnames from imported ones using only the register. There is one, it is a single division, and it works because of an accident in Spanish naming law.
The mechanism, in one paragraph
A Spanish person carries two surnames: the first (apellido 1º) from the father, the second (apellido 2º) from the mother. INE publishes them as separate columns. A surname reaches the second column only when a woman carrying it has had a child registered in Spain, and that child is now a counted resident. A surname that arrived recently can be common as ap1 and near-absent as ap2, because the second column takes a generation to fill.
So the ratio
c2 / c1
is not an estimate of anything. It is a count of how far a surname has propagated through the maternal line inside the Spanish registry. We covered why that works, and what it does and does not mean, in How Spain's Surname Registry Accidentally Measures Integration. This article is about turning it into a decision procedure with a threshold, which is a different and harder problem.
Where to put the threshold
A ratio is only a filter once you know where to cut. Picking a round number would be fitting; the honest way is to look at the shape of the distribution first.
Across the 2,222 hand-curated Spanish surnames — a set assembled without ever consulting this column — the distribution is a spike:
c2/c1 band | Surnames | Shape |
|---|---|---|
| 0.0 – 0.1 | 16 | # |
| 0.1 – 0.2 | 2 | # |
| 0.2 – 0.3 | 5 | # |
| 0.3 – 0.4 | 1 | # |
| 0.4 – 0.5 | 2 | # |
| 0.5 – 0.6 | 2 | # |
| 0.6 – 0.7 | 6 | # |
| 0.7 – 0.8 | 13 | # |
| 0.8 – 0.9 | 77 | #### |
| 0.9 – 1.0 | 1,358 | #################################################### |
| 1.0 – 1.1 | 712 | ############################ |
| 1.1 – 1.2 | 23 | ## |
| above 1.2 | 5 | # |
2,070 of 2,222 — 93.2% — land between 0.9 and 1.1. Median 0.987, p10 0.924, p90 1.036. A settled Spanish surname is transmitted through both parental lines at almost exactly the same rate, and it has been for centuries, so the ratio pins itself to 1.0 with very little scatter.
Now the same measurement over the register itself, restricted to surnames with c1 ≥ 500 (7,991 of them). Here the distribution is not a spike. It is bimodal:
c2/c1 band | Surnames | Shape |
|---|---|---|
| 0.0 – 0.1 | 76 | #### |
| 0.1 – 0.2 | 251 | ############# |
| 0.2 – 0.3 | 159 | ######## |
| 0.3 – 0.4 | 179 | ######### |
| 0.4 – 0.5 | 128 | ####### |
| 0.5 – 0.6 | 59 | ### |
| 0.6 – 0.7 | 55 | ### |
| 0.7 – 0.8 | 128 | ####### |
| 0.8 – 0.9 | 835 | ########################################## |
| 0.9 – 1.0 | 3,617 | #####…##### (the peak) |
| 1.0 – 1.1 | 2,187 | ########################################################## |
| above 1.1 | 317 | ################ |
Two populations, and a valley between them. Just 114 surnames out of 7,991 — 1.43% — sit in the band from 0.5 to 0.7. Below it, 793 surnames (9.9%). Above it, 7,084 (88.6%).
That is where the threshold goes. 0.70, in the empty part. A threshold placed in a dense region is a judgement call that moves your results when you nudge it; a threshold placed in a valley barely moves them at all. Ours could have been 0.60 or 0.75 and the outcome would have changed by a handful of names.
What the filter actually caught
At the working threshold — surnames with c1 ≥ 500 proposed for accent restoration — the ratio test rejected 876 candidates. Seventy of them ended in -ON, -AN or -IN, meaning the accent rule was about to fire on them:
| Register surname | c1 | c2 | c2/c1 | What the rule would have written |
|---|---|---|---|---|
MOLDOVAN | 3,073 | 348 | 0.113 | Moldován |
HASSAN | 2,948 | 883 | 0.300 | Hassán |
STAN | 2,719 | 347 | 0.128 | Stán |
CONSTANTIN | 2,652 | 371 | 0.140 | Constantín |
SERBAN | 2,099 | 270 | 0.129 | Serbán |
STEFAN | 2,094 | 272 | 0.130 | Stefán |
RAHMAN | 1,858 | 199 | 0.107 | Rahmán |
MURESAN | 1,720 | 267 | 0.155 | Muresán |
JOHNSON | 1,712 | 581 | 0.339 | Johnsón |
WILSON | 1,673 | 474 | 0.283 | Wilsón |
ION | 1,659 | 259 | 0.156 | Ión |
THOMPSON | 1,251 | 305 | 0.244 | Thompsón |
IVAN | 1,232 | 146 | 0.119 | Iván |
ROBINSON | 1,151 | 244 | 0.212 | Robinsón |
Romanian, Bangladeshi, Arabic, Chinese and English surnames, all of them legitimately in the Spanish register, all of them about to be handed a Spanish accent they have never carried. Not one of them reaches 0.35.
An earlier pass at a lower size threshold produced the same failure with different names — Feldmán (FELDMAN, 103/47 = 0.456), Hanssón (101/34 = 0.337), Donaldsón (101/14 = 0.139), Mirzoyán (100/27 = 0.270), Cirján (100/8 = 0.080), Sarhán (102/23 = 0.225), Otmán (100/48 = 0.480). Every single one falls under the threshold.
And raising the size threshold would not have saved us. MOLDOVAN at 3,073 bearers and CONSTANTIN at 2,652 sit near the top of the candidate pool, not in the tail. Depth was not the problem. The problem was that "looks Spanish in the Latin alphabet" is not a property the alphabet can express.
What it correctly did not catch
A filter that rejects everything foreign-looking is not a filter, it is a prejudice. This one keeps names that look wrong and are right.
| Surname | c1 | c2 | c2/c1 | Verdict |
|---|---|---|---|---|
HUAMAN | 2,328 | 2,614 | 1.123 | kept → Huamán |
ADRIAN | 2,611 | 2,662 | 1.020 | kept → Adrián |
HOLGUIN | 3,686 | 3,766 | 1.022 | kept → Holguín |
ESTEBAN | 43,313 | 42,416 | 0.979 | kept (documented accent exception) |
SANTISTEBAN | 2,418 | 2,187 | 0.904 | kept (same exception class) |
Huamán is the one that matters. It is Quechua in origin, it does not look Castilian, and a heuristic built on "does this look Spanish" would have thrown it out. It has been in Peru and in Spanish-speaking families for centuries, its second column is fuller than its first, and the counter keeps it without knowing a word of Quechua.
The 300 surnames that survived the whole pipeline have a minimum ratio of 0.728, a median of 0.980 and a maximum of 1.337. Nothing marginal got through.
And here is the check that makes the whole thing hang together. Recall the four surnames that broke the accent rule on the curated corpus:
| Rule-breaker | c1 | c2 | c2/c1 |
|---|---|---|---|
LIN | 14,151 | 1,194 | 0.084 |
HUSSAIN | 6,549 | 568 | 0.087 |
KHAN | 6,296 | 1,156 | 0.184 |
PAN | 4,323 | 2,025 | 0.468 |
All four, far below the threshold. Meanwhile the five surnames in the corpus that belong to the rule's documented phonological exception class — Esteban, Juan, Sanjuán, Santisteban, Holguín — sit at 0.979, 1.032, 0.977, 0.904 and 1.022.
Two completely independent instruments — a phonological rule about stress, and a division of two columns in a census file — partition the same 140 surnames the same way. Remove the four that the ratio test rejects and the accent rule scores 136 out of 136. Neither instrument can see what the other is looking at. That agreement is the strongest evidence in this article that the ratio is measuring something real.
The precision depends on size, and that is not optional
Here is the part that a threshold on its own hides. The ratio is a quotient of two counts, and small counts make bad quotients. Measured across the whole register by size band:
c1 range | Surnames | p10 of c2/c1 | Median | p90 | Share landing in the 0.5–0.7 valley |
|---|---|---|---|---|---|
| 20 – 99 | 51,107 | 0.214 | 0.589 | 1.200 | 16.21% |
| 100 – 499 | 19,657 | 0.250 | 0.874 | 1.133 | 7.33% |
| 500 – 1,999 | 5,572 | 0.391 | 0.955 | 1.070 | 1.81% |
| 2,000 – 9,999 | 1,861 | 0.903 | 0.981 | 1.037 | 0.64% |
| 10,000+ | 558 | 0.962 | 0.994 | 1.018 | 0.18% |
At ten thousand bearers and up, the p10–p90 range is 0.962 to 1.018 — a span of six hundredths. At twenty to ninety-nine bearers it is 0.214 to 1.200, a span of one whole unit, and the median has fallen to 0.589, below the threshold.
Two things are happening at once and they must not be confused. Part of it is genuine composition: the tail of the Spanish register really is dominated by recently arrived surnames, so a low median down there is a fact about Spain, not noise. But the spread is noise, and it widens exactly as you would expect from small denominators. Whatever the mix, the consequence is the same: below a few hundred bearers, a single surname's ratio is not evidence about that surname.
That is why the size floor and the ratio threshold are one decision, not two. Applying c2/c1 < 0.70 to a surname with 30 bearers is arithmetic, not measurement.
Where the test is blind
Four failure modes, all of them measured on the same file, none of them fixable by moving the threshold.
1. It is blind to countries that already use two surnames. Brazil, Bolivia, Peru and most of Spanish America give their citizens two surnames at birth. Their residents arrive in Spain with a second surname already attached, so the mechanism that makes the ratio meaningful — that ap2 can only be earned by being registered here — does not apply to them at all:
| Surname | c1 | c2 | c2/c1 |
|---|---|---|---|
MAMANI | 3,341 | 4,042 | 1.210 |
HUAMAN | 2,328 | 2,614 | 1.123 |
QUISPE | 6,392 | 6,840 | 1.070 |
For our purpose this is fine — those surnames are rooted in the Spanish-language naming system, which is exactly what we were filtering for. As a measure of "how long has this surname been in Spain", it is simply wrong for them, and no threshold repairs it.
2. It rejects real Spanish surnames that happen to be small. Three surnames in the curated corpus fall below 0.70 despite being unambiguously Iberian:
| Surname | c1 | c2 | c2/c1 |
|---|---|---|---|
PASQUAL | 162 | 101 | 0.623 |
BAHAMONTES | 134 | 85 | 0.634 |
SUNYER | 495 | 326 | 0.659 |
All three have small c1, which is precisely the regime the previous section says is unreliable. The filter's job here was to admit new surnames safely, so these false rejections cost nothing — they were already in the corpus. Used as a general-purpose "is this surname Spanish" oracle, the same three would be errors.
3. It rejects real surnames that are genuinely recent. KAUR sits at 11,587 / 7,608 = 0.657, below the threshold, and it is a completely real Punjabi surname carried by thousands of Spanish residents. The filter is not saying Kaur is fake. It is saying Kaur's spelling is not governed by Castilian phonology — which is true, and which is the only question the filter was asked. Read the output as the question you posed, not as a verdict on the name.
4. It is not portable. This works because Spain publishes two columns that other registries do not have. It cannot be carried to France, Germany or the United States by analogy. What is portable is the shape of the idea: look for a column whose value can only be acquired by time spent inside the system. Spain's is the maternal surname. Brescia's municipal register does the same job with a citizenship flag — SINGH is 75% non-citizen there, FERRARI 0% — which we look at in The Registry That Does Not Exist. France's INSEE file does the inverse by excluding people born abroad, which makes the same layer visible by its absence.
One last curiosity, because it shows how far "imported" is from "foreign". Italian surnames in the Spanish register:
| Surname | c1 | c2 | c2/c1 |
|---|---|---|---|
ROSSI | 2,378 | 1,348 | 0.567 |
FERRARI | 1,532 | 774 | 0.505 |
LOCATELLI | 128 | 50 | 0.391 |
Two European surnames from a neighbouring Romance-language country, spelled in the same alphabet, with no accents to restore and nothing exotic about them — and by this measure they are as recently arrived as Traoré (5,212 / 916 = 0.176) is by a wider margin. The column is not reading language, or region, or plausibility. It is reading the registry, and only the registry.
The general lesson
Our filter was originally built out of two things that felt like knowledge: a phonological rule that was genuinely correct, and a hand-written test for whether a surname "looks Spanish" — no K, no W, no foreign consonant clusters. The rule survived. The hand-written test was worthless, and it was worthless in the most dangerous possible way: it produced confident, plausible, permanently wrong output, and nothing downstream would ever have flagged Johnsón as an error.
What replaced it was one column, one division and one threshold placed in a valley. It knows nothing about Spanish, nothing about Romanian, nothing about Punjabi. It has never heard of Huamán. It rejected 876 candidates and got every one of the twenty-odd we checked by hand right.
When a judgement can be replaced by a counter, replace it. And when the registry hands you a column nobody meant as a signal, check what it counts before deciding it is metadata.
Data as of 2026-07-21
Every figure above was recomputed on 21 July 2026 directly from the INE file and from the corpus files on disk, not quoted from working notes.
Sources:
- Spain — INE (Instituto Nacional de Estadística), Frecuencias de apellidos, Censo de población as of 01/01/2025. Columns used:
Apellido 1º / Total(c1) andApellido 2º / Total(c2). <ine.es/daco/daco42/nombyapel/…; - File scope: 86,512 surnames; sum of
c1= 46,541,135; sum ofc2= 43,595,570. The published caption states a cut-off of 100 on the first surname, but the file runs down toc1= 20. 7,757 rows carry..in thec2column instead of a count; those rows are excluded from every ratio computed here rather than treated as zero. - Accents: the INE file stores 1,832 surnames containing
Ñand zero containing an acute accent or diaeresis. That asymmetry is the reason the accent-restoration problem exists at all. Accented forms in this article are the corpus spellings;CAPITALSdenote the register's own unaccented form. - Curated reference set: 2,222 hand-spelled Spanish surnames held on disk before this work, all 2,222 of which matched the register. Used as a held-out test for both the accent rule and the threshold. The rule check reported here (136 / 140) was recomputed independently for this article rather than taken from the build log.
- Outcome of the run described: 300 surnames added, corpus 2,222 → 2,522, weight sum 34,026,331; 5,490 candidates rejected in total, of which 876 by the ratio test.
On what the ratio measures. Spain does not collect ethnicity data, by law, and nothing in this article is derived from any such source. Every figure is a count of surname positions in a civil register. The ratio cannot tell you anything about any individual, it should not be aggregated into claims about communities, and at least four distinct situations produce a low value which it cannot distinguish. It is used here for one narrow purpose: deciding whether a Castilian spelling rule may be applied to a string.