Repository navigation
Add a feature to work with Hispanic and Portuguese names (latinos) #103
Description
Activity
The current parser logic basically takes the first name it gets and sticks it in the first name list, then sticks all other names into the middle name list until it reaches the last name.
It seems feasible that the logic could be changed to switch from middle names to last names at some other point as it proceeds from beginning to end, but if there can be an arbitrary number of both middle and last names how would the parser know when to switch to last names?
Also if middle names are not really a separate concept like in English, maybe the surnames attribute is what you want:
https://nameparser.readthedocs.io/en/latest/modules.html#nameparser.parser.HumanName.surnames
It's basically all names except the first name or any titles or suffixes, so middle and last concated together.
In the "Cristiano Ronaldo dos Santos Aveiro" example, what would be a useful way to separate out the name parts in Portuguese? You referred to "second" and "surname". I know in Portuguese people can have like 10 names, but I don't know what's a useful way to group them. If more names were added would you want "third" name and then "surname" for the just final name part? Would a numerical index be useful?
Also might want to look at #72 for some related conversation. I think that's when I added the surnames attribute.
I know in Portuguese people can have like 10 names, but I don't know what's a useful way to group them.
yes, same in Spanish. Maybe a good way would be to have a field called
names(similar to surnames) that would include all the names, in the case of "Cristiano Ronaldo dos Santos Aveiro" it would be "Cristiano Ronaldo". In Spanish people can also have multiple names, an example:"Dr. Miguel Ángel González-Fierro Palacios":
- title: Dr
- names: Miguel Ángel
- surnames: González-Fierro Palacios
Is this ready to use
Much of what's requested here is now supported, and the remainder is a deterministic-ambiguity limit worth being explicit about.
The common Spanish case (one given name + multiple surnames) works with the opt-in
middle_name_as_last(#194):>>> from nameparser import HumanName >>> from nameparser.config import Constants >>> HumanName("Rafael Nadal Parera", Constants(middle_name_as_last=True)) # first="Rafael", last="Nadal Parera"
And you can already access the pieces without any flag:
surnames/surnames_list— all surnames (middle + last), e.g."Nadal Parera"given_names/given_names_list— all given names, first + middle (feat: add given_names attribute (closes #157) #180)
What isn't solvable deterministically is a name with multiple given names AND multiple surnames — e.g. your
"Cristiano Ronaldo dos Santos Aveiro"(Cristiano + Ronaldo given, the rest surnames). The parser has no way to know thatRonaldois a second given name rather than a first surname —"Rafael Nadal Parera"has the identical shape but the opposite split. So no positional rule can place that boundary correctly for both.The only partial signal is that Iberian surnames often begin with a particle (
de/da/dos/del), which the parser already keeps together — sodos Santosstays intact — but particle-less surnames (Nadal,Parera,Garcia) offer no such cue.Closing since the tractable pieces are in and the aggregate accessors cover the "several names" use case. If you have a specific dataset pattern in mind (e.g. a surname-particle boundary heuristic as an opt-in), a focused issue for that would be welcome.
The library doesn't work for latino names, in Spanish and Portuguese we have 2 surnames, ie:
Spanish: Rafael Nadal Parera
Portuguese: Cristiano Ronaldo dos Santos Aveiro (in this case Cristiano is the first name, Ronaldo de second name and the rest surnames)
Also, we can have several names, they are not called first and middle name, but fist, second, third names...
I guess it not an easy thing to implement, but there are around 700M people on that group :-)