Viser opslag med etiketten statistics. Vis alle opslag
Viser opslag med etiketten statistics. Vis alle opslag

lørdag den 17. januar 2015

Studying the Nahuan family with Methods of Molecular Biology


In this blog post I will give a sneak-peek on some ongoing work that I am doing which involves learning the techniques biologists use to build family trees (called phylogenies) for species or other groups of organisms and applying those techniques to understanding the relations between Nahuan dialects.

As we know, languages form "families", which basically means that a group of languages can trace their lineages back to a shared ancestor. Language families appear because most languages split into daughter languages each of which diverge from the parent in a different way. This process is what we can observe when we look at dialects of a single language. Each dialect is a variety of the language that is in the process of breaking away from the parent language and form its own language. Divergence like this happens because all languages innovate, by introducing new ways of speaking, that didn't exist in the parent language.

The same is of course what happens in biological organisms where new mutations accumulate and spread in the DNA of a population causing daughter populations to diverge from the parent population. Biologists working with molecular data analyze how these mutations can be compared and traced so that we can reconstruct the genealogy of an entire family of related organisms by grouping those organisms that share specific mutations together.

Because of each new way of speaking that arises and spreads in a community can be compared to a mutation in the DNA that spreads in a population, we can use the same methods that biologists use to trace biological phylogenies to understanding the divergence of dialects.

Nahuan Dialects:

I am in the process of understanding how all the different regional varieties (dialects) of Nahuatl are related to each other. This would help us understand where the Nahuan parent language (called proto-Nahua by linguists) originated and how it spread. In order to do this I have to find out what the shared innovations that characterize each dialect area are, and then trace them back to build a family tree. I am well underways with this project, and I presented a preliminary version of this at the Meeting of the Friends of Uto-Aztecan in Nayarit this summer (Here is a link to the paper).

But what I am doing now is mapping this information into the same kind of model that biologists use, in order to apply statistical methods to better understand the relations. Doing this allows me to generate nice graphic trees using a program called Mesquite, and this is what I want to let you peek at. 


Here is an example of how each dialect (identified with a three letter code)
is plotted into a matrix of shared innovations.
These are only some of the characters I am using.

The data is entered in binary form: each innovation present in a dialect is entered as a 1, and the lack of a specific innovation is entered as a zero (i.e. when the original form is retained). Each position of 1/0 is called "a character". By listing all the characters that I am tracking, each dialect is identified with a string of ones and zeros (called a Markov chain). And by analyzing all the characters together a phylogeny is formed. For example the dialect area of Morelos is identified with the string: 011110101 - 100000100001 – 1000000000 Where each 1 is an innovation the dialects in Morelos have, and each 0 an innovation they don't have. 


I am working with 22 distinct dialects, most of them are regional, though a few are limited to only a single community. The communities and the codes I use to identify them are: Durango/Mexicanero (DUR). Michoacan (around Pomaro) (MCH), Mexico State (around Toluca) (MEX), North Guerrero (around Coatepec Costales), Federal District (mainly Texcoco)  (DFE), North Puebla (around Huauhchinango) (NPU), Morelos (MOR), Tetelcingo, Morelos (TET), SOuth Puebla (around Tehuacan) (SPU), Zongolica, Veracruz (ZON), Tlaxcala (TLX), Western Huasteca (Hidalgo/San Luis), Eastern Huasteca (Hidalgo/Veracruz), Sierra de Puebla (SDP), Isthmus (Area north of Coatzacoalcos, Veracruz) (IST),  Pajapan, Veracruz (Town in the Isthmus area), Tabasco Nawat (TAB), Chiapas Nahuat (extinct) (CHS), Pipil Nawat of El Salvador (PIP), Central Guerrero Nahuatl (GRO), Oapan (a town in Central Guerrero) and Southern Guerrero Nahuatl (CGR). Each of these have been assigned a specific binary chain, although I still need to double check some of them in the literature to be sure I am assigning the right values, and perhaps adding more characters to the chains.

The preliminary result of my analysis is a tree that looks like this, which is interesting because it has a much more treelike branching structure than most of the previous classifications. Otherwise it does not diverge in major ways from the recent classifications by Canger (1988) or Kaufman (2001). That is partly because I have used their analyses to identify the innovations that have relevance for classification.

The bushiness comes from the fact that I have been able to identify chains of innovations in the eastern branch, where particularly the development of the negation provides a signal that I believ can can be traced along several steps. Previous classifications have been shallow bushes because they have not to the same extent tried to trace independent innovations in different areas. Canger and Dakin introduced the basic split between eastern and western branches, based on a single innovation. I have been able to add a few further innovations to the basic split. But this is about the only deep isogloss in the standard classification - the rest of it is basically a synchronic classification of the distinct areas assigning them to one or the other of the two branches. By looking at the relations between the areas within the two branches I have been able to get a bushier tree.

The tree has two main branches an Eastern and a Western branch. The Western branch is divided into a Periphery (including the dialects of Durango, Michoacan, Mexico state and North Guerrero), and a Center (including the dialects of the D.F. Morelos, North Puebla, South Puebla, Zongolica and Tlaxcala). The Eastern branch is divided into the Eastern Periphery (including the dialects of the Huasteca, the Isthmus, and Central America), and Guerrero (Central and South). The Eastern periphery is subdivided into a Huastecan and an Isthmian branch, with the Sierra de Puebla dialects in a kind of intermediate position between the two.

The next step after double and triple checking the data for the tree is to map each dialect to its geographical correlates and use a phylogeographic mapping program to calculate the probable paths that each community took to arrive to their current locations.

Meanwhile, here is the pretty tree for you to take a look at:




tirsdag den 18. november 2014

Of Statistics, Lies and Genocide: How many Nahuas lived in Morelos before the Revolution?

[This post is based on work in progress, so if you would like to cite the material or argument, please contact me by email first to get the most recent version of the argument, and my permission]

The way that statistics can be used to create reality is well known, and so is the way that censuses can be used to make inconvenient segments of the population look less significant than their actual numbers suggest. Many Native American scholars have made incredible efforts trying to create realistic estimates of indigenous American populations at different times in history (Russell Thornton's work is particularly excellent). But this endeavor is always difficult due to the challenges of finding out how well census data actually represents native populations. 

A group of Zapatista soldiers.
Pedro Lavana came from the Nahua
community of Hueyapan, Morelos.
Courtesy of the Casasola Collection.
In this post (based on a part of the history chapter of my dissertation), I look at how the indigenous Nahua population of Morelos has been counted and represented before during and after the Mexican Revolution. This is an important topic because it speaks to the question of how much indigenous involvement was a part of the Zapatista movement. Since the seminal study of John Womack "Zapata and the Mexican Revolution" (1969), the consensus has been that the Zapatista movement was primarily a mestizo agrarian rebellion. But in this post I aim to demonstrate this to be completely false. Womack based his idea of the rebellion on a misreading of census data that caused him to vastly underestimate the indigenous element of the population of Morelos in 1910. 


Womack (1969:71) dedicates but a footnote to the ethno-linguistic composition of Morelos at the turn of the 20th century, and to the question of Zapata's possible relation with the Nahuatl language. He cites a 1962 UNAM master's thesis in geography that analyses census data in Morelos from 1900 to 1930. From this work, which I have not been able to consult, he extracts the information that Nahuatl speakers only made up 9.29% of the population of Morelos at the time the Revolution broke out. He also cites Sotelo Inclán's description of Zapata traveling to the village priest in Tetelcingo to get his help in deciphering the ancient Nahuatl titles of Anenecuilco, as evidence that Zapata did not know a word of Nahuatl. He claims that when the morelenses heard Madero's statement that he would return the lands appropriated by the haciendas to the indians, they interpreted “indian” to be simply the way city people referred to the rural peasantry but that they otherwise did not recognize their state as particularly Indian.

The 1900 Mexican census did collect data about indigenous languages spoken. So far Womack is on the right track. The census questionnaire (which is available online here) provided a field with the title “Idioma nativo o lengua hablada”, the instructions to the person administrating the census stated clearly the procedure for filling out the field: 

En la columna 11 debe escribirse el nombre de la lengua nativa ó hablada comunmente, como castellano, francés, inglés, etc., ó bien el nombre del idioma indígena, como por ejemplo el mexicano ó nahuatl, el zapoteco, el otomí, el tarasco, el maya, el tzendal, el huasteco, el totonaco, etc., etc. A la persona que hable el castellano y un idioma indígena, como el otomí ó el mexicano ó cualquier otro, se le anotará de preferencia el castellano.” [In column 11 should be noted the name of the native or commonly spoken language, such as Spanish, French English etc. Or also the name of the indigenous language, such as Mexicano or Náhuatl, Zapotec, Otomí Tarascan, Maya, Tzeltal, Huastec, Totonac etc. For the person who speaks Spanish and an indigenoys language such as Otomi, Mexicano or any other, Spanish will be noted by preference. (my emphasis).]

 These instructions meant that for bilingual persons only Spanish should be noted, which in turn means that the percentage figure given for speakers of Nahuatl includes only monolingual Nahuatl speakers, whereas bilingual Nahuas (and any ethnic Nahuas who did not speak the language) are counted as Spanish speakers. In 1900 using this way of counting, the number of speakers of indigenous languages was 16,9% monolingual Nahuatl speakers. Today, there are few communities with percentages of monolingual speakers of indigenous languages as high as 16% and in those communities the vast majority of inhabitants tend to speak Nahuatl as a first language and Spanish as a second language. Towns with similar numbers of monolinguals are found in for example in the Zongolica region, where census figures today suggest that a breakdown of 10% monolinguals would correspond well to a demographic composition with 10-20% monolingual speakers of Spanish and 70-80% Spanish/Nahuatl bilinguals. Given that the state of Morelos had 161,000 inhabitants in 1900, that would suggest a composition with approximately 16,000 monolinguals, and at probably least 100,000 bilingual Nahuas in the state.

However in the 1910 census, which seems to have used the same questionnaire, for some reason the number of Nahuatl speakers in Morelos declined to 9%, only to jump back up to 14% in the 1930 census, the first one after the revolution. There is no record of any events in Morelos in the period that would have plausibly caused the Indigenous population to drop by almost 40% in this ten year period. The same abrupt jump in the reporting of indigenous people is found in most of the states in the 1910 census. This seems to suggest some kind of irregularity with the 1910 census. Probably this means that the census for practical or logistical reasons did not adequately sample the rural population at this time. In any case, the figure of 9% is an anomaly that seems to under represent Nahuatl speakers by about 5%. And at the same time, contrary to what Womack clearly believes, it does not pretend to provide the total number of Nahuatl speakers, only the number of monolingual speakers. 

This of course means that when Womack takes the percentage of monolinguals to refer to the total number of speakers he is vastly underestimating the number of Nahuatl speakers of Morelos. And in contrast to his glib assertion that there were hardly any Indians there, we would be justified in considering at least 70% of the population of 161,000 people to have been Nahuatl speakers. This suggests that contrary to Womack's assertion, it is quite likely that Zapata spoke Nahuatl, and the eyewitness testimony of Doña Luz Jiménez which Womack also ignores, corroborates that he did. 

In the 1930 census the questionnaire gave the possibility of recording two languages, first whether the respondent spoke the national language or not, and then in the second slot which other language they spoke. This means that for 1930 the figure of 14% Nahuatl speakers includes both monolingual and bilingual speakers. The total population of Morelos in 1930 was 130,000, 30,000 less than before the Revolution. Based on the percentages of Nahuatl speakers we can estimate the indigenous population of Morelos at ca. 100,000 in 1910 (possibly more, including both bi- and monolingual speakers), and we can show that after the war it had been reduced to less than 20,000 (also including both mono and bilinguals).

Given the relatively modest decline in the total population from 1910 to 1930 this figure of an 80% indigenous population loss may seem exaggerated. But the population loss is hidden in the censuses because they don't take into account the influx of out-of-state people in the 7 years following the Revolution. The fact that indigenous population loss was much greater than what the raw population figure suggests is also shown by cohort analyses that show that the people counted in 1910 are not the same as the ones counted in 1930. For example of the 90,000 women counted in Morelos in 1910 only 35,000 were counted again in 1930 (McCaa 2003). This points to a drastic decline in native born (mostly Nahuatl speaking) Morelenses and their replacement of people from other states after the war. The argument could be further supported if the portion of Morelos residents born out of state could be shown to have increased drastically from 1910 to 1930, but unfortunately I have not been able to find this piece of information in the census even though the census did ask for state of birth.

This is a clear example of how census data can be used to mask what was essentially a genocidal event, and to mask the participation of indigenous peoples in National history.

*Womack, J. (1969). Zapata and the Mexican revolution. Random House LLC.
*McCaa, R. (2003). Missing millions: the demographic costs of the Mexican revolution. Mexican Studies19(2), 392-93