An inside view of the daily working life of a linguist interested in language description and linguistic theory. (Click on any image below to enlarge it.)
Showing posts with label thom smith stark. Show all posts
Showing posts with label thom smith stark. Show all posts
Friday, May 15, 2015
Thursday, March 28, 2013
New instances of progressive aspect in Colonial Valley Zapotec
I was very pleased to find some more instances of the progressive aspect marker in Colonial Valley Zapotec texts this morning, since it is not generally known that this morpheme was present in the language at this stage. Smith Stark (2008), for example, doesn't include any mention of it in an otherwise comprehensive overview of the Colonial Valley Zapotec aspect morphology.
These new examples involve the verb 'do' and 'be sitting'. It's interesting that with the first verb, the form in the text is
ca-g-oni
where the /g/ looks like a reflex of the potential aspect. Thom Smith Stark argued in a (2004) paper that the progressive evolved from a construction that involved a verb of position followed by a verb in the potential.
These new examples involve the verb 'do' and 'be sitting'. It's interesting that with the first verb, the form in the text is
ca-g-oni
where the /g/ looks like a reflex of the potential aspect. Thom Smith Stark argued in a (2004) paper that the progressive evolved from a construction that involved a verb of position followed by a verb in the potential.
Labels:
aspect,
colonial,
positional verb,
progressive,
thom smith stark,
Zapotec
Saturday, March 23, 2013
Central Zapotec languages
The tree below is the current version of something that I am working on. It's intended to show major branches of Central Zapotec (according to the classification of Smith Stark 2007), and it also shows (in brackets and italics) where we currently have Colonial Valley Zapotec documents in our FLEx corpus.
I know that there are many other Colonial Valley Zapotec documents out there, so the tree will get 'bushier' over time, but I found this a useful graphic view of where the documents come from in terms of distribution.
I know that there are many other Colonial Valley Zapotec documents out there, so the tree will get 'bushier' over time, but I found this a useful graphic view of where the documents come from in terms of distribution.
Labels:
central,
classification,
colonial,
Cordoba,
Feria,
FLEx,
thom smith stark,
tree,
Zapotec
Thursday, June 7, 2012
Culling blanks from the Colonial Zapotec
In the process of importing entries from Cordova dictionary of Colonial Valley Zapotec, I forgot that I needed to weed out the cases where the Zapotec Lexeme Form is blank.
This is generally because the original entry in Cordova is a cross reference. For example:
Where we don't have a Zapotec word for the entry under Cepo de animales. The underlying file from Thom Smith Stark's group has all of these listed as records with no Zapotec form.
That doesn't make much sense in a Zapotec- Spanish - English dictionary, so I weeded those out tonight. That results in a dictionary with 46,712 entries. (There are still plenty of duplicates in there also...)
Here's a screenshot with the new number:
Labels:
colonial,
Cordoba,
dictionary,
FLEx,
import,
thom smith stark,
Zapotec
Friday, June 1, 2012
Cordova import complete!
I finished importing entries from the massive Cordova dictionary of Colonial Zapotec into the FLEx project today. Lots of entries will need editing and clean-up, but it's great to have them all there finally. At this point there are about 47,000 entries in the lexicon for Colonial Valley Zapotec -- lots of them probably will need to be merged, since each entry represents a different Spanish language translation in the original (with perhaps several different entries corresponding to the same Zapotec word). But on the other hand, a lot of entries also need to be split, since more than one Zapotec word shows up in the entry.
Here's the screenshot of the last form that was imported. Note the nice big number at the bottom on the left :)
Here's the screenshot of the last form that was imported. Note the nice big number at the bottom on the left :)
Labels:
colonial,
dictionary,
FLEx,
import,
thom smith stark,
Zapotec
Wednesday, May 30, 2012
Scilicet in Cordova dictionary entries
In the Cordova dictionary of Colonial Valley Zapotec, the abbreviation scillicet means 'that is, for example'. In the Spanish orthography of the time, this is the long s followed by a period. In Thom Smith Stark's markup, this is {¥s[cilicet].}
Here is a screenshot of a sample Cordova entry that uses this.
Almoçadas means something like 'handfuls' in Spanish. The FLEx entry that shows my best guess of the correct interpretation of the Cordova entry is as follows:
I think the o in front of the word xoopa is probably Spanish o 'or'.
This example shows some of the difficulties of deciphering all the information that Cordova includes in his entries!
Here is a screenshot of a sample Cordova entry that uses this.
Almoçadas means something like 'handfuls' in Spanish. The FLEx entry that shows my best guess of the correct interpretation of the Cordova entry is as follows:
I think the o in front of the word xoopa is probably Spanish o 'or'.
This example shows some of the difficulties of deciphering all the information that Cordova includes in his entries!
Labels:
colonial,
Cordoba,
dictionary,
FLEx,
import,
thom smith stark,
Zapotec
Bulk editing Cordova Zapotec entries
Continuing on my efforts to make the enormous Cordova dictionary of colonial Zapotec.
My general goal is to have the Lexeme Form of each verb show the verb root. Cordova's usual practice was to cite a verb in the habitual aspect for the 1st person singular.
The habitual has several allomorphs, written in the following way by Cordova (with my best guess at the intended phonemic?
to+chìba-ya ticha-pitào, 'bendezir algo o consagrar' ('bless something or consecrate')
So the stem should be
chiba
with /to-/ and /-ya/ stripped away. This entry also shows that the /-ya/ is not necessarily final in the entry, since Cordova often includes a typical object along with the verb. Here the object is ticha pitào
'word of God'.
Here is the bulk replace setup screen:
And here is an example of an entry after the Bulk Replace has removed the to- prefix.
(I also changed the part of speech for all items with the to+ pattern to Verb.)
More difficult are the entries where Thom did not do the analysis. It's not correct to remove every initial to sequence, since some of the resulting items are just nouns that start with to. For example, the noun tola 'sin' shouldn't be changed to la with a to prefix.
What I tried here was searching for the pattern to...a or to...ya, then inspecting the results to make sure that the Spanish gloss seems to be a verbal form (generally cited in either the infinitive or the past participle form). I changed the part of speech for all of the good instances to Verb, and make a few manual changes to other parts of speech when I could figure it out.
After the valid instances of verbal to...(y)a were identified, I filtered the data to show only verbs and then used the same Bulk Replace method to delete the prefixes and suffixes from the Lexeme Form.
My general goal is to have the Lexeme Form of each verb show the verb root. Cordova's usual practice was to cite a verb in the habitual aspect for the 1st person singular.
The habitual has several allomorphs, written in the following way by Cordova (with my best guess at the intended phonemic?
- <to> /ru-/
- <ti> /ri-/
- <t> /r-/
The first person is usually written as
- <a> /=a/ after a consonant
- <ya> /=ya/ after a vowel
The version of the database that we have inherited from Thom Smith Stark often has these separated from the root as follows:
to+chìba-ya ticha-pitào, 'bendezir algo o consagrar' ('bless something or consecrate')
So the stem should be
chiba
with /to-/ and /-ya/ stripped away. This entry also shows that the /-ya/ is not necessarily final in the entry, since Cordova often includes a typical object along with the verb. Here the object is ticha pitào
'word of God'.
I've worked through most of the verbs in the 5000 imported Cordova entries at this point.
My first step was to copy all the information in the original entry to the Citation Form field, so that I always have the original form available. Then I word on the Lexeme Form field to a.) remove the habitual aspect prefixes b.) remove the 1st singular suffix, b.) put the information about completive and potential aspects into special fields.
The procedure uses the Bulk Edit function of FLEx, generally searching for various allomorphs of the habitual and 1sg and replacing them with nothing. This is easiest for the entries where Thom Smith Stark's analysis, where + separates the prefix , - precedes the suffix. I can search for entries with to+, ti+, t+, -a, and -ya pretty easily.
Here are some screen shots, first filtering to find all the examples of the pattern. The search uses regular expressions, so ^ means at the beginning of the record and the \ makes the following + be interpreted literally as + (not some function).
Here is the bulk replace setup screen:
What I tried here was searching for the pattern to...a or to...ya, then inspecting the results to make sure that the Spanish gloss seems to be a verbal form (generally cited in either the infinitive or the past participle form). I changed the part of speech for all of the good instances to Verb, and make a few manual changes to other parts of speech when I could figure it out.
After the valid instances of verbal to...(y)a were identified, I filtered the data to show only verbs and then used the same Bulk Replace method to delete the prefixes and suffixes from the Lexeme Form.
Labels:
colonial,
Cordoba,
dictionary,
FLEx,
thom smith stark,
Zapotec
Monday, May 14, 2012
More on Importing Cordova into the FLEx database
For many years, Thom Smith Stark and his students worked on an electronic version of the giant Cordova dictionary of Colonial Valley Zapotec. I've been working for the last few days on the process for importing this into FLEx.
I have a few different versions of the electronic Cordova. One is a MS-Access database, and the advantage of this version is that each of the multiple Zapotec words listed as the translation of the Spanish gets its own record. In the design of the MS-Access, the ID number is keyed to the entry in the Cordova dictionary and there are two linked tables. (One with entry, folio number and Spanish, the other with the Zapotec, notes, etc.) To see both you construct a MS-Access query, which looks like this:
This can be output as an Excel file. (Then from Excel =--> Sheet swiper --> FLEx.)
After the export to Excel, you need to add a row to the top with \lx over the column that will be the Lexical Form and \gn (=Gloss National) for the Spanish. The original database has a Folio field that tells you what page of the book the word is located on. I called this field \cordova page number. And for all the rows, I added a field \source with the value [Cordova import].
In order to get Sheet swiper to work properly, the \lx column needs to be first. So I reordered the Excel columns. I also sorted the Zapotec column to remove blanks. (In the original book, these are cross-references from one Spanish entry to another.)
The original database also has two Zapotec fields, one with diacritics (ZAP_COMP) and one without (ZAP). I decided not to import the version without diacritics, so I didn't put a backslash code over that column in the Excel. The result looks something like this
Then I ran Sheet Swiper, which converts the Excel file into a standard format dictionary file. That is fairly easy.
Within FLEx, you use the Lexicon | Import | Standard Format lexical data dialogue to pull this into FLEx.
You have to go through a few steps here. Most are pretty easy to understand, but a few are not completely intuitive. In the old Shoebox/Toolbox format the languages were divided into Vernacular, National, Regional, English. I had used Spanish as the equivalent of National in earlier versions, and that is why I put the \gn tag over that column. So in the dialogues, you need to tell FLEx that for this project National = Spanish.
For the various fields, you also need to tell FLEx where they will go in the entry. If it doesn't know where they go, it will put the information in a field called "Import Residue". That's okay, since it doesn't lose the information, but it is better to specify where it will go, if you know. I wanted the Folio number to go in the Source field in each entry so I specified it in that way.
Running through this all let me import the first 483 test entries from Cordova to FLEx.
Within FLEx, I also made a few global changes. Since the Zapotec form listed has all kinds of information in it, I wanted to segregate this out into other fields and work towards having just the root of the word as the Lexeme Form. But because I didn't want to lose any information, I copied all of the information from the original entry into the Citation Form field via the Bulk Edit Entries dialogues. A few examples:
The Zapotec field has information about things that are corrected somewhere (presumably in the pages of errata at the beginning.) This entry has a correction in the original. I changed the Lexical Form to reflect the correction and show the original form + correction in the Citation Form field.
The original also contains information about the prefix that a verb takes in the completive and potential aspects. (The form is cited in the habitual.) In the TSS database, this follows the field marker /cv/
This entry shows how the information got processed. I created custom fields for the completive, potential, and habitual. Then I filtered the imported entries to find /cv/ and used the "Click copy" feature in Bulk Edit to populate these fields. After the fields had been populated, I used the "Delete" feature in Bulk Edit to remove this from the Lexeme Form:
I have a few different versions of the electronic Cordova. One is a MS-Access database, and the advantage of this version is that each of the multiple Zapotec words listed as the translation of the Spanish gets its own record. In the design of the MS-Access, the ID number is keyed to the entry in the Cordova dictionary and there are two linked tables. (One with entry, folio number and Spanish, the other with the Zapotec, notes, etc.) To see both you construct a MS-Access query, which looks like this:
This can be output as an Excel file. (Then from Excel =--> Sheet swiper --> FLEx.)
After the export to Excel, you need to add a row to the top with \lx over the column that will be the Lexical Form and \gn (=Gloss National) for the Spanish. The original database has a Folio field that tells you what page of the book the word is located on. I called this field \cordova page number. And for all the rows, I added a field \source with the value [Cordova import].
In order to get Sheet swiper to work properly, the \lx column needs to be first. So I reordered the Excel columns. I also sorted the Zapotec column to remove blanks. (In the original book, these are cross-references from one Spanish entry to another.)
The original database also has two Zapotec fields, one with diacritics (ZAP_COMP) and one without (ZAP). I decided not to import the version without diacritics, so I didn't put a backslash code over that column in the Excel. The result looks something like this
Then I ran Sheet Swiper, which converts the Excel file into a standard format dictionary file. That is fairly easy.
Within FLEx, you use the Lexicon | Import | Standard Format lexical data dialogue to pull this into FLEx.
You have to go through a few steps here. Most are pretty easy to understand, but a few are not completely intuitive. In the old Shoebox/Toolbox format the languages were divided into Vernacular, National, Regional, English. I had used Spanish as the equivalent of National in earlier versions, and that is why I put the \gn tag over that column. So in the dialogues, you need to tell FLEx that for this project National = Spanish.
For the various fields, you also need to tell FLEx where they will go in the entry. If it doesn't know where they go, it will put the information in a field called "Import Residue". That's okay, since it doesn't lose the information, but it is better to specify where it will go, if you know. I wanted the Folio number to go in the Source field in each entry so I specified it in that way.
Running through this all let me import the first 483 test entries from Cordova to FLEx.
Within FLEx, I also made a few global changes. Since the Zapotec form listed has all kinds of information in it, I wanted to segregate this out into other fields and work towards having just the root of the word as the Lexeme Form. But because I didn't want to lose any information, I copied all of the information from the original entry into the Citation Form field via the Bulk Edit Entries dialogues. A few examples:
The Zapotec field has information about things that are corrected somewhere (presumably in the pages of errata at the beginning.) This entry has a correction in the original. I changed the Lexical Form to reflect the correction and show the original form + correction in the Citation Form field.
The original also contains information about the prefix that a verb takes in the completive and potential aspects. (The form is cited in the habitual.) In the TSS database, this follows the field marker /cv/
This entry shows how the information got processed. I created custom fields for the completive, potential, and habitual. Then I filtered the imported entries to find /cv/ and used the "Click copy" feature in Bulk Edit to populate these fields. After the fields had been populated, I used the "Delete" feature in Bulk Edit to remove this from the Lexeme Form:
Saturday, May 12, 2012
Working with Thom Smith Stark's material on colonial Zapotec
I've undertaken the task of seeing whether it is possible to incorporate into our project the enormous work that Thom Smith Stark and his group put into creating an electronic and searchable version of the massive Cordova (1567) dictionary of Colonial Valley Zapotec. The big problem for modern researchers is that the Cordova dictionary is only Spanish - Zapotec, so it cannot be used to read documents.
There is tremendous potential here, but one initial challenge is figuring out the various notes, abbreviations, conventions, and file formats involved.
For example, I have one set of Word documents which are a reversal of the Cordova, now alphabetized Zapotec to Spanish. Here is an image of part of one of them:
Here are my guesses about what some of the mark-up means. In the Spanish column, I think the asterisk must indicate a word for which a new entry ought to be created. So for the second word, I think this means make an entry for this Zapotec word under 'dar cuenta o razon' and also under 'razon, dar cuenta o'.
In the Zapotec, I think the | after /ti/ is showing that this is a prefix. I don't know exactly why there is also a + symbol at this point. The ÷ precedes a clitic /=a/. (Cordova's convention was to list all the verbs in the 1st person habitual.)
In Cordova's dictionary, at the entry for a verb he lists the prefixes that are used for the preterite (or completive). So when we see prt> in TSS, that means that the completive prefix is what follows. Sometimes the completive attaches to a different form of the root.
For example, with ti-bee=a quij 'fuego sacar con yslabon o assi', the prt> co+lè means that the completive is co-lèe=a quij
For comparison, here is Cordova's entry for this verb
Although Córdova writes this all as one word tibèeaquij, Thom's morphological analysis is that the /=a/ is the 1st person. So the verb must end after this, and quij is a separate word. Thus - in Thom's dictionary seems to mean 'the following is a separate word'.
Subscribe to:
Posts (Atom)





