Tuesday, August 7, 2012

Working a bit with the Advanced Search feature at Dadegan today.  I'm trying to make sure I understand all the menu options.



This search looks for the verb خوردن ' eat'  with غذا 'food' as its object.

Tuesday, July 10, 2012

Persian Dependency analyses

I've been exploring the uses of Dadegan,  a database of Persian sentences that have been manually annotated with a syntactic analysis.

Here is a fairly simple example, where we search for the valence pattern of خوردن /khordan/ 'eat'




The first line shows the valence pattern.  Pink is the subject, yellow is the object (and the  را accusative marker is highlighted), and blue is the verb.

The screenshot doesn't show this, but when you put your mouse on the example sentences, various words in Farsi show the same highlights as the valence pattern.  In the example selected, من غذا خوردهام,  the subject is من /man/ 'I', the object is غذا /qaza/ 'food', and the verb is an inflected form of خوردن.

In the dependency analysis, the arc between the final verb and the subject is labelled فاعل ٫ /fa'al/ 'subject', and the arc between the object and the verb is labelled مفعول /mafa'ul/ 'object.

Each dependency analysis also has an abstract root, labelled ریشه ٫ /rishe/ 'root'.

Here is a more complex example of a screen shot from the Dadegan project:


This is a bit hard to interpret for a non-Persian speaker!

The analysis style is Dependency Grammar, in the style of Tesnĩere.   Here is what I understand so far:

The search is for the verb رسیدن /rasidan/  'arrive' .  The results are grouped by the type of valence pattern that appears with this verb.  When you click on a sentence, you see the dependency-tree analysis of it.

The first constituent is سنگئ که زورت نمی رسد, approximately 'the stone that you cannot move (?)'.
 This is connected by an arc to the word بکن, which is the first verb of the sentence.  The arc is labelled مفعول  'object/patient', so the first constituent is the object of the clause.
The relationship between بوشه and بکن is labelled فعل یار, which I think means something like 'co-verb'.  (It seems to be literally 'friend of the verb'.) بوشه کردن /bushe kardan/ is a compound 'do a kiss';

The last two words are linked by و 'and'.  The label between 'and' and its left sister is همپايه, which means 'cohort, equal'.

The main verb is بگذر, and I think this is the imperative of the verb گذر 'pass'.  A helpful person at Wordreference suggests the translation 'If you cannot drag a stone from your way, kiss it and change your path.' 

Tuesday, July 3, 2012

Persian chrestomathy

One of the ways in which I'm trying to teach myself Farsi is by working with texts which already have an analysis available.   I stumbled on an old, but rather nice resource on Google Play (the stupid new name for Google Books).

The book was published in 1857 by Arthur Henry Bleeck.A concise grammar of the Persian language: containing dialogues, reading lessons, and a vocabulary: together with a new plan for facilitating the study of languages 

 It starts with a brief grammatical overview, but then there are 100+ pages of text/analysis.  Typically every other page has a paragraph of Persian, with translation, and every word is given in the vocabulary below and facing this paragraph.
Here is an example:



I've been adding this text to my FLEx project as a learning exercise.  Here is a little of the text in my preliminary analysis:



This aid to learning a language was called a chrestomathy;  I don't think that they are much used anymore, but they work pretty well for my purposes.

It would be nice if the Persian were in electronic format, so that I didn't have to retype it.  But it has the side advantage of making me pay more attention to the text as it goes in.  More easily accessible electronic texts don't have the advantage of an analysis...

Wednesday, June 27, 2012

Farsi language learning software

I am trying to learn Farsi this summer for one of our research projects.

I've been exploring two online language learning systems.  One is Headstart2, developed by the military for soldiers.  But it seems that anyone can register and use the system for free.

Some features for Headstart2

  • In each unit, there is an initial vocabulary introduction section, then a subsequent group of exercises in increasing difficulty that involve using this vocabulary.
  • For example, in the following screenshot, the system is helping the user learn words that describe people, such as 'old', 'fat', 'man', 'blonde', 'tall', etc.


The question at the bottom asks "Who is the short man?"   The pictures alone are often enough to answer, but the text underneath the first picture ("Amir is short") also helps, especially when it is not completely obvious from the picture.
  • Some tasks can be rather complex, as in the following screen, where you need to be able to read the names of cities, days, and times in order to find the right flight.


Some likes and dislikes:
  • You can always go back and forth in the lessons to double check vocabulary.
  • You can play sound clips of all the words
  • The system also teaches you to read the Farsi alphabet
  • The lessons get complex fairly quickly.  I think it might be a bit overwhelming for some users.
  • There is a link for a 'glossary', but this doesn't work for me.  So there is no single place that all the vocabulary is collected
  • There are tests at the end of each chapter, so you can check your progress.

The second system is Mango .  My library system has a subscription for patrons.

Mango has a simpler rotation of activities.  It is always along the lines of 

English1   ---  Farsi1
English1  ---  ?

where at the prompt, you are supposed to say the Farsi word.  The system is based on a spaced repetition model, so it is similar overall to Pimsleur, but it presents the material in both written and audio form.

Here is a sample screenshot


If you mouse over the words, you can see their transcription or play any word individually.  The play button at the left plays the whole thing.   

The general approach for each section is to present a long phrase, then break it apart into constituent parts, drill each separately (interleaving them), then recombine them into the whole phrase at the end of the section.

Likes and dislikes
  • As with Headstart2, there is no central vocabulary that collects all the words.
  • All the words are presented in Farsi orthography.  Since I know how to read this alphabet, this works for me.  But I also tried a little of the Mandarin course, and without knowing any characters, it was much, much more difficult to do the lessons.  (Essentially it then works without any readable written portion, if you exclude the mouse-over transcriptions.)
  • The variety of tasks is low.
A possible weakness of both programs -- I see drilling/repetition of words within a chapter.  But once you've left the chapter on colors (for instance), the color words don't seem to reappear in the subsequent chapters, so it would be easy to forget them.

Both systems are far advanced above what is available for the endangered languages I know.  Would it be possible for one of these authoring systems to be made available to people working on small/endangered languages?  

I know that Rosetta Stone is working with Navajo and Chitimacha on language preservation and revitalization projects.  Unfortunately, I don't have access to their product on Farsi for a comparison to the other two.

Monday, June 18, 2012

Homograph numbers as an aid to merging entries

In a previous post I mentioned the issue of merging what are separate entries in Cordova's dictionary of Colonial Valley Zapotec.

Today I experimented with using the Homograph Numbers feature in FLEx to help me locate these entries.  In the configure columns menu, one of the options is Homograph Numbers.  I chose to restrict that entries with a homograph number greater than 0.  (Ordinary entries don't have homograph numbers; they are only assigned when two entries have identical Lexeme Forms.)



This locates about 6000 entries where there is a homograph.  In this screenshot, zèni is both 'tomar en la mano' and 'tener en la mano', so these entries should be merged.

Unfortunately, that still leaves me with lots of entries to look at -- some of them should be merged and some should not, but you need to look at the translations to decide.

I think the homograph number method would also fail to catch entries that are mostly identical, but differ by the placement of an accent, e.g. if there were an entry zéni that means the same thing as zèni.

Friday, June 15, 2012

Merging entries from Cordova's colonial Zapotec dictionary

A couple of weeks ago, I managed to import all of Cordova's 1567 dictionary of Spanish into my FLEx project.

Since that book is Spanish -- Zapotec, one thing that is revealed by created a project that can sort on the Zapotec is how many times Zapotec words are listed under multiple translations in Spanish.  This gives us a much better sense of the range of meanings of the Zapotec word.

Consider the following merged entry for aba-queta 'fall', which combines information from six different Spanish - Zapotec entries

.  

Here it would not have been too difficult to see this, since all but one are listed under caer 'fall' in Spanish.

But the value is much clearer when the Spanish glosses are more various.   Consider the following entry for aate 'harden',  where the various Spanish glosses are not alphabetically adjacent.  Nor is the idea that the same word is used for freezing and hardening obvious.


Thursday, June 14, 2012

Senses of psychological verbs in Copala Triqui

It is very difficult to accurately translate all the different psychological and emotional verbs into English or Spanish, since Triqui divides the world up somewhat differently.

Case in point is the verb aráya'anj , which has a range of meanings that includes 'be amazed at something; be concerned about someone; be alarmed at something.'   Working through the large corpus with this verb allows us to find examples that show the different shades of meaning.



Here's an example of the 'be concerned' sense:


Here is an example of the 'be amazed' sense


And here is an example of the 'alarmed'  sense



The common semantics of this verb seem to include an aspect of worry/lack of knowledge about the future and concern for people affected in the future.

We also see this verb in combination with the ni'yanj psych marker to show an attitude of amazement toward someone.


Wednesday, June 13, 2012

Accidentally working with different project versions in FLEx -- how to fix the problem

I was away for a couple of weeks and my graduate student was working with our Triqui FLEx database.  She and our native speaker colleague were looking for entries where there was no example sentence and he was creating examples for these entries.

Unfortunately, when I got back and we looked at the project together, we saw that she had accidentally opened and modified an old version of the project.

At first, I thought we would just have to look for the entries modified in the old project and cut and paste them to the new, but after a bit of thought, I found a much easier way.  Since this problem is likely to arise in any collaborative FLEx project, it's possibly useful to others as well.

Since I knew the dates that she had worked with our speaker while I was away, the first step was to go to Bulk Edit and add the Date Modified field to the Column Choices.  The Restrict choice lets you select the dates that you want to look at.


Once I had done this, I could see that there were 14 entries modified during this time.

So I exported these entries from the older version of the project via LIFT.  Then I imported the same entries into the current project version.



This resulted in a little bit of duplication.  Luckily the import log that is produced showed me all the entries where this was an issue:


What I needed to then was to look at the entries for the listed conflicts.  The first one involved duplicated entries, so I used the Merge Entry feature to combine the two forms of ananj chij.  The other three conflicts involved duplicated senses, so I went to each entry and merged the senses.  (Pull-down menu to the left of the Sense label.)

It's still best to try to avoid working on different version of a project, and I don't know what the solution would be if the interlinear texts had been modified.  But if the accidental use of different versions only affects lexical entries, then
  • filtering by Date Modified, 
  • exporting entries from Project A, 
  • importing entries to Project B, and 
  • checking the Import Log for conflicts
results in a much simpler solution than cutting and pasting.

Thursday, June 7, 2012

Culling blanks from the Colonial Zapotec

In the process of importing entries from Cordova dictionary of Colonial Valley Zapotec, I forgot that I needed to weed out the cases where the Zapotec Lexeme Form is blank.

This is generally because the original entry in Cordova is a cross reference.  For example:


Where we don't have a Zapotec word for the entry under Cepo de animales.  The underlying file from Thom Smith Stark's group has all of these listed as records with no Zapotec form.

That doesn't make much sense in a Zapotec- Spanish - English dictionary, so I weeded those out tonight. That results in a dictionary with 46,712 entries.  (There are still plenty of duplicates in there also...)

Here's a screenshot with the new number:


Sunday, June 3, 2012

The mercies of the Spanish man in the 16th century

From the Zapotec Doctrina of 1567.  There are still some parts of the Zapotec analysis to be worked out,  but of interest for its horrifying gender politics!

(Click to enlarge, if you dare...)


Friday, June 1, 2012

Cordova import complete!

I finished importing entries from the massive Cordova dictionary of Colonial Zapotec into the FLEx project today.  Lots of entries will need editing and clean-up, but it's great to have them all there finally.   At this point there are about 47,000 entries in the lexicon for Colonial Valley Zapotec -- lots of them probably will need to be merged, since each entry represents a different Spanish language translation in the original (with perhaps several different entries corresponding to the same Zapotec word).  But on the other hand, a lot of entries also need to be split, since more than one Zapotec word shows up in the entry.

Here's the screenshot of the last form that was imported.  Note the nice big number at the bottom on the left :)


Wednesday, May 30, 2012

Scilicet in Cordova dictionary entries

In the Cordova dictionary of Colonial Valley Zapotec, the abbreviation scillicet means 'that is, for example'.  In the Spanish orthography of the time, this is the long s followed by a period.  In Thom Smith Stark's markup, this is {¥s[cilicet].}

Here is a screenshot of a sample Cordova entry that uses this.




Almoçadas means something like 'handfuls' in Spanish.  The FLEx entry that shows my best guess of the correct interpretation of the Cordova entry is as follows:


I think the o in front of the word xoopa is probably Spanish o 'or'.

This example shows some of the difficulties of deciphering all the information that Cordova includes in his entries!

Bulk editing Cordova Zapotec entries

Continuing on my efforts to make the enormous Cordova dictionary of colonial Zapotec.

My general goal is to have the Lexeme Form of each verb show the verb root.  Cordova's usual practice was to cite a verb in the habitual aspect for the 1st person singular.

The habitual has several allomorphs, written in the following way by Cordova (with my best guess at the intended phonemic?
  • <to> /ru-/
  • <ti> /ri-/
  • <t> /r-/
Also sometime as <te>, though I'm not sure about /re-/ as an allomorph of the habitual in modern Valley Zapotec.

The first person is usually written as 
  • <a> /=a/  after a consonant
  • <ya> /=ya/ after a vowel


The version of the database that we have inherited from Thom Smith Stark often has these separated from the root as follows:


to+chìba-ya ticha-pitào,  'bendezir algo o consagrar' ('bless something or consecrate')




So the stem should be

chiba


with /to-/ and /-ya/ stripped away.  This entry also shows that the /-ya/ is not necessarily final in the entry, since Cordova often includes a typical object along with the verb.  Here the object is ticha pitào
'word of God'.

I've worked through most of the verbs in the 5000 imported Cordova entries at this point.  

My first step was to copy all the information in the original entry to the Citation Form field, so that I always have the original form available.  Then I word on the Lexeme Form field to a.) remove the habitual aspect prefixes b.) remove the 1st singular suffix, b.) put the information about completive and potential aspects into special fields.  

The procedure uses the Bulk Edit function of FLEx, generally searching for various allomorphs of the habitual and 1sg and replacing them with nothing.  This is easiest for the entries where Thom Smith Stark's analysis, where + separates the prefix , - precedes the suffix.  I can search for entries with to+, ti+, t+, -a,  and -ya pretty easily.

Here are some screen shots, first filtering to find all the examples of the pattern.  The search uses regular expressions, so ^ means at the beginning of the record and the \ makes the following + be interpreted literally as + (not some function).

Here is the bulk replace setup screen:

And here is an example of an entry after the Bulk Replace has removed the to- prefix.


(I also changed the part of speech for all items with the to+ pattern to Verb.) More difficult are the entries where Thom did not do the analysis.   It's not correct to remove every initial to sequence, since some of the resulting items are just nouns that start with to.   For example, the noun tola  'sin'  shouldn't be changed to la with a to prefix.


What I tried here was searching for the pattern to...a or to...ya, then inspecting the results to make sure that the Spanish gloss seems to be a verbal form (generally cited in either the infinitive or the past participle form).  I changed the part of speech for all of the good instances to Verb, and make a few manual changes to other parts of speech when I could figure it out. 


After the valid instances of verbal to...(y)a  were identified, I filtered the data to show only verbs and then used the same Bulk Replace method to delete the prefixes and suffixes from the Lexeme Form.

Monday, May 14, 2012

More on Importing Cordova into the FLEx database

For many years, Thom Smith Stark and his students worked on an electronic version of the giant Cordova dictionary of Colonial Valley Zapotec.  I've been working for the last few days on the process for importing this into FLEx.

I have a few different versions of the electronic Cordova.  One is a MS-Access database, and the advantage of this version is that each of the multiple Zapotec words listed as the translation of the Spanish gets its own record.  In the design of the MS-Access, the ID number is keyed to the entry in the Cordova dictionary and there are two linked tables. (One with entry, folio number and Spanish, the other with the Zapotec, notes, etc.) To see both you construct a MS-Access query, which looks like this:



This can be output as an Excel file.  (Then from Excel =--> Sheet swiper --> FLEx.)

After the export to Excel, you need to add a row to the top with \lx over the column that will be the Lexical Form and \gn (=Gloss National) for the Spanish.  The original database has a Folio field that tells you what page of the book the word is located on.  I called this field \cordova page number. And for all the rows, I added a field \source with the value [Cordova import].

In order to get Sheet swiper to work properly, the \lx column needs to be first.  So I reordered the Excel columns.  I also sorted the Zapotec column to remove blanks.  (In the original book, these are cross-references from one Spanish entry to another.)

The original database also has two Zapotec fields, one with diacritics (ZAP_COMP) and one without (ZAP).  I decided not to import the version without diacritics, so I didn't put a backslash code over that column in the Excel.  The result looks something like this


Then I ran Sheet Swiper, which converts the Excel file into a standard format dictionary file.  That is fairly easy.


Within FLEx, you use the Lexicon | Import | Standard Format lexical data dialogue to pull this into FLEx.


You have to go through a few steps here.  Most are pretty easy to understand, but a few are not completely intuitive.  In the old Shoebox/Toolbox format the languages were divided into Vernacular, National, Regional, English.  I had used Spanish as the equivalent of National in earlier versions, and that is why I put the \gn tag over that column.  So in the dialogues, you need to tell FLEx that for this project National = Spanish.

For the various fields, you also need to tell FLEx where they will go in the entry.  If it doesn't know where they go, it will put the information in a field called "Import Residue".  That's okay, since it doesn't lose the information, but it is better to specify where it will go, if you know.  I wanted the Folio number to go in the Source field in each entry so I specified it in that way.

Running through this all let me import the first 483 test entries from Cordova to FLEx.

Within FLEx, I also made a few global changes.  Since the Zapotec form listed has all kinds of information in it, I wanted to segregate this out into other fields and work towards having just the root of the word as the Lexeme Form.  But because I didn't want to lose any information, I copied all of the information from the original entry into the Citation Form field via the Bulk Edit Entries dialogues.  A few examples:

The Zapotec field has information about things that are corrected somewhere (presumably in the pages of errata at the beginning.)  This entry has a correction in the original.  I changed the Lexical Form to reflect the correction and show the original form + correction in the Citation Form field.



The original also contains information about the prefix that a verb takes in the completive and potential aspects.  (The form is cited in the habitual.)  In the TSS database, this follows the field marker /cv/

This entry shows how the information got processed.  I created custom fields for the completive, potential, and habitual.  Then I filtered the imported entries to find /cv/ and used the "Click copy" feature in Bulk Edit to populate these fields.  After the fields had been populated, I used the "Delete" feature in Bulk Edit to remove this from the Lexeme Form:


Sunday, May 13, 2012

Two things you should listen to with attention

In this bit of the Feria doctrina, I was interested in the verb /zoba..tiyaga/, which means 'listen'.  The first part is 'set', and the second part is 'ear'.  I think these function together as a single (complex) root in modern Zapotec.  The evidence is that the subject agreement follows 'ear' only.  Since modern Valley Zapotec languages are not generally pro-drop, that's good evidence that the pronoun after 'ear' is the subject of the preceding.

So it is not 'you set your ear'  but the equivalent of 'you ear-set'.  (Zapotec-like equivalents:
Not Set=you ear=you but set-ear=you).

We could call this a kind of incorporation in modern Valley Zapotec (though confined to a small set of V + N combinations).

But in this colonial document, the two parts are separated by an adverbial element chahui.  That seems to imply that the two parts were less lexicalized as a compound 500 years ago.


I remember that Pam Munro, John Foreman, (and Aaron Sonnenschein?) talked about something like this in the colonial Valley Zapotec documents at a SSILA meeting some number of years ago.

Saturday, May 12, 2012

Working with Thom Smith Stark's material on colonial Zapotec

I've undertaken the task of seeing whether it is possible to incorporate into our project the enormous work that Thom Smith Stark and his group put into creating an electronic and searchable version of the massive Cordova (1567) dictionary of Colonial Valley Zapotec.  The big problem for modern researchers is that the Cordova dictionary is only Spanish - Zapotec, so it cannot be used to read documents.

There is tremendous potential here, but one initial challenge is figuring out the various notes, abbreviations, conventions, and file formats involved.

For example, I have one set of Word documents which are a reversal of the Cordova, now alphabetized Zapotec to Spanish.   Here is an image of part of one of them:



Here are my guesses about what some of the mark-up means.   In the Spanish column, I think the asterisk must indicate a word for which a new entry ought to be created.  So for the second word, I think this means make an entry for this Zapotec word under 'dar cuenta o razon'  and also under 'razon, dar cuenta o'.

In the Zapotec, I think the | after /ti/ is showing that this is a prefix.  I don't know exactly why there is also a + symbol at this point.   The ÷ precedes a clitic /=a/.  (Cordova's convention was to list all the verbs in the 1st person habitual.)

In Cordova's dictionary, at the entry for a verb he lists the prefixes that are used for the preterite (or completive).  So when we see prt>  in TSS, that means that the completive prefix is what follows.  Sometimes the completive attaches to a different form of the root.

For example, with ti-bee=a quij 'fuego sacar con yslabon o assi', the prt> co+lè means that the completive is co-lèe=a quij

For comparison, here is Cordova's entry for this verb


Although Córdova writes this all as one word tibèeaquij, Thom's morphological analysis is that the /=a/ is the 1st person.  So the verb must end after this, and quij is a separate word.  Thus - in Thom's dictionary seems to mean 'the following is a separate word'.


Thursday, May 10, 2012

Proceedings of CILLA V

The proceedings of the

Conference on Indigenous Languages of Latin America-V


They include my paper on Negation as raising, as well as many other interesting papers!

Tuesday, May 8, 2012

More on comparatives in Triqui

In previous posts 1 and 2, I discussed comparatives in Copala Triqui.  The following sentence shows a comparative-like structure that I haven't noticed before, which uses síj 'reach, arrive' before another verb:


I don't see or can't find another example like this in my corpus, so I will want to check with our language consultant for his thoughts.

Monday, May 7, 2012

Orthography worries in San Dionisio Zapotec

Unfortunately, the meetings in Oaxaca left me less certain, rather than more certain, about what the best practical orthography to use in Zapotec is.  I have been using an orthography which is essentially the same as that used in Isthmus and Mitla Zapotec, but the meeting made it fairly clear to me that no one agrees on what to use, especially including speakers of the various kinds of Valley Zapotec.

In my current dictionary database, I've got too many spellings floating around, possibly confusing me.  The old one that I used is the top one given, but I've now added an Americanist style phonetic field just to avoid too much confusion.

The citation field was initially composed from the old practical, but deleting tones, vowel length, and breathiness, along the lines used in the Cali chiu simplified orthography for San Lucas Quiavini Zapotec (Munro and Lillehaugen).  The other puzzling/difficult question is how to represent the difference between plain vowels and diphthongs in the practical orthography.

Finally, the difference between ʃ, ʒ, tʃ, and dʒ is a plague.  <ch> for /tʃ/ is the only simple solution.  But every other way of writing these seems to cause reading problems.  Currently, I am leaning toward a solution where /ʃ/ is <x>, /ʒ/ is /zh/, and /dʒ/ is <dx>.

I am getting indications from my speaker that she is finding the simplified orthography too simplified in some areas (the diphthong issue being a prominent difficulty in her reading).  I still don't know how to solve this issue.

Sunday, May 6, 2012

Negative focus in Colonial Zapotec

The following example shows an interesting case of a preverbal negative focus in Colonial Valley Zapotec.  This is very much like what would occur in San Dionisio Ocotepec Zapotec in the same context.

Look at the part about 'no one can rise', where we get aca ru-ti benni zoaca chapi...

Here in SDOZ, we would have ru-te'ca biiny 'no person', where the /ru-/ prefix on the negative matches the animacy of the noun that follows.