Showing posts with label dictionary. Show all posts
Showing posts with label dictionary. Show all posts

Sunday, October 29, 2017

New lexical entries

Here are two of the most recent lexical entries.  I am rather proud of having figured out (most) of what these words mean.



The second of these words, pileno, seems to refer to a food which would be considered a delicacy for a feast.  Maybe something like chicken livers?

Friday, January 31, 2014

Combining Cordova entries for more explanatory value in a Colonial Valley Zapotec dictionary.

Below are all the verb roots with the shape chi(j)ll{a|e} in the 1567 Cordova dictionary:



Spanish gloss
Cordova original
Location
Abrir carta cerrada
tochìllea, tixàllea
f4
Abrir la mano vazia
Tochille ñaaya
f4
Abrir los braços este[n]die[n]dolos.
Tochilleñaya, tocigàañaaya.
f4
Adiuinar por suertes o agueros Tochillaya, tibeea. {vel.} tebeea pijci. /cv/ col.
f10
Agudo de ingenio nachìlla f15
Agudo ser assi tichìllalachi
f15
Apartar casados o a dos que estan juntos
Tochilléa, tochillaquiçòa, til láaya. /prt/ co
f32
apartar casados o a dos que estan juntos Tochilléa, tochillaquiçòa, til láaya. /prt/ co
f32
Bullicioso sin sosiego, o trauiesso nachílla
f62
Casados apartar Tochillaya f74
Casamientos sortear antiguamente el sortilegio para ver si eran para en vno tochillaya
f74
Crucificar a otro tochilleñaaya làni
f99
Cuydado tener o estar con el de lo q[ue] te[n]go de hazer Tichìllalàchia
f102
Derribar edificio o casa tochillea. {[ve]l.} tochijllaya
f119
Desarrollar
Tochíllea
f121
Desarrugado ser Tichílle
f121
Desarrugar lo arrugado
Tochíllea f121
Desassosegado assi nachíllaláchi
f121
Desassosegado ser o estar, o enfermo assi tichílla
f121
Desassosegamiento nachíllalàchi f121
Desbaratada ser ge[n]te assi tichílla
f121
Desbaratar algo como edificio o casa tochíllaya
f121
Descasarse ellos [los casados]
Tichíllea
f122
Desembuelto ser [serle quitada la envoltura] Tichíllea
f125
Desemparejado ser tichillea
f125
Desemparejado ser tichillaya
f125
Desemparejar dos cosas
tochillea
f125
Desempegar
tochillaya
f125
Desemperezar tichíllalàchia f125
Deshecho ser assi [casa o edificio tichílla
f127
Deslizarse assi delas manos Tichille
f129
Desmandarse en hablar chillaya f129v
Despartido ser tichillea
f131
Despartido ser assi alos que esta[n] juntos tichillaya f131
despartir o dividir tochillaya f131
Despedirse el que se parte tochillaya ticha làoa çàa-ya. {[ve]l.} teòchillaya f132
Despedirse el que se parte Tochillaya ticha çàaya
f132
Distinguir o apartar vno de otro tochillea
f142
Enconada cosa assi nachillaquij
f162
Enconado ser assi tichillaquija
f162
Enconado ser assi [herida, llaga o hinchazon] Tichillaquija, /prt/ pi
f162
Estenderme en luengo o lo encogido Tichillea
f189
Despartir o diuidir Tochijlea, f131
Estender otra cosa o ami mesmo Tochijlea f189
Estendida cosa assi. s. el braço o pierna Nachijleñaa
f189
Estender el braço. s. estendido ser tichijleñaaya, tigáanaaya, tilijñaaya.
f189
Descasar los casados tochijllaya f122
Sortear echar suerte Tochijllaya f387
Deshazer casa o edifficio Tochijllaya
f127
Descasarse ellos [los casados] tichijllaya f122
Suerte hechar tochijllaya f390
Despoblarle [al pueblo] tochijllea
f133
Desdoblado ser Tichijllea f125
Despartir ruydo o renzilla tochijllea f131
Desplegados ser assi [quitados los pliegues] Tichijllea f133
Destechada ser assi [casa]. Tichijllea f134
Despartido ser assi alos que esta[n] juntos. Tichijllea f131
Desenhetrar como los cabellos. Tochijllea
f126
Descasar los casados Tochijlléa f122
Entesar Tochijllea
f173
Desdoblar Tochijllea
f125
Tender lo encogido Tochijllea, tochijllaya f396
Estender la mano vazia. Tochijlleláchiñaaya. {[ve]l.} ñaaya {solum.}
f189
Agudamente huachillalàchi
f15
Despoblador huechijlle
f113
Estendedor Huechijlle
f189
Agudo o dilige[n]te. chilla f15
diligente Nachijlla, naciña, natiti, naquèñe, nanij. f140
Diuidir Tochillea, tochij-çòoa f143







In our dictionary project, we are trying to extract the roots of the verbs and make lexical entries that reflect the range of meanings and orthographic representations of these verbs.

Here are the entries for chilla (2) 'divide (intr)', chilla (3) 'decide by chance or casting lots', and chilla (4) 'divide (trans)'.




These entries are still in draft form.

Sunday, September 1, 2013

Creating an audio dictionary with Webonary -- A web version of the Copala Triqui dictionary

For the last several months, I've been working on a web version of the Copala Triqui dictionary, hosted by the kind people at webonary.org  (part of SIL).  They provided the space for the page and the audio files, and Philip was very helpful in giving me advice.

The current version is available at http://copalatriqui.webonary.org/  (This is a preliminary draft, and not everything works correctly.  Comments welcome!)

Steps in getting this to work:
1) Within FLEx, you output XHTML.  There are a few issues about how to set up the dictionary configuration to deal with multiple writing systems.  The hardest thing was working to get all the audio files to display.

The current configuration of the dictionary prior to export is as in the following entry:

Here the names of the audio files for the lemma and the examples come directly after these elements of the entry.

Because I had originally recorded all the audio in .wav format, I needed to convert all the .wav files to .mp3 format for the web.

If you use the Send/Receive function, then normally the audio files will be located at the following place (substituting the name of your own project):



To hold the file names for the new .mp3 files, I created two custom fields -- Audio for Web and Audio for Web Examples.

The configuration for the Dictionary looks like this, with the Audio fields directly after the Headword and the Example:



2.) To upload the audio, you use a special entry to webonary.org via the AjaXplorer interface.  The interface looks like this:

To get the files onto the system, you use the upload button, and then you see an interface like the following:


3.) To convert the .wav to .mp3, I used Audacity.   You can select a group of files in Windows Explorer and drag them onto the Audacity window.  This has the effect of opening them all at once.


Then you can use the File | Export Multiple command to convert bunches of them to .mp3



4.)  The main XHTML output from FLEx needs to be uploaded in Webonary.  After you log in, you go to the Dashboard, then to Tools | Import | SIL FLEx dictionary, where you see a screen like the following:


Here you use the Choose File button to select the XHTML files generated by the Export function of FLEx.

5.) After doing all this, you should get a web page with audio links.  It is necessary to do a certain amount of editing of your webonary dictionary to get the title, menus, etc. to accurately reflect the name of this dictionary project (and not some other Webonary dictionary for a different language).

Here's a picture of the browse screen for entries beginning with m:


Thursday, May 16, 2013

New draft of the San Dionisio Ocotepec Zapotec dictionary posted

I'm more and more adopting the philosophy that we shouldn't keep our data buried in our field notes and computers until it is perfect.  It's more likely to be useful to others if we make it available in intermediate draft stages, so I'm starting to post draft versions of various projects on my page at Academia.edu.   (Inspired by the wisdom of Peter Austin, who posts draft grammars and dictionaries on his page!)

Today I've posted the current, imperfect state of the San Dionisio Ocotepec Zapotec dictionary. (Ethnologue ZTU) at http://www.academia.edu/3546909/San_Dionisio_Ocotepec_Zapotec_--_Spanish_--_English_dictionary_interim_draft_version_

Here's a screenshot of one part of it:



Thursday, November 8, 2012

Phrase-book material in a lexicon

Native speakers or learners of a language often find it useful to have a set of common phrases in the language.  But do these go in the dictionary?

It seems awkward to have entries like the following:


But an alternative is to use the Publications field in FLEx to define a separate Phrasebook.  You configure this via Lists | Publications and add a new Publication called Phrasebook. The default publication is Main Dictionary, so if you haven't changed anything, that is where all the entries will appear.

The default seems to be that all entries are in all publications.   I used the Bulk Edit commands to remove Phrasebook from the Publication lists of all the entry, then went back in and selected the entries I wanted for the Phrasebook.

For this entry, this means adjusting the Publications field to say only Phrasebook and not Main Dictionary.  After changing the field, it now looks like this:


In the lexicon pane, we can pick which dictionary to display.  The default will be the main dictionary:


From the pull-down arrow next to the words Main Dictionary Entries, we can also select other publications, such as the Phrasebook.  The following screen shows a partially populated Phrasebook:


Monday, June 18, 2012

Homograph numbers as an aid to merging entries

In a previous post I mentioned the issue of merging what are separate entries in Cordova's dictionary of Colonial Valley Zapotec.

Today I experimented with using the Homograph Numbers feature in FLEx to help me locate these entries.  In the configure columns menu, one of the options is Homograph Numbers.  I chose to restrict that entries with a homograph number greater than 0.  (Ordinary entries don't have homograph numbers; they are only assigned when two entries have identical Lexeme Forms.)



This locates about 6000 entries where there is a homograph.  In this screenshot, zèni is both 'tomar en la mano' and 'tener en la mano', so these entries should be merged.

Unfortunately, that still leaves me with lots of entries to look at -- some of them should be merged and some should not, but you need to look at the translations to decide.

I think the homograph number method would also fail to catch entries that are mostly identical, but differ by the placement of an accent, e.g. if there were an entry zéni that means the same thing as zèni.

Friday, June 15, 2012

Merging entries from Cordova's colonial Zapotec dictionary

A couple of weeks ago, I managed to import all of Cordova's 1567 dictionary of Spanish into my FLEx project.

Since that book is Spanish -- Zapotec, one thing that is revealed by created a project that can sort on the Zapotec is how many times Zapotec words are listed under multiple translations in Spanish.  This gives us a much better sense of the range of meanings of the Zapotec word.

Consider the following merged entry for aba-queta 'fall', which combines information from six different Spanish - Zapotec entries

.  

Here it would not have been too difficult to see this, since all but one are listed under caer 'fall' in Spanish.

But the value is much clearer when the Spanish glosses are more various.   Consider the following entry for aate 'harden',  where the various Spanish glosses are not alphabetically adjacent.  Nor is the idea that the same word is used for freezing and hardening obvious.


Thursday, June 7, 2012

Culling blanks from the Colonial Zapotec

In the process of importing entries from Cordova dictionary of Colonial Valley Zapotec, I forgot that I needed to weed out the cases where the Zapotec Lexeme Form is blank.

This is generally because the original entry in Cordova is a cross reference.  For example:


Where we don't have a Zapotec word for the entry under Cepo de animales.  The underlying file from Thom Smith Stark's group has all of these listed as records with no Zapotec form.

That doesn't make much sense in a Zapotec- Spanish - English dictionary, so I weeded those out tonight. That results in a dictionary with 46,712 entries.  (There are still plenty of duplicates in there also...)

Here's a screenshot with the new number:


Friday, June 1, 2012

Cordova import complete!

I finished importing entries from the massive Cordova dictionary of Colonial Zapotec into the FLEx project today.  Lots of entries will need editing and clean-up, but it's great to have them all there finally.   At this point there are about 47,000 entries in the lexicon for Colonial Valley Zapotec -- lots of them probably will need to be merged, since each entry represents a different Spanish language translation in the original (with perhaps several different entries corresponding to the same Zapotec word).  But on the other hand, a lot of entries also need to be split, since more than one Zapotec word shows up in the entry.

Here's the screenshot of the last form that was imported.  Note the nice big number at the bottom on the left :)


Wednesday, May 30, 2012

Scilicet in Cordova dictionary entries

In the Cordova dictionary of Colonial Valley Zapotec, the abbreviation scillicet means 'that is, for example'.  In the Spanish orthography of the time, this is the long s followed by a period.  In Thom Smith Stark's markup, this is {¥s[cilicet].}

Here is a screenshot of a sample Cordova entry that uses this.




Almoçadas means something like 'handfuls' in Spanish.  The FLEx entry that shows my best guess of the correct interpretation of the Cordova entry is as follows:


I think the o in front of the word xoopa is probably Spanish o 'or'.

This example shows some of the difficulties of deciphering all the information that Cordova includes in his entries!

Bulk editing Cordova Zapotec entries

Continuing on my efforts to make the enormous Cordova dictionary of colonial Zapotec.

My general goal is to have the Lexeme Form of each verb show the verb root.  Cordova's usual practice was to cite a verb in the habitual aspect for the 1st person singular.

The habitual has several allomorphs, written in the following way by Cordova (with my best guess at the intended phonemic?
  • <to> /ru-/
  • <ti> /ri-/
  • <t> /r-/
Also sometime as <te>, though I'm not sure about /re-/ as an allomorph of the habitual in modern Valley Zapotec.

The first person is usually written as 
  • <a> /=a/  after a consonant
  • <ya> /=ya/ after a vowel


The version of the database that we have inherited from Thom Smith Stark often has these separated from the root as follows:


to+chìba-ya ticha-pitào,  'bendezir algo o consagrar' ('bless something or consecrate')




So the stem should be

chiba


with /to-/ and /-ya/ stripped away.  This entry also shows that the /-ya/ is not necessarily final in the entry, since Cordova often includes a typical object along with the verb.  Here the object is ticha pitào
'word of God'.

I've worked through most of the verbs in the 5000 imported Cordova entries at this point.  

My first step was to copy all the information in the original entry to the Citation Form field, so that I always have the original form available.  Then I word on the Lexeme Form field to a.) remove the habitual aspect prefixes b.) remove the 1st singular suffix, b.) put the information about completive and potential aspects into special fields.  

The procedure uses the Bulk Edit function of FLEx, generally searching for various allomorphs of the habitual and 1sg and replacing them with nothing.  This is easiest for the entries where Thom Smith Stark's analysis, where + separates the prefix , - precedes the suffix.  I can search for entries with to+, ti+, t+, -a,  and -ya pretty easily.

Here are some screen shots, first filtering to find all the examples of the pattern.  The search uses regular expressions, so ^ means at the beginning of the record and the \ makes the following + be interpreted literally as + (not some function).

Here is the bulk replace setup screen:

And here is an example of an entry after the Bulk Replace has removed the to- prefix.


(I also changed the part of speech for all items with the to+ pattern to Verb.) More difficult are the entries where Thom did not do the analysis.   It's not correct to remove every initial to sequence, since some of the resulting items are just nouns that start with to.   For example, the noun tola  'sin'  shouldn't be changed to la with a to prefix.


What I tried here was searching for the pattern to...a or to...ya, then inspecting the results to make sure that the Spanish gloss seems to be a verbal form (generally cited in either the infinitive or the past participle form).  I changed the part of speech for all of the good instances to Verb, and make a few manual changes to other parts of speech when I could figure it out. 


After the valid instances of verbal to...(y)a  were identified, I filtered the data to show only verbs and then used the same Bulk Replace method to delete the prefixes and suffixes from the Lexeme Form.

Monday, May 14, 2012

More on Importing Cordova into the FLEx database

For many years, Thom Smith Stark and his students worked on an electronic version of the giant Cordova dictionary of Colonial Valley Zapotec.  I've been working for the last few days on the process for importing this into FLEx.

I have a few different versions of the electronic Cordova.  One is a MS-Access database, and the advantage of this version is that each of the multiple Zapotec words listed as the translation of the Spanish gets its own record.  In the design of the MS-Access, the ID number is keyed to the entry in the Cordova dictionary and there are two linked tables. (One with entry, folio number and Spanish, the other with the Zapotec, notes, etc.) To see both you construct a MS-Access query, which looks like this:



This can be output as an Excel file.  (Then from Excel =--> Sheet swiper --> FLEx.)

After the export to Excel, you need to add a row to the top with \lx over the column that will be the Lexical Form and \gn (=Gloss National) for the Spanish.  The original database has a Folio field that tells you what page of the book the word is located on.  I called this field \cordova page number. And for all the rows, I added a field \source with the value [Cordova import].

In order to get Sheet swiper to work properly, the \lx column needs to be first.  So I reordered the Excel columns.  I also sorted the Zapotec column to remove blanks.  (In the original book, these are cross-references from one Spanish entry to another.)

The original database also has two Zapotec fields, one with diacritics (ZAP_COMP) and one without (ZAP).  I decided not to import the version without diacritics, so I didn't put a backslash code over that column in the Excel.  The result looks something like this


Then I ran Sheet Swiper, which converts the Excel file into a standard format dictionary file.  That is fairly easy.


Within FLEx, you use the Lexicon | Import | Standard Format lexical data dialogue to pull this into FLEx.


You have to go through a few steps here.  Most are pretty easy to understand, but a few are not completely intuitive.  In the old Shoebox/Toolbox format the languages were divided into Vernacular, National, Regional, English.  I had used Spanish as the equivalent of National in earlier versions, and that is why I put the \gn tag over that column.  So in the dialogues, you need to tell FLEx that for this project National = Spanish.

For the various fields, you also need to tell FLEx where they will go in the entry.  If it doesn't know where they go, it will put the information in a field called "Import Residue".  That's okay, since it doesn't lose the information, but it is better to specify where it will go, if you know.  I wanted the Folio number to go in the Source field in each entry so I specified it in that way.

Running through this all let me import the first 483 test entries from Cordova to FLEx.

Within FLEx, I also made a few global changes.  Since the Zapotec form listed has all kinds of information in it, I wanted to segregate this out into other fields and work towards having just the root of the word as the Lexeme Form.  But because I didn't want to lose any information, I copied all of the information from the original entry into the Citation Form field via the Bulk Edit Entries dialogues.  A few examples:

The Zapotec field has information about things that are corrected somewhere (presumably in the pages of errata at the beginning.)  This entry has a correction in the original.  I changed the Lexical Form to reflect the correction and show the original form + correction in the Citation Form field.



The original also contains information about the prefix that a verb takes in the completive and potential aspects.  (The form is cited in the habitual.)  In the TSS database, this follows the field marker /cv/

This entry shows how the information got processed.  I created custom fields for the completive, potential, and habitual.  Then I filtered the imported entries to find /cv/ and used the "Click copy" feature in Bulk Edit to populate these fields.  After the fields had been populated, I used the "Delete" feature in Bulk Edit to remove this from the Lexeme Form: