Showing posts with label lexicography. Show all posts
Showing posts with label lexicography. Show all posts

Friday, May 17, 2013

Possible output of Colonial Valley Zapotec dictionary

Maybe it is because I have been working with output formats for Copala Triqui and San Dionisio Ocotepec Zapotec dictionaries lately, but I spent some time today thinking about what an eventual output format for a Colonial Valley Zapotec dictionary might look like.

This is one draft of a format to consider, where the entries have the form:

CVZ orthography  (part of speech) English gloss Spanish gloss (Examples) {Cordova page numbers, references to other sources such as Whitecotton} {Cordova original entry/entries} (alternate spellings)

Here is one entry fleshed out with examples, with several Cordova entries merged together …



I also put the different languages in different colors to make it easier to distinguish the different parts of the entry from each other.  Examples here are chosen to illustrate some of the range of different morphology that the root chono appears with.  I pulled them from the corpus using the 'Find Example' feature in FLEx and edited them for length.  I also italicized the word or phrase that contains chono and the corresponding part of the translation to make it easier for the reader.

(One thing I can see already is that I have been inconsistent in the abbreviations for the names of the texts -- sometimes Feria and sometimes Doctrina!)

A posted version of a dictionary like this would allow users who don't have access to the full FLEx project to nevertheless search through it and be able to advance over what is currently available to them.  (Of course, I think what we have still needs lots of clean up before we would have a version that would be ready for posting.)

Thursday, May 16, 2013

New draft of the San Dionisio Ocotepec Zapotec dictionary posted

I'm more and more adopting the philosophy that we shouldn't keep our data buried in our field notes and computers until it is perfect.  It's more likely to be useful to others if we make it available in intermediate draft stages, so I'm starting to post draft versions of various projects on my page at Academia.edu.   (Inspired by the wisdom of Peter Austin, who posts draft grammars and dictionaries on his page!)

Today I've posted the current, imperfect state of the San Dionisio Ocotepec Zapotec dictionary. (Ethnologue ZTU) at http://www.academia.edu/3546909/San_Dionisio_Ocotepec_Zapotec_--_Spanish_--_English_dictionary_interim_draft_version_

Here's a screenshot of one part of it:



Thursday, December 20, 2012

'Goody goody' in Trique

One of my favorite lexical entries in a while is the difficult to translate interjection arsínj, said when a bad person, or a person who has rejected your good advice, gets into trouble.  The closest we could come up with in English was 'Goody! Goody!'



Here is an English expression of the same sentiment.

Thursday, September 6, 2012

Puzzling verb in Copala Triqui

It puzzles me that the verb achríj rá seems to mean both 'realize' and 'suspect'.  These have different properties of factivity in English -- the complement of 'realize' must be a fact, while the complement of 'suspect' does not have to be.  The same seems to be true for Spanish darse cuenta and sospechar.

We explained this to our speaker, who understands the difference between the English and Spanish, but still thinks the Triqui might mean both.  It's interesting that the Triqui lexicon divides up the semantic space of cognition differently that English.

Here's our current (imperfect!) entry for achríj rá:


Wednesday, June 13, 2012

Accidentally working with different project versions in FLEx -- how to fix the problem

I was away for a couple of weeks and my graduate student was working with our Triqui FLEx database.  She and our native speaker colleague were looking for entries where there was no example sentence and he was creating examples for these entries.

Unfortunately, when I got back and we looked at the project together, we saw that she had accidentally opened and modified an old version of the project.

At first, I thought we would just have to look for the entries modified in the old project and cut and paste them to the new, but after a bit of thought, I found a much easier way.  Since this problem is likely to arise in any collaborative FLEx project, it's possibly useful to others as well.

Since I knew the dates that she had worked with our speaker while I was away, the first step was to go to Bulk Edit and add the Date Modified field to the Column Choices.  The Restrict choice lets you select the dates that you want to look at.


Once I had done this, I could see that there were 14 entries modified during this time.

So I exported these entries from the older version of the project via LIFT.  Then I imported the same entries into the current project version.



This resulted in a little bit of duplication.  Luckily the import log that is produced showed me all the entries where this was an issue:


What I needed to then was to look at the entries for the listed conflicts.  The first one involved duplicated entries, so I used the Merge Entry feature to combine the two forms of ananj chij.  The other three conflicts involved duplicated senses, so I went to each entry and merged the senses.  (Pull-down menu to the left of the Sense label.)

It's still best to try to avoid working on different version of a project, and I don't know what the solution would be if the interlinear texts had been modified.  But if the accidental use of different versions only affects lexical entries, then
  • filtering by Date Modified, 
  • exporting entries from Project A, 
  • importing entries to Project B, and 
  • checking the Import Log for conflicts
results in a much simpler solution than cutting and pasting.

Tuesday, January 31, 2012

Root vs stem in FLEx

My friend Danny asked me a question about morpheme types in FLEx that prompted some thought.  FLEx allows you to specify a pretty wide range of morpheme types, including a distinction between roots and stems.

The documentation resources that come with FLEx describe the distinction in the following terms:


I can imagine that there are projects where the root/stem distinction has some practical application, but I have not found much use for it my own work.

I don't use the automatic parsing function of FLEx, so I do manual parsing instead.  Possibly something in the set-up of project with morphological rules might want the rules to be sensitive to whether affixes attach to roots vs stems.

Apart from that, it has seemed to me like I have been able to do pretty much everything I need without really using the distinction very much.

For my own purposes, I use 'stem' for nearly everything, and 'compound' for most of the rest and 'bound stem' for a small number of morphemes that appear in compounds but not independently.  The possible advantage of 'stem' as the default, is that you can use the 'Components' field to show when stems contain other morphemes.

These two screen shots show the Triqui verbs achén 'pass' and its causative tacuachén 'deliver'.  They show that tacuachén has achén as a component and achén is part of a complex form tacuachén.




(The default order of the fields in the dictionary puts components right after the lemma.  I don't much like this, so I think I would edit the dictionary configuration to put them closer to the end.)


This isn't a problem with the design of FLEx, but one of practical implications for a person in language documentation is that sometimes the software gives you the ability to make more distinctions that are really needed for a dictionary and text collection.  At least it gives me more options than I seem to need so far...