Showing posts with label import. Show all posts
Showing posts with label import. Show all posts

Monday, June 18, 2012

Homograph numbers as an aid to merging entries

In a previous post I mentioned the issue of merging what are separate entries in Cordova's dictionary of Colonial Valley Zapotec.

Today I experimented with using the Homograph Numbers feature in FLEx to help me locate these entries.  In the configure columns menu, one of the options is Homograph Numbers.  I chose to restrict that entries with a homograph number greater than 0.  (Ordinary entries don't have homograph numbers; they are only assigned when two entries have identical Lexeme Forms.)



This locates about 6000 entries where there is a homograph.  In this screenshot, zèni is both 'tomar en la mano' and 'tener en la mano', so these entries should be merged.

Unfortunately, that still leaves me with lots of entries to look at -- some of them should be merged and some should not, but you need to look at the translations to decide.

I think the homograph number method would also fail to catch entries that are mostly identical, but differ by the placement of an accent, e.g. if there were an entry zéni that means the same thing as zèni.

Friday, June 15, 2012

Merging entries from Cordova's colonial Zapotec dictionary

A couple of weeks ago, I managed to import all of Cordova's 1567 dictionary of Spanish into my FLEx project.

Since that book is Spanish -- Zapotec, one thing that is revealed by created a project that can sort on the Zapotec is how many times Zapotec words are listed under multiple translations in Spanish.  This gives us a much better sense of the range of meanings of the Zapotec word.

Consider the following merged entry for aba-queta 'fall', which combines information from six different Spanish - Zapotec entries

.  

Here it would not have been too difficult to see this, since all but one are listed under caer 'fall' in Spanish.

But the value is much clearer when the Spanish glosses are more various.   Consider the following entry for aate 'harden',  where the various Spanish glosses are not alphabetically adjacent.  Nor is the idea that the same word is used for freezing and hardening obvious.


Wednesday, June 13, 2012

Accidentally working with different project versions in FLEx -- how to fix the problem

I was away for a couple of weeks and my graduate student was working with our Triqui FLEx database.  She and our native speaker colleague were looking for entries where there was no example sentence and he was creating examples for these entries.

Unfortunately, when I got back and we looked at the project together, we saw that she had accidentally opened and modified an old version of the project.

At first, I thought we would just have to look for the entries modified in the old project and cut and paste them to the new, but after a bit of thought, I found a much easier way.  Since this problem is likely to arise in any collaborative FLEx project, it's possibly useful to others as well.

Since I knew the dates that she had worked with our speaker while I was away, the first step was to go to Bulk Edit and add the Date Modified field to the Column Choices.  The Restrict choice lets you select the dates that you want to look at.


Once I had done this, I could see that there were 14 entries modified during this time.

So I exported these entries from the older version of the project via LIFT.  Then I imported the same entries into the current project version.



This resulted in a little bit of duplication.  Luckily the import log that is produced showed me all the entries where this was an issue:


What I needed to then was to look at the entries for the listed conflicts.  The first one involved duplicated entries, so I used the Merge Entry feature to combine the two forms of ananj chij.  The other three conflicts involved duplicated senses, so I went to each entry and merged the senses.  (Pull-down menu to the left of the Sense label.)

It's still best to try to avoid working on different version of a project, and I don't know what the solution would be if the interlinear texts had been modified.  But if the accidental use of different versions only affects lexical entries, then
  • filtering by Date Modified, 
  • exporting entries from Project A, 
  • importing entries to Project B, and 
  • checking the Import Log for conflicts
results in a much simpler solution than cutting and pasting.

Thursday, June 7, 2012

Culling blanks from the Colonial Zapotec

In the process of importing entries from Cordova dictionary of Colonial Valley Zapotec, I forgot that I needed to weed out the cases where the Zapotec Lexeme Form is blank.

This is generally because the original entry in Cordova is a cross reference.  For example:


Where we don't have a Zapotec word for the entry under Cepo de animales.  The underlying file from Thom Smith Stark's group has all of these listed as records with no Zapotec form.

That doesn't make much sense in a Zapotec- Spanish - English dictionary, so I weeded those out tonight. That results in a dictionary with 46,712 entries.  (There are still plenty of duplicates in there also...)

Here's a screenshot with the new number:


Friday, June 1, 2012

Cordova import complete!

I finished importing entries from the massive Cordova dictionary of Colonial Zapotec into the FLEx project today.  Lots of entries will need editing and clean-up, but it's great to have them all there finally.   At this point there are about 47,000 entries in the lexicon for Colonial Valley Zapotec -- lots of them probably will need to be merged, since each entry represents a different Spanish language translation in the original (with perhaps several different entries corresponding to the same Zapotec word).  But on the other hand, a lot of entries also need to be split, since more than one Zapotec word shows up in the entry.

Here's the screenshot of the last form that was imported.  Note the nice big number at the bottom on the left :)


Wednesday, May 30, 2012

Scilicet in Cordova dictionary entries

In the Cordova dictionary of Colonial Valley Zapotec, the abbreviation scillicet means 'that is, for example'.  In the Spanish orthography of the time, this is the long s followed by a period.  In Thom Smith Stark's markup, this is {¥s[cilicet].}

Here is a screenshot of a sample Cordova entry that uses this.




Almoçadas means something like 'handfuls' in Spanish.  The FLEx entry that shows my best guess of the correct interpretation of the Cordova entry is as follows:


I think the o in front of the word xoopa is probably Spanish o 'or'.

This example shows some of the difficulties of deciphering all the information that Cordova includes in his entries!

Monday, May 14, 2012

More on Importing Cordova into the FLEx database

For many years, Thom Smith Stark and his students worked on an electronic version of the giant Cordova dictionary of Colonial Valley Zapotec.  I've been working for the last few days on the process for importing this into FLEx.

I have a few different versions of the electronic Cordova.  One is a MS-Access database, and the advantage of this version is that each of the multiple Zapotec words listed as the translation of the Spanish gets its own record.  In the design of the MS-Access, the ID number is keyed to the entry in the Cordova dictionary and there are two linked tables. (One with entry, folio number and Spanish, the other with the Zapotec, notes, etc.) To see both you construct a MS-Access query, which looks like this:



This can be output as an Excel file.  (Then from Excel =--> Sheet swiper --> FLEx.)

After the export to Excel, you need to add a row to the top with \lx over the column that will be the Lexical Form and \gn (=Gloss National) for the Spanish.  The original database has a Folio field that tells you what page of the book the word is located on.  I called this field \cordova page number. And for all the rows, I added a field \source with the value [Cordova import].

In order to get Sheet swiper to work properly, the \lx column needs to be first.  So I reordered the Excel columns.  I also sorted the Zapotec column to remove blanks.  (In the original book, these are cross-references from one Spanish entry to another.)

The original database also has two Zapotec fields, one with diacritics (ZAP_COMP) and one without (ZAP).  I decided not to import the version without diacritics, so I didn't put a backslash code over that column in the Excel.  The result looks something like this


Then I ran Sheet Swiper, which converts the Excel file into a standard format dictionary file.  That is fairly easy.


Within FLEx, you use the Lexicon | Import | Standard Format lexical data dialogue to pull this into FLEx.


You have to go through a few steps here.  Most are pretty easy to understand, but a few are not completely intuitive.  In the old Shoebox/Toolbox format the languages were divided into Vernacular, National, Regional, English.  I had used Spanish as the equivalent of National in earlier versions, and that is why I put the \gn tag over that column.  So in the dialogues, you need to tell FLEx that for this project National = Spanish.

For the various fields, you also need to tell FLEx where they will go in the entry.  If it doesn't know where they go, it will put the information in a field called "Import Residue".  That's okay, since it doesn't lose the information, but it is better to specify where it will go, if you know.  I wanted the Folio number to go in the Source field in each entry so I specified it in that way.

Running through this all let me import the first 483 test entries from Cordova to FLEx.

Within FLEx, I also made a few global changes.  Since the Zapotec form listed has all kinds of information in it, I wanted to segregate this out into other fields and work towards having just the root of the word as the Lexeme Form.  But because I didn't want to lose any information, I copied all of the information from the original entry into the Citation Form field via the Bulk Edit Entries dialogues.  A few examples:

The Zapotec field has information about things that are corrected somewhere (presumably in the pages of errata at the beginning.)  This entry has a correction in the original.  I changed the Lexical Form to reflect the correction and show the original form + correction in the Citation Form field.



The original also contains information about the prefix that a verb takes in the completive and potential aspects.  (The form is cited in the habitual.)  In the TSS database, this follows the field marker /cv/

This entry shows how the information got processed.  I created custom fields for the completive, potential, and habitual.  Then I filtered the imported entries to find /cv/ and used the "Click copy" feature in Bulk Edit to populate these fields.  After the fields had been populated, I used the "Delete" feature in Bulk Edit to remove this from the Lexeme Form:


Thursday, February 16, 2012

SayMore to FLEx

I've been experimenting today with SayMore, which is a way of organizing language documentation materials.

I think I'm probably typical of many people doing this work in that my materials are scattered across several locations -- on the hard drive at my office, on my laptop, in Dropbox -- and they are in several different formats (sound files, text files, some images, etc.)  One of the potential advantages of SayMore would be using it to organize links to all these things and keep the metadata about speakers, permissions, dates, formats, in a single place.  This screen, for example, has information about the various contributors and their roles.  There is a central repository for these, where you can keep background info (age, native language, contact information, permission form).  In these screen, Román Vidal López is the speaker in this audio clip.




I've also been interested in the new Annotation feature (present in the Alpha release), which provides a way to do first-pass transcription of audio and then export the annotations into a format that FLEx can read, for further analysis and correction.

To test this out today, I got an audio portion of the Address to the Triqui people and added it to SayMore.  (I did something wrong because the file comes out with the name NewEvent...)


When you have this in a format that you like, you can export it:


It ends up in a format called FLEx Interlinear XML.

You can now import this into a FLEx project.  Here is what the screen looks like after File | Import |FLExText Interlinear.


You browse to wherever the export file was located.  One little thing is that the default extension for an import file is .flextext, so at first you won't see your file.  You need to click on the FileType button and select .xml.


After you do that, (if everything is working right), you should see your transcription as an Interlinear Text in FLEx:


Overall, I was fairly pleased with the whole process.  There were one or two little glitches that may improve in the later releases:

a.) Before you do the annotation, you need to do a segmentation of the audio into little chunks.  You can then listen to these chunks at any speed you like, and they repeat until you've filled in the Transcription line to your satisfaction.   However -- if you start doing the transcription and discover you've accidentally put the segment boundary in a bad place (perhaps you cut off the last vowel), it does not seem like you can ever edit the segmentation.  At least, it wasn't obvious to me how to do this.

b.) To transcribe the Triqui properly, I needed to be able use the underscore diacritic (for low register tone).  I have a Keyman keyboard that allows me to do this properly in FLEx, but it did not seem that it would work in SayMore.  Possibly I missed some dialogue that would allow me to pick the font to a Unicode font that supports this diacritic?  Or something that allows me to pick my keyboard?

Still, these are fairly minor problems, and I think I could easily see using this for my next text transcription session.

Compared to ELAN, Praat, or Transcriber, the SayMore annotation tool is less elaborate.  (In my view ELAN is way too elaborate and unwieldy for my needs.  Praat is more a tool more suited to a phonetician's needs, and Transcriber works fine but its output does not import into FLEx in any straightforward way.)

Since my working style relies heavily on FLEx, anything that imports smoothly into that program has a huge advantage.  It is possible that if I were more focused on phonological, intonational, or gestural properties of the texts, it would be worth spending the time on something like Praat or Elan.  But since my interest is more keyed to morphosyntax and lexicon, I like a fairly light transcription tool that will help me do first-pass transcription, and SayMore looks promising for that.