Language Technology at UiT

The Divvun and Giellatekno teams build language technology aimed at minority and indigenous languages

View GiellaLT on GitHub divvungiellatekno/giellalt.uit.no

Meeting setup

Agenda

Cf. one of the following, depending on context:

Opening, agenda review, participants

Opened at 13:45.

Present: Børre, Ciprian, Lene, Maja, Sjur, Thomas, Tomi

Absent: Trond

Agenda accepted as is.

Updated task status since last meeting

Børre

Ciprian

Maja

Sjur

Thomas

Tomi

Trond

Oahpa!

Started to update official Oahpa! with Finnish Leksa, and changed the code to also get a Swedish, English and German versions.

We need a meeting to find a way to easily get the Leksa word lists for new languages.

Lene: Cooperation with teacher in Finland to create a Finnish version of Oahpa!.

Meeting memos can be found at [http://giellatekno.uit.no/ped/index.html#Meeting+memos]

TODO

Corpus gathering

Maja has sent several letters to Julie Eira to get them into the official letter template.

Børre is working on converting all corpus files to the new svn repository. He is also planning a trip to Karasjok to meet the Ávvir people (Per Christian Biti). He will try to get the rest of the old-format articles, as well as discussing how we can get the articles in their present platform.

Lene will get the Riddu riddu newspaper (in InDesign format? that is ok.)

TODO:

Promoting Divvun

TODO:

Future plans, directions and ideas

See a separate document in plan/strat/5year.jspwiki.

Northern areas project

TODO:

Infrastructure

Corpus infra remake

TODO:

License

TODO:

Corpus interface

TODO:

Makefile + tag simplification

TODO:

  1. test that the output from the new transducers is identical to the old one (Tomi)
  2. write new build commands (Sjur, Tomi)
  3. when the new build infrastructure works as it should, delete the old ones (Sjur, Tomi)

General list

Trond will write an e-mail to the fit group explaining our situation.

To accommodate future enhancements in different directions (in rough order of importance):

  1. test bench for all parts of our language technology efforts
    1. test bench enhanced, but not yet complete
  2. improve Forrest i18n support with static sites
  3. reorganise the documentation:
    1. differ between target groups
    2. get better grouping
    3. decide what to write in Forrest and what in wiki (cf. Apertium and [http://xixona.dlsi.ua.es/apertium/]) for a similar split)
    4. update/add missing parts
  4. migrate lexc lexicons to XML, splitting the task
    1. Name lexica (the Name project)
    2. Dictionaries (already in XML, task is to integrate them)
    3. At least migrate the lexc open POSes (Komi as a pilot case)
  5. change the look of the documentation web
  6. corpus content moved to Max Planck repositories? Norsk språkbank?
  7. update infrastructure to allow content-restricted spellers for special target groups

TODO:

Linguistics

North Sámi

(nothing new, see proofing bugs below)

Lule Sámi

(nothing new, see proofing bugs below)

South Sámi

What about possessives in sma? Needs to be checked/corrected, they seems to be in use in print, at least.

TODO:

Name lexicon/risten.no infrastructure

TODO:

Dictionaries

TODO:

Proofing tools

South Sámi

TODO:

HFST-based proofing tools

The work with Voikko+HFST is moving forward.

Testing

Spelling Error Markup

TODO:

Speller testing

TODO:

Testing open-source Norwegian spellers

Sjur has invited the open-source group to test their spell-checker using our test bench. The response has been positive, we’ll see what happens.

We should go to their developer meetings, and present our work and how to work with language technology.

Speller bugs

List of bugs returned from Polderland:

Tag reordering for abbreviations have caused a lot of problems:

smj:
hr.
hr.	hr+ABBR+Acc
cand.philol.
cand.philol.	cand.philol+ABBR+N+Acc
Per
Per	Per+N+Prop+Mal+Sg+Attr

sme:
hr.
hr.	hr+N+ABBR+Acc
Per
Per	Per+N+Prop+Mal+Sg+Attr

Open issues based on test results:

sme

Version: Davvisámi, version 1.2, 2009-09-18

smj

Version: Julevsáme, version 1.2, 2009-09-20

TODO:

Hyphenator bugs

Open issues based on test results :

sme

Lexicon version: Davvisámi, version 1.2, 2009-09-18

No known issues!

smj

Lexicon version: Julevsáme, version 1.2, 2009-09-20

sma

Command to test the hyphenator:

preprocess dev/corp/pressemelding.txt | lookup bin/hyph-sma.fst | cut -f2 | \
lookup bin/hyphrules-sma.fst | grep -v '^$' | cut -f2 | uniq | see

TODO:

Installer changes

TODO:

User documentation

TODO:

1.2 release

Content:

Other

Thursday inhouse seminar

A short (less than 1h) seminar every Thursday at 10 AM. Possible topics:

Spring planning

Topics:

Dates:

Text to speech

TODO:

CAT

A-ITE is working, and seems to do its job ok. The testing has involved too little text to really test the TM functions, but simple text matching seems to work. Also Apertium integration is possible, making it possible to use MT directly within A-ITE.

TODO:

Next meeting, closing

The next meeting is 12.04.2010, 13:00 Norwegian time.

The meeting was closed at 14:41.

Appendix - task lists for the next week

Boerre

Ciprian

Maja

Sjur

Thomas

Tomi

Trond