Language Technology at UiT

The Divvun and Giellatekno teams build language technology aimed at minority and indigenous languages

View GiellaLT on GitHub divvungiellatekno/giellalt.uit.no

Meeting setup

Agenda

Cf. one of the following, depending on context:

Opening, agenda review, participants

Opened at 10:32.

Present: Børre, Ciprian, Sjur, Thomas, Tomi, Trond

Absent: Maja (on study leave)

Agenda accepted as is.

Updated task status since last meeting

Børre

Ciprian

Maja

Sjur

Thomas

Tomi

Trond

Oahpa!

Nothing new. Possibly Sahka is not working. Lene would like a captcha for the Oahpa! feedback form, to avoid too much spam.

TODO

Corpus gathering

Thomas is still continuously scanning more texts.

Trond has been to the new Språkbanken meeting in Geilo. Nasjonalbiblioteket has been scanning and digitising enormous amounts of books. Everything in the public domain or free for our purposes is texts for authors dead before 1940, as well as all texts published 1990-1999, by agreement with the publishers and authors union. The free or available texts also includes all Sámi languages.

The text produced by the scanning process is automatically evaluated for OCR quality, and they are looking for ways to automatically improve the OCR recognition rate.

TODO:

Promoting Divvun

TODO:

Future plans, directions and ideas

See a separate document in plan/strat/5year.jspwiki.

Northern areas project

Trip to Syktyvkar forthcoming.

TODO:

Infrastructure

Updated corpus online

See Ciprian´s document about the corpus content in $GTPRIV/plan/corpus/oslo_corpus_update_todo.txt.

Issues:

TODO:

Corpus infra remake

TODO:

License

TODO:

Corpus interface

This depends on the infrastructure cleanup.

TODO:

Makefile + tag simplification

Problems with the proofing tools compilation, now solved.

TODO:

  1. test latest proofing tools, compare results with previous version (Tomi)
    1. there are some perl tools missing
      1. Børre (or someone) installed them
  2. write new build commands (Sjur, Tomi)
    1. make new targets in parallell to the old ones, not by remaking them
  3. when the new build infrastructure works as it should, delete the old ones (Sjur, Tomi)

General list

Meänkieli adaptions in our infrastructure.

Requirements:

Tentative task list

To accommodate future enhancements in different directions (in rough order of importance):

  1. test bench for all parts of our language technology efforts
    1. test bench enhanced, but not yet complete
  2. improve Forrest i18n support with static sites
  3. reorganise the documentation:
    1. differ between target groups
    2. get better grouping
    3. decide what to write in Forrest and what in wiki (cf. Apertium and [http://xixona.dlsi.ua.es/apertium/]) for a similar split)
    4. update/add missing parts
  4. migrate lexc lexicons to XML, splitting the task
    1. Name lexica (the Name project)
    2. Dictionaries (already in XML, task is to integrate them)
    3. At least migrate the lexc open POSes (Komi as a pilot case)
  5. change the look of the documentation web
  6. corpus content moved to Max Planck repositories? Norsk språkbank?
  7. update infrastructure to allow content-restricted spellers for special target groups

TODO:

Linguistics

North Sámi

(nothing new, see proofing bugs below)

Lule Sámi

(nothing new, see proofing bugs below)

South Sámi

TODO:

Name lexicon/risten.no infrastructure

TODO:

Dictionaries

Released:

Other things dictionary-related:

TODO:

Proofing tools

South Sámi

Beta release: June 1. We should be getting the Polderland (now Knowledge Concepts - hereafter KC) binaries in the second half of May.

TODO:

HFST- and Voikko-based proofing tools

TODO:

Testing

Updates and changes to the speller test bench result presentation: two tables combined into one.

Useful feedback.

Spelling Error Markup

TODO:

Testing open-source Norwegian spellers

Sjur has invited the open-source group to test their spell-checker using our test bench. The response has been positive, we’ll see what happens.

Speller bugs

List of bugs returned from Polderland:

Tag reordering for abbreviations have caused a lot of problems:

smj:
hr.
hr.	hr+ABBR+Acc
cand.philol.
cand.philol.	cand.philol+ABBR+N+Acc
Per
Per	Per+N+Prop+Mal+Sg+Attr

sme:
hr.
hr.	hr+N+ABBR+Acc
Per
Per	Per+N+Prop+Mal+Sg+Attr

Open issues based on test results:

sme

Version: Davvisámi, version 1.2, 2009-09-18

smj

Version: Julevsáme, version 1.2, 2009-09-20

TODO:

Hyphenator bugs

Open issues based on test results :

sme

Lexicon version: Davvisámi, version 1.2, 2009-09-18

No known issues!

smj

Lexicon version: Julevsáme, version 1.2, 2009-09-20

sma

Command to test the hyphenator:

preprocess dev/corp/pressemelding.txt | lookup bin/hyph-sma.fst | cut -f2 | \
lookup bin/hyphrules-sma.fst | grep -v '^$' | cut -f2 | uniq | see

TODO:

Installer changes

TODO:

User documentation

TODO:

1.2 release

Content:

Other

Thursday inhouse seminar

Next time suggestion list:

  1. introduction to xslt - Ciprian to start out
    1. relevant xslt issues:
    2. basic principles of xslt …
    3. sorting in xslt … (have a look at the dictionary sort xslt script)
    4. converting from one xml format to another wilt xslt (sugg: convert from DivvunGT dictionary dtd to MacDict xml)

Future seminars:

  1. XQuery
  2. More XML (needs concretisation)
  3. UML
  4. other suggestions?

Spring planning

Topics:

Dates:

Text to speech

TODO:

CAT

Autshomato has been installed for some of us. It is up and running, wit UTF-8 problems for 10.5 and not for 10.6.

Sjur will go to Karasjok to present state-of-art for Divvun, and Autshomato, with MT. We have a problem specifying what language to translate, but we got feendback from South Africa. Priority: Change Autshomato to do MT for the language pair under consideration. Deadline for this feature is

TODO:

Norsk språkbank

Sámi will form a part of Norsk språkbank. We should consider candidates for pilot content (preferably terminology or bilingual text) for a first presentation of Språkbanken.

Summer vacations

Name Dates
Børre Dates
Ciprian Dates
Maja Dates
Sjur Dates
Thom 5 weeks between 21/6-13/8
Tomi Dates
Trond Dates

Next meeting, closing

The next meeting is 18.5.2010, 09:30 Norwegian time.

The meeting was closed at 12:00.

Appendix - task lists for the next week

Boerre

Ciprian

Maja

Sjur

Thomas

Tomi

Trond