Mine as well! Icelandic is an interesting language. Icelandic (at least in written form) is mutually intelligible with Old Norse, which means that knowing Icelandic can allow you to read 1000 year old Norse sagas. Super cool, if you ask me.
Keep in mind that the similarity to old Norse was intentional. It was a form of linguistic nationalism to revert some of the dialect changes that had built up over the centuries during icelandic independence. They started teaching kids using the old books and used them to standardize the orthography. The vocal changes were a bit more sticky though.
An interesting tidbit (hinted at in the article) is that there’s an active effort to reuse old (no longer used) words for new meanings. An example is sími (telephone), an old word for cord/rope. I wonder how well GPT-4 copes with the issue of loanword vs coining new/reusing old terms.
That’s really interesting as the as the word ’sim’ in modern farsi (persian) also mean cord or line.
It seems according to wiktionary (https://en.m.wiktionary.org/wiki/s%C3%ADma#Icelandic) that the Word has Proto-Indo-European roots which is the common ancestor of farsi and European languages.
I wonder if it’s possible to introspect or query the model and find out about the etymologi of words. I suspect more context like the phonetic version of each word is needed to make that final connection not just the semantic embedding.
I suspect (though I may be wrong) that GPT-4 wouldn't provide more insight than what's in the text it was trained on. If prompted by examples of regular sound change or reconstruction of proto-indo-european (PIE) roots, it may be able deduce other PIE word forms.
This made me want to look up LLM projects on minority indegenous languages, makes me wonder what kinds of effects it would have on places like the Philippines with many different dialects and languages that aren't all mutually intelligible and need a common lingua franca.
As someone who works in and with one, likely not. Languages are inherently physical things, and they need to be passed on in the real world. If the real world still works through [insert regional majority language] that's what's get passed on due to pure economic pressure.
I haven't looked at GPT-4 yet, but for Irish the previous versions were bad. I expect 4 to be bad too simply because most the training data (Irish used online) is quite shit. But I don't see how this really changes anything about the situation on the ground with minority languages and intergenerational transmission.
We still lack insight into the nature of GPT's hallucinations.
Without the supervision of an expert human, we don't know when these hallucinations are happening.
Given that, is this really the right tool to "preserve" icelandic especially since there are probably fewer experts available to judge the results?
Related(?), I would expect that an LLM would be very useful in helping people learn a conlang [constructed language] such as Interlingua or Esperanto, since the LLM might come closer to a "native" proficiency than any human tutor.
Or even so-called "dead" languages. I'm looking forward to the day where there will be a good locally hosted LLM-based assistant for Latin… both for learning and for helping to formulate "native" quality texts.
esperanto is easy enough to learn that you don't need a tutor with native proficiency. and you if you practice enough you can easily achieve proficiency in less than a year.
the grammar is simple so the main work is picking up the vocabulary and the rest is conversational practice and learning about edge cases
OpenAI is partnered with https://mideind.is/english.html to help train GPT-4. I wonder how that partnership came to fruition.
I guess this explains it.
On the initiative of the country’s President, HE Guðni Th. Jóhannesson, and with the help of private industry, Iceland has partnered with OpenAI to use GPT-4 in the preservation effort of the Icelandic language —and to turn a defensive position into an opportunity to innovate.
Not just that, but there has been a public data gathering effort[0] running now for a few years, specifically with development of speech and text recognition in mind. According to national news today, when OpenAI was approached about this project last year, they were offered this data to train their models on.
Are there any analyses how GPT fares with adapting to regional dialects, or i.e british vs. american spellings? I am asking because the article uses computadora (instead of ordenador) as example for an adapted (only latin) spanish word.
(It's as always, you notice one mistake and it makes you skeptical about the entire article, even though one shouls know better)
18 comments
[ 1.6 ms ] story [ 57.3 ms ] threadhttps://en.wikipedia.org/wiki/Icelandic_language
I wonder if it’s possible to introspect or query the model and find out about the etymologi of words. I suspect more context like the phonetic version of each word is needed to make that final connection not just the semantic embedding.
I haven't looked at GPT-4 yet, but for Irish the previous versions were bad. I expect 4 to be bad too simply because most the training data (Irish used online) is quite shit. But I don't see how this really changes anything about the situation on the ground with minority languages and intergenerational transmission.
the grammar is simple so the main work is picking up the vocabulary and the rest is conversational practice and learning about edge cases
I guess this explains it.
On the initiative of the country’s President, HE Guðni Th. Jóhannesson, and with the help of private industry, Iceland has partnered with OpenAI to use GPT-4 in the preservation effort of the Icelandic language —and to turn a defensive position into an opportunity to innovate.
[0] https://almannaromur.is/
(It's as always, you notice one mistake and it makes you skeptical about the entire article, even though one shouls know better)