Does anyone here know if recent advances in ML are relevant at all to this kind of linguistic study? Are there attempts being made to adapt them? I would imagine there's enough of a literary corpus in Basque to compare it with various other existing languages, and maybe (based on no evidence at all, admittedly, though the recent developments with AlphaGo seem encouraging, albeit in a different direction) human researchers have missed some patterns that would reveal themselves to a well-trained neural network.
It's just that the answer to the question posed in the headline is a definite 'no'. This type of silly inquiry into "is Azeri related to Sumerian", "is Georgian related to Chinese" etc. is very common in that part of the world, and is largely pseudo-scientific at its core.
As for ML methods, I know for a fact that modern Bayesian inference techniques have been successfully applied in comparative linguistics and proto-language reconstruction.
>It's just that the answer to the question posed in the headline is a definite 'no'. This type of silly inquiry into "is Azeri related to Sumerian", "is Georgian related to Chinese" etc. is very common in that part of the world, and is largely pseudo-scientific at its core.
I should have made this clearer, but the 'question' I had in mind was not whether Basque is related to Georgian, but the more general one of whether it's related to anything at all.
I haven't heard of neural network approaches, but people are using and working on improving automated comparisons to detect relationships. A recent article discusses its application to word lists: The Potential of Automatic Word Comparison for Historical Linguistics http://journals.plos.org/plosone/article?id=10.1371/journal....
I admittedly understand very little about ML and how Google is using it in Translate, but this seems relevant.
In “Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation”, we address this challenge by extending our previous GNMT system, allowing for a single system to translate between multiple languages. Our proposed architecture requires no change in the base GNMT system, but instead uses an additional “token” at the beginning of the input sentence to specify the required target language to translate to. In addition to improving translation quality, our method also enables “Zero-Shot Translation” — translation between language pairs never seen explicitly by the system.
As a Georgian (from the country of Georgia) I was really surprised to find out that Stallman is into our folk stuff. He seems a bit too uh, cerebral for it. I believe he even visited our country around 2010 just to research the music.
Yeah, I was also really surprised. My comment was somewhat off-topic in this thread, but I couldn't resist the excuse to put the link out on HN and see if anyone else thought it was amazing.
>Typological similarities certainly exist between Basque and Georgian.
Typological similarities are not considered to be all too important for establishing a genetic relationship between languages. What's much more relevant are word cognates among the most basic words (because there's less of a tendency for languages to borrow those). And those have to be made with phonological changes in mind, meaning you can't just compare today's words but rather their forms at an earlier point in time.
It's true that cognates in basic vocabulary are very important. But the beauty of the comparative method is that you don't need the know earlier forms in order to infer cognacy (although they can provide further evidence).
IIRC, basque has a lot of grammatical parallels with both Hungarian and Finnish, both of which also have no nearby related languages, and both of which are nearly impossible to learn proficiently as a second language. I'd love to read more about these languages' histories. Are there any introductory reads on the subject?
Is Finno-Ugric a magnet for this kind of speculation? I've talked with Koreans who are adamant about the existence of a relationship between Korean and Finnish.
> Various versions included the Turkic, Mongolic, Tungusic and sometimes the Korean and Japonic languages.
I was surprised of a kind of "UEH-ll" sound that you can hear in both modern turkish (their goodbye) and Korean language (don't have an example) and wondered if they had some kind of connection. Apparently they do!
Turkish ü is similar to German umlauted u and the French u. Here the theory that Turkish is part of a big Altaic-Uralic family that may include also Korean, Mongolic and Japanese is taught at schools. IANAL but AFAIK the general view is that Altaic and Uralic are separate, and Korean and Japanese are not Altaic languages.
Note that Uralic-Altaic and the like have been the fruit of an Anatolian Turkish nationalism and tend to speculate for nobilising the language.
My favorite (even more insane) extrapolation of this is the Dene-Caucasian language family, which proposes a common origin for Basque, Caucasian languages, Yeniseian languages, Chinese, Sumerian, and Navajo!
Because it is a mathematical model, and a reasonable one. When you have p( 211 to 220 ) = .0002 as in the article you don't need to consider negative values. The power of mathematical models comes from their ability to abstract these details.
29 comments
[ 2.7 ms ] story [ 73.8 ms ] threadAs for ML methods, I know for a fact that modern Bayesian inference techniques have been successfully applied in comparative linguistics and proto-language reconstruction.
I should have made this clearer, but the 'question' I had in mind was not whether Basque is related to Georgian, but the more general one of whether it's related to anything at all.
In “Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation”, we address this challenge by extending our previous GNMT system, allowing for a single system to translate between multiple languages. Our proposed architecture requires no change in the base GNMT system, but instead uses an additional “token” at the beginning of the input sentence to specify the required target language to translate to. In addition to improving translation quality, our method also enables “Zero-Shot Translation” — translation between language pairs never seen explicitly by the system.
https://research.googleblog.com/2016/11/zero-shot-translatio...
The related research paper:
https://arxiv.org/abs/1611.04558
https://stallman.org/RMSGeorgianMusicWUOG.ogg
Typological similarities are not considered to be all too important for establishing a genetic relationship between languages. What's much more relevant are word cognates among the most basic words (because there's less of a tendency for languages to borrow those). And those have to be made with phonological changes in mind, meaning you can't just compare today's words but rather their forms at an earlier point in time.
Citation needed.
The remaining connections of Korean still fascinated me when I first heard about it: https://en.wikipedia.org/wiki/Altaic_languages
> Various versions included the Turkic, Mongolic, Tungusic and sometimes the Korean and Japonic languages.
I was surprised of a kind of "UEH-ll" sound that you can hear in both modern turkish (their goodbye) and Korean language (don't have an example) and wondered if they had some kind of connection. Apparently they do!
Note that Uralic-Altaic and the like have been the fruit of an Anatolian Turkish nationalism and tend to speculate for nobilising the language.
Oh my yes.
https://en.wikipedia.org/wiki/Den%C3%A9%E2%80%93Caucasian_la...
https://en.wikipedia.org/wiki/Den%C3%A9%E2%80%93Yeniseian_la...
How can it be a normal distribution when the values on the "x"-axis cannot be negative?