Interesting project. I would love if you could switch fonts to something like New Athena Unicode.
I built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
I would love to learn more / collab (for personal use) if there's a way to connect with you. I'm on proton mail if you want to reach out, henrystafford77.
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.
Suggestion: For the pop-up text, bold the meaning of the word so it pops out a bit better. Due to the formatting of word definitions, you have to go hunting for it in many cases.
Clicking out to close a dictionary popup only works if you click empty space within the layout center. This is annoying. On a wide monitor, most empty space is on the sides (where clicking doesn't close).
It seems some books are missing chapter markings. Revelation, for instance. The verses are numbered, and you can notice when chapters change, but one can be lost when searching for a specific chapter.
Dictionary entries are cool, but I would want the in context meaning at least highlighted, so I don't have to read the full entry and do that myself.
Overall, nice idea but it seems like a very barebones implementation that needs a tremendous amount of polishing to be useful.
I built myself such a simple lexicon for technical stuff (concurrency vocab - invariant, genetics, stuff like that) - can only recom the practice as vocabulary is clearly a big step difference
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
I'm always surprised to see that a large enough portion of the HN community are interested in classics that posts like this make it to the front page. Who are you all? What are you doing here?
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
Philosophy and CS double major here, love reading old books. Read a lot of theology and ancient literature. Homeschooling my kids, part of the curriculum is learning Latin.
In some countries classics are still widely taught in middle school/high school (as in about 10% of the population taking up Latin). Where I was brought up (Flanders) I'd guess over 50% of people with an advanced degree took up Latin, maybe 5-10% Greek. In certain circles not taking up Latin and/or Greek means you were either not smart enough to do so or you were uncultured. So anyone "cultured" with a CS/engineering degree took up Latin.
I generated a read-along version of Athenaze read by a real native language Greek speaker and I met some of the problems that this web has: the vast amount of vocabulary makes managing a dictionary rather complicated.
For tbos version, I gather that a bilingual presentation would be more than necessary: keep Ancient Greek text on a side and display a scholar translation into a selection of switcheable languages
The Greek font is terrible—it’s using the Modern Greek “acute” accent which looks weird, and commas aren’t supposed to have spaced before them. It’s cool that Perseus did the work that makes this possible though. More info about what edition we’re looking at would be nice too.
An interesting project. Ancient languages have always fascinated me, but the effort required to read even a single page of Greek or Latin can be a serious obstacle
It's nice to see tools that simplify this process. I might even want to give it another try
I'm a bit out of energy at present but I'll try to put together a comment.
From the point of view of a learner it's unfortunate that macrons (diacritical marks for long vowels) have not been added to the text, and that 'u's have not been altered to 'v's where appropriate.
The word-by-word treatment of translation makes this broadly a part of the category of interlinear texts. For comparison here's one of the early pages of Max Müller's 1864 Sanskrit-to-English interlinear of the https://en.wikipedia.org/wiki/Hitopadesha : https://archive.org/details/firstbookofhitop00ml/page/2/mode... . This is part of a revival of interlinear texts, as a resource for language learners, which traces back to John Locke but really caught on in the English-speaking world in the early nineteenth century thanks to James Hamilton. It's unusual for "traditional" interlinears in having one row which gives a(n explicit) morphological gloss of each word in the original (the "da, 3 sg. Pres. Par" and so on) in the same column as the original word, an English translation of the word, and (in this case) a transliteration. But morphological glosses are a standard feature of modern "interlinear morphemic glosses" https://www.christianlehmann.eu/ling/ling_meth/ling_descript... made by and for linguists who want to analyse and compare languages, rather than learn them. Linguists' adoption of modern IMGs seems to have taken off in the 1960s (though some use has been made of interlinears for "serious" linguistic purposes since at least the 1890s beginnings of the Linguistic Survey of India https://en.wikipedia.org/wiki/Linguistic_Survey_of_India ).
I tried to find a way to read the entire translation of a given work but I found only being able to highlight a word to see it's meaning, not the entire translation. Maybe I missed how to get that from this site? Or is there a site that provides translations of such ancient works?
I love this. I try to read different works of philosophy by getting chatgpt live to narrate it paragraph by paragraph to fully digest it. But newest works it refuses to do that unfortunately. It would be quite amazing to do a second pass on the original Latin works in the same way as well.
I checked out the page on Caesar's The Gallic War and looked at its famous opening line "Gallia est omnis diuisa in partes tres , quarum unam incolunt Belgae , aliam Aquitani , tertiam qui ipsorum lingua Celtae , nostra Galli appellantur".
Apart from the UI around the word parsing pop-up being kind of buggy (it doesn't close readily, the scroll position seems to get saved differently depending on the word), I'm pretty sure there's an analytical error in that sentence.
The word "nostra" is glossed as Nom Plur Neut; that is, the possessive adjective "noster", "ours", in its nominative plural neuter form. "Nostra", with a short final -a, is indeed the nominative plural neuter form of that possessive adjective. However in that sentence, the word is being used as an ablative singular feminine, with a long final -â (which is unmarked because the text as they present it doesn't seem to mark vowel length at all).
archive.org has this (https://archive.org/details/caesarscommentar07caes/page/n11/...) 1918 out-of-copyright literal word-by-word translation of the text, which does mark vowel length, and you can see that the translator words the last clause as "the third, (those) who in (the) language of themselves are called Celtae, in ours, Gauls". nostrâ is modifying an implied linguâ there, which is a singular feminine noun in the ablative case, so nostrâ has to inflect to agree with it - nostrâ [linguâ] translates to "in ours [our language]".
My Latin is self-taught and mediocre at best, but even I was able to spot that error. And The Gallic War is a very famous text that is routinely taught fairly early on in traditional Latin education - a lot of students of Latin over the centuries have read that exact line of Caesar's. The fact that the LatinCy neural language model for Latin, which the site claims performed the analysis, screwed this up doesn't leave me very confident about the quality of the grammatical analysis in more difficult corners of Latin grammar.
I host a parallel translation of the original book on Peano’s Axioms. If you have comments on either the Latin or English translation, they’d be very welcome!
(My 2 years of high school Latin from a Latin-Mass priest is helpful, but insufficient.)
This is very close to a problem I have also been working on. I built Kevilex for Ancient Greek readers: https://kevilex.com
It also provides one-click definitions and grammar analysis. The main difference from this project is that it remembers your vocabulary across texts. You can mark words as new, learning, or known. It then recommends texts based on the vocabulary you already know.
It also includes flashcard review, sentence explanations, text import, and an Old English library.
35 comments
[ 0.22 ms ] story [ 45.0 ms ] threadI built something similar to this by cloning the Diogenes repo and getting Claude to re-implement it in Python (it’s a very old battle tested Perl code base, so a great reference implementation) and using the TLG database for the Greek and Latin texts. You can take this even further by integrating it with the Barrington Atlas (there are scans on Anna’s Archive) for looking up ancient place names, so you have dictionary + map lookups. If you’ve read any of the Landmark series books you’d known what I mean.
Also better if you can generate chapter by chapter critical apparatus on difficult grammar and Anki decks. It’s an annoying part of learning these languages to have to stop and look up words, I like spending a few days learning vocab before reading and it makes it so much more pleasant than having to stop and look things up the whole time.
Obviously all this stuff is copyright so it can never be shared, but I don’t care it’s for my own personal use. I also bought like 6 hours worth of Ionnis Strattakis’ recordings where he reads a bunch of Ancient Greek in reconstructed Attic pronunciation and fine-tuned a text to speech model (styletts2) with full accent markings and breathings. It’s extremely natural sounding. My long term goal is to have a personal tutor that I can speak Attic to and basically have lessons with everyday (speech to text -> llm-> text to speech). All the pieces are there to actually do this.
LLMs have been an absolute game changer for me who is a hobby classicist but also knows how to build software. They are so amazingly good at Attic Greek and Latin, it’s like having the best teachers in the word at your fingertips for some really niche topics that it wouldn’t be possible with otherwise. Also extremely good at managing, building, cleaning up, deduplicating Anki decks.
I don't get a definition for 'fulgere', third word first entence, just a reference to 'fulgo'. I can guess what it means though from the more common 'fulgur', maybe something adjacent to flashing or lightning, but translations seem a bit sketchy.
It seems some books are missing chapter markings. Revelation, for instance. The verses are numbered, and you can notice when chapters change, but one can be lost when searching for a specific chapter.
Dictionary entries are cool, but I would want the in context meaning at least highlighted, so I don't have to read the full entry and do that myself.
Overall, nice idea but it seems like a very barebones implementation that needs a tremendous amount of polishing to be useful.
The way the Greek is displayed isn't very helpful, though. Specifically, any vowel with a grave accent displays the accent as a separate letter, which makes it very distracting to read. I don't think it's a problem with the encoding, as I can copy the inline text and it displays just fine - though with some weird spaces before commas or periods:
ΕΝ ΑΡΧΗ ἦν ὁ λόγος , καὶ ὁ λόγος ἦν πρὸς τὸν θεόν , καὶ θεὸς ἦν ὁ λόγος .
My undergrad was in classics and viticulture, then did grad school in atmospheric science, which led me to tech and hackernews.
How did you get here?
For tbos version, I gather that a bilingual presentation would be more than necessary: keep Ancient Greek text on a side and display a scholar translation into a selection of switcheable languages
From the point of view of a learner it's unfortunate that macrons (diacritical marks for long vowels) have not been added to the text, and that 'u's have not been altered to 'v's where appropriate.
The word-by-word treatment of translation makes this broadly a part of the category of interlinear texts. For comparison here's one of the early pages of Max Müller's 1864 Sanskrit-to-English interlinear of the https://en.wikipedia.org/wiki/Hitopadesha : https://archive.org/details/firstbookofhitop00ml/page/2/mode... . This is part of a revival of interlinear texts, as a resource for language learners, which traces back to John Locke but really caught on in the English-speaking world in the early nineteenth century thanks to James Hamilton. It's unusual for "traditional" interlinears in having one row which gives a(n explicit) morphological gloss of each word in the original (the "da, 3 sg. Pres. Par" and so on) in the same column as the original word, an English translation of the word, and (in this case) a transliteration. But morphological glosses are a standard feature of modern "interlinear morphemic glosses" https://www.christianlehmann.eu/ling/ling_meth/ling_descript... made by and for linguists who want to analyse and compare languages, rather than learn them. Linguists' adoption of modern IMGs seems to have taken off in the 1960s (though some use has been made of interlinears for "serious" linguistic purposes since at least the 1890s beginnings of the Linguistic Survey of India https://en.wikipedia.org/wiki/Linguistic_Survey_of_India ).
https://theamericanscholar.org/the-new-old-way-of-learning-l... is a very incomplete history of interlinears (the French are especially shortchanged) but probably still the best thing out there in English. See also https://www.reddit.com/r/interlinear . (I have to release some things about the history of interlinears myself but I'm years late at this point.)
https://shop.hyplern.com (might as well add an affiliate link! https://invi.tt/N5VZXP7g ) and https://interlinearbooks.com/ are two publishers offering modern interlinears aimed at language learners. The polished but expensive Legentibus service https://legentibus.com/ offers some interlinears too.
Guess who Vincent F. Hopper https://archive.org/details/chaucerscanterbu0000chau_k0e4 was for a time married to!
Apart from the UI around the word parsing pop-up being kind of buggy (it doesn't close readily, the scroll position seems to get saved differently depending on the word), I'm pretty sure there's an analytical error in that sentence.
The word "nostra" is glossed as Nom Plur Neut; that is, the possessive adjective "noster", "ours", in its nominative plural neuter form. "Nostra", with a short final -a, is indeed the nominative plural neuter form of that possessive adjective. However in that sentence, the word is being used as an ablative singular feminine, with a long final -â (which is unmarked because the text as they present it doesn't seem to mark vowel length at all).
archive.org has this (https://archive.org/details/caesarscommentar07caes/page/n11/...) 1918 out-of-copyright literal word-by-word translation of the text, which does mark vowel length, and you can see that the translator words the last clause as "the third, (those) who in (the) language of themselves are called Celtae, in ours, Gauls". nostrâ is modifying an implied linguâ there, which is a singular feminine noun in the ablative case, so nostrâ has to inflect to agree with it - nostrâ [linguâ] translates to "in ours [our language]".
My Latin is self-taught and mediocre at best, but even I was able to spot that error. And The Gallic War is a very famous text that is routinely taught fairly early on in traditional Latin education - a lot of students of Latin over the centuries have read that exact line of Caesar's. The fact that the LatinCy neural language model for Latin, which the site claims performed the analysis, screwed this up doesn't leave me very confident about the quality of the grammatical analysis in more difficult corners of Latin grammar.
(My 2 years of high school Latin from a Latin-Mass priest is helpful, but insufficient.)
https://github.com/mdnahas/Peano_Book/
It also provides one-click definitions and grammar analysis. The main difference from this project is that it remembers your vocabulary across texts. You can mark words as new, learning, or known. It then recommends texts based on the vocabulary you already know.
It also includes flashcard review, sentence explanations, text import, and an Old English library.