Garbage in… Greek out? Experiments with Deepseek using OCR’d Italian containing embedded Greek.

The letters of the 6th century sophist Aeneas of Gaza have been sitting in a folder on my desktop for a month or two now, and I want to make some progress with making a translation into English.

It’s not a big text.  Each letter is only a short paragraph, and there are only twenty-five letters.  So the whole text would fill less than a dozen pages perhaps.  I have the 1962 edition with Italian translation by Lidia Massa Positano, which is more than a hundred pages.  There seems to be a rule that editions of tiny texts can be obese!  I have never forgotten the Gerlo edition of Tertullian’s “De Pallio” – a very short text of a page or two – which filled two lengthy volumes.  But Gerlo published in 1940 in the Netherlands, and it may have been expedient for him to be engaged in such a project at that time.  Anyway the Positano edition is not that daft, and consists of an introduction, the Greek text, the Italian translation of each letter with commentary and footnotes.

At some earlier point I seem to have run the Positano book through Abbyy Finereader 15 software to create a Word document, which is 77k in size.  So today I extracted the portion to do with the letters.  The OCR language was Italian, so the portions of the commentary that contained quotes from the Greek were gibberish at those points.

Anche Procopio scrive lettere riguardanti prestiti di libri. Per esempio, dall’ep. LXIII, diretta a Pizio, risulta che Procopio si rammarica di non possedere un libro chiestogli in prestito da co­stui. Nell’ep. CHI, diretta a Stefano, Procopio vivamente lo rim­provera di un grave ritardo nella restituzione di un libro (p. 572,46 : èpuO-piàv cpiQaec^ w; napa^à; tVjv òrtoo/eaiv xaì xà? auv^xaj uirepi- Swv xal tò pi^Xfov i/wv Tpkov v) ré-capTov toutI ó p^SèrpiTOV xaftéÉEiv è7raYYeiXàp,Evo<;) e cita come esempio di solerzia e di zelo proprio un Giovanni (p. 573,12 àXX’ où/ 3 y£ aocpuiTaxo; ’IwàvvTj;

The printed text is as follows:

I thought, as a first step, that I would load the Word document into ChatGPT and get an AI translation of the lot, commentary, footnotes etc.  The idea was only to allow me to skim the material, and decide what to focus on.

But ChatGPT started complaining that it was a long file, a very very long file, etc etc:

The document you uploaded is very large (roughly 130,000 characters / 60+ pages). Translating the entire file accurately would be too long for a single response.

Then it offered various ways to make my life difficult if I carried on.

Well, I didn’t.  I popped it into Deepseek instead.  My prompt was “Translate this Italian text into English” – nothing exotic.   Deepseek made no objections and speedily output an English version of the whole thing.  I did have to copy and paste the output to Word, but that was not too burdensome.

But as I copied and pasted, I noticed something strange.  My eye was drawn to the gibberish sections of the Greek.  Here is the same passage, converted from exactly that gibberish above:

Procopius also writes letters concerning loans of books. For example, from ep. LXIII, addressed to Pyzius, it appears that Procopius regrets not possessing a book requested from him by the latter. In ep. CIII, addressed to Stephen, Procopius severely reproaches him for a serious delay in returning a book (p. 572,46: ἐρυθριᾷν φῂς ὡς παρελθὼν τὴν ὑπόσχεσιν καὶ τὰς συνθήκας ὑπερβὰς καὶ τὸ ἐμὸν χρέος τεταρτεῦσαι τοῦθ᾽ ὁ μηδὲ τρίτον καθέξειν ἐπαγγειλάμενος) and cites as an example of diligence and zeal a certain John (p. 573,12 ἀλλ᾽ οὐχ ὁ γε σοφώτατος Ἰωάννης…

It has recognised that the text is Greek – I did not tell it so – and it produced that accented Greek output.  The Greek is not actually completely correct.  But it’s very close!

In computing, there is no magic.  If it looks like magic, it means only that you don’t understand what is going on.  The input that I gave it was not enough to produce that Greek output.  It was just the attempts of an OCR engine to make Italian out of Greek.  So the Greek was retrieved from elsewhere, and the garbage string used to search for it.

I then tried the following prompt with a page of the garbled Greek:

This comes from a book in Italian, displaying Greek. Correct the Greek. ”’….”’

This it proceeded to do, with an interesting commentary underneath:

So the garbled text is being used for a look-up of some sort.

We know that the databases in these “AI” engines – the Large Language Models (LLMs) – are essentially a search database made by pirating vast amounts of books and everything else. In this case it is perhaps using the garbled strings to look up stuff in the Thesaurus Linguae Graecae database.

But, as ever with AI, you just cannot trust what it gives you.  You have to check, and checking can take longer than doing it yourself.

Interesting, and frustrating.  As ever!

5 thoughts on “Garbage in… Greek out? Experiments with Deepseek using OCR’d Italian containing embedded Greek.

  1. Have you tried giving the image directly to the AI and asking for a transcription? I just told Gemini “transcribe and translate the text on this image” + copypaste the image, and I got:

    Anche Procopio scrive lettere riguardanti prestiti di libri. Per esempio, dall’ep. LXIII, diretta a Pizio, risulta che Procopio si rammarica di non possedere un libro chiestogli in prestito da costui. Nell’ep. CIII, diretta a Stefano, Procopio vivamente lo rimprovera di un grave ritardo nella restituzione di un libro (p. 572,46: ἐρυθριᾶν φήσεις ὡς παραβὰς τὴν ὑπόσχεσιν καὶ τὰς συνθήκας ὑπεριδὼν καὶ τὸ βιβλίον ἔχων τρίτον ἢ τέταρτον ἔτος τουτὶ ὃ μηδὲ τρίτον μῆνα καθέξειν ἐπαγγειλάμενος) e cita come esempio di solerzia e di zelo proprio un Giovanni (p. 573,12 ἀλλ’ οὐχ ὅ γε σοφώτατος Ἰωάννης
    Procopius also writes letters regarding book loans. For example, from letter LXIII, addressed to Pithius, it appears that Procopio regrets not owning a book requested from him as a loan by the latter. In letter CIII, addressed to Stephen, Procopio sharply rebukes him for a serious delay in returning a book (p. 572,46: “you will say you blush, having violated your promise and disregarded the agreements, keeping the book this third or fourth year, when you promised not to hold it even for a third month”) and cites a certain John as an example of diligence and zeal (p. 573,12 “but not so the most wise John…”).

    I don’t see any errors in the transcription. I guess I should have told it to keep the Greek as is. In any case this seems a more direct way of getting results from the AI.

    On the other hand, I once had a very negative experience using the AI as an OCR engine. I fed it a page in English that I needed to transcribe. The OCR looked perfect at first, with no missing or extraneous characters. Only after a while of working with the transcription I realized that it had hallucinated whole words and phrases on me, even though the image was perfectly legible. I had to proofread the whole thing, spending even longer than if I had used traditional OCR from the start. The problem was that the hallucinations made sense and so were easier to overlook. So I gave up on using AI for that kind of job.

  2. I never even thought of that! I’m so used to scanning stuff and then working with it using whatever tools are available, that I did not think of omitting the first step. The Greek was very good indeed – only a ὃ with the wrong accentuation.

    I must try Gemini for translation.

    Good point about the OCR using AI! Thank you. I think AI is great so long as you can check every word. As soon as you start using it for long chunks unchecked, you’re at risk.

  3. Thank you for posting about your experience translating ancient texts using modern tools such as AI. Have you tried more specialized models that are open source such as qwen? We have seen incredible results at VerbaPatrum.com using large qwen models. Granted these were on very expensive hardware running for days, but given we are building a complete patristic concordance with 78 church fathers and translating into native languages for our English and Russian speaking orthodox users it required a lot of careful engineering. However, again, the results are impressive.
    Is this your first time engaging with AI to translate? If not, you’ve probably seen the speed at which open source models are improving. Given this, I am curious to follow your experiments with AI in this area to see whether your opinions change of its opportunities.

Leave a Reply

Your email address will not be published. Required fields are marked *