Searchable texts
Text and OCR files are delivered to partner institutions for the agreed parts of their collections, making search and reader services easier.
RESEARCH INITIATIVE · AUTh
Greek texts from 1821–1922.
A common ground for language, history
and artificial intelligence.
An interdisciplinary initiative to build a corpus of Greek historical texts and to study the written heritage with modern computational methods.
The Greek language, wherever it was written.Greece · Eastern Mediterranean · Europe · diaspora communities
01 / THE INITIATIVE
The written heritage of an era lies in large collections, local archives and personal documents. We want to create the connections that will make it more accessible.
Our aim is to build a documented corpus of Greek texts from the period 1821–1922 and to develop tools that make them easier to search, study and compare.
In the longer term, the corpus can become the basis for specialised AI language models, adapted to the language and the sources of the period. This prospect opens new ways of studying and accessing the material for the research community and the wider public, with clear reference to its provenance and terms of use.
Seeking collections and partnerships, while focusing on training optical character recognition (OCR) models for historical Greek texts of the period 1821–1922.
02 / THE MATERIAL
We are looking for texts written or published in Greek between 1821 and 1922, regardless of where they are kept today and whether they have already been digitised.
We are interested in the whole range of Greek writing: the major editions, but also the print that circulated in only a few copies.
The words that were read, taught and travelled.
The press preserves public debates, news and small stories. We are looking for Greek-language publications from Greece and from the communities abroad.
From the front page to the smallest notice.
The archives of municipalities, communities, associations and public bodies offer a different view of the language and the social organisation of the period.
Local collections of great research value.
Letters, diaries and handwritten notes complement the printed output. Their automatic recognition is a research challenge in its own right.
Every hand, a different voice.
Small and lesser-known collections can fill the gaps left by large archives and broaden the picture of Greek writing. Digitised or not, every collection has its own path to collaboration.
03 / FROM PAGE TO TEXT
OCR turns the image of a page into editable text. For historical Greek, polytonic script and manuscripts, adaptation and quality control are part of the research.
3 matches · The search ignores accents and breathings.
Καὶ τάχα δὲν μποροῦσαν νὰ εἶναι κι αὐτοὶ ἐκεῖ ψηλά;
Πολλὲς φορὲς ὁ δάσκαλος τοὺς εἶχε πεῖ στὸ μάθημα, πὼς τὰ παιδιὰ ποὺ εἶναι στὴν τελευταία τάξη τοῦ ἑλληνικοῦ, μποροῦν νὰ πᾶνε μόνα τους στὸ βουνό. Πὼς ἅμα ἔχουν θάρρος καὶ πειθαρχία, μποροῦν νὰ κατοικήσουν μόνοι τους ἐκεῖ ἕνα δυὸ μῆνες. Φτάνει νὰ ἔχουν τὴν ἄδεια τοῦ πατέρα τους, τὴν κατοικία καὶ τὴν τροφή.
«Πόσα πράματα, τοὺς εἶπε, θὰ μάθετε ὅταν πᾶτε τόσο ψηλά· οὔτε τὸ βιβλίο μπορεῖ νὰ σᾶς τὰ πῇ οὔτε ἐγώ». Κι ἔφυγε γιὰ τὴν πατρίδα του, γιὰ νὰ περάσῃ τὶς διακοπές. Ἦταν βέβαιος πὼς ἅμα θέλουν τὰ παιδιά, θὰ τὸ κατορθώσουν.
Τὰ παιδιὰ παρακάλεσαν τοὺς γονεῖς τους, νὰ τοὺς ἀφήσουν νὰ πᾶνε. Ἐκεῖνοι ἀντιστάθηκαν στὴν ἀρχή.
An edited transcription of a passage, for trying out the search. It is not an output or an accuracy measurement of the OCR the team is developing.
04 / THE VALUE OF COLLABORATION
Collaboration is designed so that research results strengthen the use of the collections, both by the institutions that preserve them and by the research community.
Text and OCR files are delivered to partner institutions for the agreed parts of their collections, making search and reader services easier.
Processing documentation, quality assessment and, where one is developed, an adapted OCR model remain available to the institutions for future use.
Explicit reference to the library or archive, citation of the source and acknowledgement of the collaboration in the research results.
05 / FORMS OF COLLABORATION
The framework of each collaboration is shaped together with the institutions, according to their needs, the characteristics of their collections and the agreed terms of access.
Collaboration starts from a representative sample, so that the team and the institution can assess the needs and define a feasible first step.
The team and the institution identify Greek documents of the period and examine images, metadata and any existing OCR.
The team processes the agreed sample and checks recognition quality on selected pages.
The institution receives searchable text and documentation in the agreed formats.
The scale of processing and the timetable are set after the pilot phase, once recognition quality is known.
The absence of digitisation does not rule out collaboration. The initial exploration concerns the content of the collection and the institution's capabilities.
The team and the institution review document types, dates, volume of material and the priorities of the collection.
Together we examine how a suitable digital sample can be produced, with respect for preservation needs.
Quality, available resources and technical requirements are assessed before a wider programme of work is defined.
Full digitisation of a collection requires separate planning and agreement on the available resources.
A collection can contribute to research without being made public. The framework of use is agreed with the institution before access is granted.
The authorised members of the team, the processing environment and the retention period of the material are specified.
Work is limited to the agreed analysis and OCR adaptation. Any further use for developing other AI models requires separate, explicit permission.
The agreed results are delivered to the institution. Public release or transfer to third parties takes place only under a separate agreement.
The terms for the original material, the texts produced and the adapted models are agreed separately.
Every collaboration begins with understanding the collection.
The first step is a conversation with the institution.
06 / THE TEAM
A collaboration between two Schools of the Aristotle University of Thessaloniki with complementary roles. Computational linguistics takes on the compilation, documentation and linguistic analysis of the corpus; high-performance computing, the methods and the infrastructure for training models at scale. The team is joined by students and graduates who work on authentic historical material and gain experience where philology meets computational research.

Associate Professor of Text Linguistics
and Computational Linguistics

Associate Professor of High-Performance
Computing Systems
They annotate historical texts, which are used to train an OCR model for the Greek of the period 1821–1922.



07 / FREQUENTLY ASKED QUESTIONS
The scope of the research, the use of the sources and the principles of collaboration.
CONTACT
The team welcomes contact from researchers, academics and institutions who wish to learn about the initiative, exchange ideas or explore possibilities for collaboration.