RESEARCH INITIATIVE · AUTh

The era
through its texts.

Greek texts from 1821–1922.
A common ground for language, history
and artificial intelligence.

An interdisciplinary initiative to build a corpus of Greek historical texts and to study the written heritage with modern computational methods.

Computational
Linguistics
+High-Performance
Computing
ONE DOCUMENT, MANY POSSIBILITIES Historical page from the book Ta psila vouna (The High Mountains), 1918 edition.
The page keeps its voice.1918
Z. Papantoniou · The High Mountains IEP / Hellenic Parliament Library
1821 1922

The Greek language, wherever it was written.Greece · Eastern Mediterranean · Europe · diaspora communities

01 / THE INITIATIVE

Access to a
world of written sources.

The written heritage of an era lies in large collections, local archives and personal documents. We want to create the connections that will make it more accessible.

Our aim is to build a documented corpus of Greek texts from the period 1821–1922 and to develop tools that make them easier to search, study and compare.

In the longer term, the corpus can become the basis for specialised AI language models, adapted to the language and the sources of the period. This prospect opens new ways of studying and accessing the material for the research community and the wider public, with clear reference to its provenance and terms of use.

Our work today

Seeking collections and partnerships, while focusing on training optical character recognition (OCR) models for historical Greek texts of the period 1821–1922.

02 / THE MATERIAL

Every type of text.
Every place where Greek was written.

We are looking for texts written or published in Greek between 1821 and 1922, regardless of where they are kept today and whether they have already been digitised.

01

From literature to the school book.

We are interested in the whole range of Greek writing: the major editions, but also the print that circulated in only a few copies.

  • Literature and translations
  • School and scientific books
  • Religious and political texts
  • Pamphlets, guides and almanacs

The words that were read, taught and travelled.

02

Everyday life, as it was recorded at the time.

The press preserves public debates, news and small stories. We are looking for Greek-language publications from Greece and from the communities abroad.

  • Daily and local newspapers
  • Literary and scientific periodicals
  • Diaspora press
  • Notices and public announcements

From the front page to the smallest notice.

03

History through the workings of institutions.

The archives of municipalities, communities, associations and public bodies offer a different view of the language and the social organisation of the period.

  • Minutes and administrative correspondence
  • Regulations and circulars
  • Community and municipal archives
  • Records of associations and societies

Local collections of great research value.

04

Personal voices deserve to be heard.

Letters, diaries and handwritten notes complement the printed output. Their automatic recognition is a research challenge in its own right.

  • Personal correspondence
  • Diaries and memoirs
  • Manuscripts of literary works
  • Notebooks and ledgers

Every hand, a different voice.

Small and lesser-known collections can fill the gaps left by large archives and broaden the picture of Greek writing. Digitised or not, every collection has its own path to collaboration.

03 / FROM PAGE TO TEXT

To be read.
To be searched. To be studied.

OCR turns the image of a page into editable text. For historical Greek, polytonic script and manuscripts, adaptation and quality control are part of the research.

01 / The digital page1918
Page 3 of the book Ta psila vouna (The High Mountains) by Zacharias Papantoniou, 1918 edition, with polytonic Greek text.
Z. Papantoniou, Ta psila vouna (The High Mountains), 1918, p. 3.
Digitisation: Institute of Educational Policy (IEP) · Original: Library of the Hellenic Parliament.Go to the source record
02 / The searchable text

3 matches · The search ignores accents and breathings.

2. Τὸ γράμμα τοῦ Ἀντρέα.

Καὶ τάχα δὲν μποροῦσαν νὰ εἶναι κι αὐτοὶ ἐκεῖ ψηλά;

Πολλὲς φορὲς ὁ δάσκαλος τοὺς εἶχε πεῖ στὸ μάθημα, πὼς τὰ παιδιὰ ποὺ εἶναι στὴν τελευταία τάξη τοῦ ἑλληνικοῦ, μποροῦν νὰ πᾶνε μόνα τους στὸ βουνό. Πὼς ἅμα ἔχουν θάρρος καὶ πειθαρχία, μποροῦν νὰ κατοικήσουν μόνοι τους ἐκεῖ ἕνα δυὸ μῆνες. Φτάνει νὰ ἔχουν τὴν ἄδεια τοῦ πατέρα τους, τὴν κατοικία καὶ τὴν τροφή.

«Πόσα πράματα, τοὺς εἶπε, θὰ μάθετε ὅταν πᾶτε τόσο ψηλά· οὔτε τὸ βιβλίο μπορεῖ νὰ σᾶς τὰ πῇ οὔτε ἐγώ». Κι ἔφυγε γιὰ τὴν πατρίδα του, γιὰ νὰ περάσῃ τὶς διακοπές. Ἦταν βέβαιος πὼς ἅμα θέλουν τὰ παιδιά, θὰ τὸ κατορθώσουν.

Τὰ παιδιὰ παρακάλεσαν τοὺς γονεῖς τους, νὰ τοὺς ἀφήσουν νὰ πᾶνε. Ἐκεῖνοι ἀντιστάθηκαν στὴν ἀρχή.

PRESENTATION EXAMPLE

An edited transcription of a passage, for trying out the search. It is not an output or an accuracy measurement of the OCR the team is developing.

04 / THE VALUE OF COLLABORATION

What we offer
our partner institutions.

Collaboration is designed so that research results strengthen the use of the collections, both by the institutions that preserve them and by the research community.

Searchable texts

Text and OCR files are delivered to partner institutions for the agreed parts of their collections, making search and reader services easier.

Knowledge and tools

Processing documentation, quality assessment and, where one is developed, an adapted OCR model remain available to the institutions for future use.

Visible provenance

Explicit reference to the library or archive, citation of the source and acknowledgement of the collaboration in the research results.

05 / FORMS OF COLLABORATION

Working with archives,
libraries and cultural
heritage institutions.

The framework of each collaboration is shaped together with the institutions, according to their needs, the characteristics of their collections and the agreed terms of access.

Making use of digitised collections.

Collaboration starts from a representative sample, so that the team and the institution can assess the needs and define a feasible first step.

  1. 01

    Selecting material

    The team and the institution identify Greek documents of the period and examine images, metadata and any existing OCR.

  2. 02

    Pilot processing

    The team processes the agreed sample and checks recognition quality on selected pages.

  3. 03

    Delivering results

    The institution receives searchable text and documentation in the agreed formats.

The scale of processing and the timetable are set after the pilot phase, once recognition quality is known.

Preparing material that is not yet digitised.

The absence of digitisation does not rule out collaboration. The initial exploration concerns the content of the collection and the institution's capabilities.

  1. 01

    Recording needs

    The team and the institution review document types, dates, volume of material and the priorities of the collection.

  2. 02

    Designing a sample

    Together we examine how a suitable digital sample can be produced, with respect for preservation needs.

  3. 03

    Planning the next phase

    Quality, available resources and technical requirements are assessed before a wider programme of work is defined.

Full digitisation of a collection requires separate planning and agreement on the available resources.

Access on agreed terms.

A collection can contribute to research without being made public. The framework of use is agreed with the institution before access is granted.

  1. 01

    Defining access

    The authorised members of the team, the processing environment and the retention period of the material are specified.

  2. 02

    Defining uses

    Work is limited to the agreed analysis and OCR adaptation. Any further use for developing other AI models requires separate, explicit permission.

  3. 03

    Returning results

    The agreed results are delivered to the institution. Public release or transfer to third parties takes place only under a separate agreement.

The terms for the original material, the texts produced and the adapted models are agreed separately.

Every collaboration begins with understanding the collection.
The first step is a conversation with the institution.

Contact us about collaborations

06 / THE TEAM

Two scientific fields.
One shared direction.

A collaboration between two Schools of the Aristotle University of Thessaloniki with complementary roles. Computational linguistics takes on the compilation, documentation and linguistic analysis of the corpus; high-performance computing, the methods and the infrastructure for training models at scale. The team is joined by students and graduates who work on authentic historical material and gain experience where philology meets computational research.

COMPUTATIONAL LINGUISTICS

Alexandros Tantos

Associate Professor of Text Linguistics
and Computational Linguistics

School of Philology · AUTh
Academic profile
HIGH-PERFORMANCE COMPUTING

Nikos Pitsianis

Associate Professor of High-Performance
Computing Systems

School of Electrical
and Computer Engineering · AUTh
Academic profile
MEMBERS OF THE TEAM

They annotate historical texts, which are used to train an OCR model for the Greek of the period 1821–1922.

  • Rania Touriki

    Graduate · School of Philology, AUTh Member since April 2026 Profile
  • Eleni Tsourilla

    Postgraduate student, Theoretical and Applied Linguistics · School of Philology, AUTh Member since April 2026
  • Konstantinos Sidiropoulos

    Undergraduate student · School of Philology, AUTh Member since April 2026
  • Georgios Chairopoulos

    Graduate · School of Philology, AUTh Member since April 2026
  • Aliki Gkounta

    Postgraduate student in Language Technology · NKUA
    Graduate · School of Philology, AUTh
    Member since April 2026

07 / FREQUENTLY ASKED QUESTIONS

The initiative
in practice.

The scope of the research, the use of the sources and the principles of collaboration.

The originals remain with the institutions that preserve them. Collaboration concerns agreed access to material and specific research tasks. Provenance, the identity of the collection and the terms of use are kept in the documentation.

The aim is for every document to carry a clear reference to the institution and to the details of the collection and, where available, a link to the original digital record. The contribution of partner institutions will be acknowledged in the related research presentations and publications.

Depending on the agreed pilot project: searchable texts, OCR files, text-to-page alignment, processing documentation and quality assessment. Where an adapted OCR model is developed, its return to the institution, with usage instructions, is agreed as well.

OCR is a first field of work. In the longer term, the corpus can become the basis for specialised AI language models, adapted to the language and the sources of the period. Every use rests on the terms of the specific collection. Access for OCR or analysis does not imply permission for further model development or public release.

The criterion is the Greek language and the period 1821–1922. Our interest covers every place where Greek was written or printed: from Constantinople and Alexandria to the communities of Europe, the Americas and Australia.

Yes. Small and lesser-known collections can offer text types and local voices that are missing from large repositories. Suitability for automatic recognition is examined on a sample; we do not assume that every manuscript or worn page can already be recognised with high accuracy.

CONTACT

A collection, an archive, a question?
Write to us.

The team welcomes contact from researchers, academics and institutions who wish to learn about the initiative, exchange ideas or explore possibilities for collaboration.

Grafes 1821–1922 research team