Skip to content

ELC-DHLR

Czech-Swiss Cooperation
of CLARIN and DARIAH Infrastructures
Enhancing Library Collections for Digital Humanities and Language Research

About

The project aims to make digitised library collections truly research-ready for digital humanities and language sciences.

Making Library Collections Research-Ready

Project Abstract

The project aims to make digitised library collections truly research-ready for digital humanities and language sciences. Building on Czech and Swiss infrastructure consortia for CLARIN and DARIAH, it will enhance the collections and create a sustainable and user-friendly ecosystem aligned with FAIR principles. ELC-DHLR will unlock the potential of centuries of cultural heritage, ensuring that digitised collections can be effectively used for advanced cross-disciplinary research, and at the same time, it will contribute to strengthening the cooperation between leading research organizations in CZ and CH.

From Digitisation to Research Data

Technical Context

The ELC-DHLR project is an ambitious collaborative initiative between Czech and Swiss infrastructures CLARIN and DARIAH that addresses the need to make the vast resources of digitised library collections truly usable for advanced scholarly research. While major digitisation programmes over the past decades have created millions of digital objects across Europe, researchers in the Digital Humanities and language sciences still encounter fundamental barriers: inconsistent OCR quality, minimal metadata, interfaces aimed primarily at casual readers, and a lack of integration with modern analytical environments.

At the same time, technological advances are opening unprecedented opportunities. Artificial Intelligence, particularly in the form of Large Language Models, has matured to the point where automated summarisation, semantic search, and intelligent exploration of collections are now achievable. Linguistic research and corpus analysis infrastructures such as the Czech National Corpus and the LiRI Corpus Platform provide powerful linguistic tools that can be connected with digitised library holdings.

Czech-Swiss Research Infrastructure Partnership

Bilateral Cooperation

The project directly addresses the objectives of the call by creating a long-term strategic partnership between the Czech and Swiss CLARIN and DARIAH infrastructures to turn digitized library collections in both countries into openly accessible and analysable data for digital humanities, language research and education.

    Bilateral collaboration across the full workflow from digitization and OCR to AI enrichment, data access, and corpus analysis. International integration of Czech and Swiss infrastructures within European and global research networks. Capacity building through know-how transfer in data enrichment, corpus methods, and user interface design. Innovation and knowledge transfer using AI-based enrichment, summarization, and analysis tools.

Main Objective

To substantially increase the usability of digitised library collections for research and unlock their potential for advanced interdisciplinary work. The project aims to enhance data both technically and conceptually, provide user-friendly access, and strengthen cooperation between research organisations in Czechia and Switzerland.

Challenges Addressed

  • Inconsistent OCR quality
  • Insufficient or minimal metadata
  • Digital library interfaces primarily designed for general readers
  • Weak integration of collections with analytical tools
  • Difficult data extraction for research purposes

Proposed Solution

  • Transformation and enrichment of digital library content
  • Integration of external services for OCR, entity recognition, and further enrichment
  • Use of local LLMs for summarisation and semantic search
  • Development of the DL-enricher platform
  • Development of Kramerius Research Workspace
  • Transfer of data to the CNC and LCP corpus infrastructures
  • AI assistant for translating natural-language queries into corpus query language

For Researchers

Easier creation of custom datasets, exports in research-ready formats, and access to enriched digital library content.

For Libraries

Validation of tools for data transformation and enrichment, integration with the Kramerius system, and transferable workflows for other digital libraries.

For Infrastructures

Connection of library collections with CLARIN and DARIAH communities, corpus tools, and analytical services for language research.

Project Framework

The five-layer concept provides a transparent, modular, and FAIR-compatible workflow from data to analysis.

DL Researcher
Workspace
WP 3
DL Researcher...
Digital Library UI
WP 3
Digital Library UI...
Text Corpora UI
WP 3
Text Corpora UI...
APIs
WP 3
APIs...
Library Digital Collections
Library Digital Collections
External API users
External API users
CLARIN Text Corpora
CLARIN Text Corpora
CLARIN Analytics tools
WP 4
CLARIN Analytics tools...
Local LLM services for
transformation and enrichment
WP 2
Local LLM services for...
Enrichment services (entity
recognition,...)
WP 2
Enrichment services (entity...
Transformation services
(OCR,...)
WP 2
Transformation services...
DL Enrichment platform
WP 2
DL Enrichment platform...

Collections and data

Digitised library collections as the primary data source.

Transformation and Enrichment

Improvement of OCR, metadata, segmentation, and content enrichment.

Access and API

Interoperable access to data through APIs and related services.

Retrieval and GUI

User interfaces for researchers and the Research Workspace.

Analysis

Analytical tools, corpus services, and AI assistants.

Work Packages

Five interlinked work packages form the complete project ecosystem.

WP1

Project Management

Coordination: CUNI, UZH
Participation: all partners

Project management, transparent communication, financial administration, deliverable monitoring, risk management, and a FAIR-compliant Data Management Plan.

The success of a complex international project depends on effective management, transparent communication, and strong governance. WP1 addresses these needs by creating a robust framework for coordination and compliance across all partners. Its objectives are to guarantee smooth collaboration, adherence to contractual and financial requirements, and proactive monitoring of progress and quality. A Steering Committee will provide oversight, including the national coordinators of the Czech and Swiss CLARIN and DARIAH nodes together with representatives of each partner institution. This body ensures that decision-making is inclusive, strategic, and aligned with broader research infrastructure goals.

Project communication is central to WP1. Regular consortium meetings, both in-person and virtual, will keep partners aligned, while online workspaces and shared repositories will provide a central hub for resources. Systematic documentation of minutes, reports, and action points ensures transparency. Administrative and financial management processes will monitor expenditure, handle legal obligations, and support partners in fulfilling reporting requirements. Progress will be tracked against milestones and deliverables, with internal reviews providing early warnings of delays or underperformance. Risk management will be implemented through a regularly updated risk register and mitigation plans, while quality assurance will rely on defined procedures and deliverable submission checklists. Finally, a Data Management Plan (DMP) will be prepared and maintained to ensure FAIR-compliant handling of project data throughout its lifecycle.

Tasks

  1. T1.1 Organisation of meetings, establishment of online collaborative platforms, systematic documentation, and coordinated communication with stakeholders and the public.
  2. T1.2 Establishment of a Steering Committee.
  3. T1.3 Monitoring of budget expenditures, handling of legal obligations, liaison with funding agencies, and support for partners in reporting.
  4. T1.4 Tracking of milestones and deliverables, preparation of progress summaries, and cross-WP timeline alignment.
  5. T1.5 Maintenance of a risk register, implementation of mitigation strategies, and enforcement of quality control procedures.
  6. T1.6 Preparation and continuous updating of a FAIR-compliant Data Management Plan.

Deliverables

  • D1.2 Establishment of the Steering Committee (internal) (M6, CUNI)
  • D1.3 Final report (for funding agency) (M36, CUNI)
  • D1.4.1 Interim report (for funding agency) (M10, CUNI)
  • D1.4.2 Interim report (for funding agency) (M22, CUNI)
  • D1.6 Project DMP (internal) (M6, CUNI)
WP2

Transformation and Enrichment

Coordination: KNAV, MZK
Participation: NKP, ETH, UZH

Development of the DL-enricher platform, integration of external services, use of LLMs for summarisation and semantic search, and entity recognition.

WP2 provides the technological backbone of the project by developing the DL-enricher platform which will enable digital library administrators to transform and enrich DL content with selected external tools and services. This open and extensible environment will allow digital libraries to integrate state-of-the-art services for OCR transformation and content enrichment directly into their workflows. By using ALTO XML outputs, the platform will take advantage of detailed word-level positioning and document structures, enabling more accurate text correction and enrichment. The platform will also support advanced enrichment tools such as named entity recognition, linguistic annotation, and detection of non-textual elements like images, tables, and graphs. In this way, WP2 significantly increases the discoverability, usability, and long-term value of digital collections. The development of the platform, including the integration of Large Language Models, will be continuously improved by the identification and rigorous testing of candidate services and tools, which will be consequently integrated to the platform. This process will utilize diverse datasets from Czech and Swiss libraries to ensure the creation of robust transformed and enriched data examples.

A major innovation of WP2 is the integration of Large Language Models. Locally hosted LLMs will be used to generate contextually relevant summaries of documents, tailored dynamically to their type and thematic focus. These summaries will be embedded into vector representations and indexed for semantic retrieval. Researchers will thus be able to interact with digital libraries using natural language queries, greatly simplifying and enhancing discovery. To better serve both human readers and advanced search algorithms, a two-tiered summarization model will be implemented. The first tier will produce a concise, readable narrative summary for general users. The second tier will use an agent-driven, sequential prompting workflow to extract and structure key entities (such as people, places, and events) into fine-grained, metadata-rich sections. This dual output delivers the best of both worlds – readable summaries for users and detailed data tables for downstream search, analytics, or integration tasks.

Tasks

  1. T2.1 Development of DL-enricher platform for calling and managing transformation and enrichment services.
  2. T2.2 Identification and testing of candidate services using datasets from Czech and Swiss libraries, creation of transformed and enriched dataset examples.
  3. T2.4 Local LLM implementation for summarization and semantic search for monographs in digital libraries, including backend, storage, APIs, and query vectorisation.

Deliverables

  • D2.1 DL-enricher - software released as open source - Result Code - R (M30, KNAV)
  • D2.2 Internal report on enrichment services and ground truth datasets (M33, KNAV)
  • D2.4 Segmented summaries for monographs, functional backend, API, and semantic search prototype with initial summaries at the Moravian Library - Result Code - G (M33, MZK)
WP3

Access and Retrieval

Coordination: KNAV, ETH
Participation: NKP, MZK, CUNI, UZH

API best practices, proxy integration, Kramerius Research Workspace, exports in TXT, CSV, JSON, TEI, and ALTO formats, and data transfer to CNC and LCP.

WP3 ensures that enriched and transformed collections become accessible to both developers and researchers. It fosters a collaborative environment for the exchange of best practices and know-how regarding API design for accessing digital collections, builds a Research Workspace with rich UI possibilities within the Kramerius digital library system, and establishes data transfer pipelines to corpus infrastructures. These components guarantee that collections are not only enriched but also discoverable, reusable, and connected to analytical environments.

The Research Workspace will be embedded in Kramerius, enabling scholars to work with digitised library collections as data. Key features will include advanced search and filtering based on named entity recognition (NER), the selection and bulk export of multiple publications in various formats (TXT, CSV, JSON, TEI, ALTO) with customizable parameters, and the option to request access to copyright-restricted content under the TDM exception for scientific research. Exports will be FAIR-aligned and will include clear licensing. Finally, WP3 establishes pipelines for bringing Czech library documents into the CNC and ETH library documents into LCP, with automation ensuring regular updates, as a bridge towards the WP4 Data analysis.

Tasks

  1. T3.1 Knowledge exchange and best practice sharing for API access to digital collections.
  2. T3.4 Development of Research Workspace UI within Kramerius, including extended search, filtering, and export options.
  3. T3.8 Import of Czech library data into CNC with post-processing, metadata adjustments and creation of workflows for data cleaning and conversion.
  4. T3.9 Import of ETH library data into LCP with post-processing, metadata adjustments and creation of workflows for data cleaning and conversion.

Deliverables

  • D3.1 Report on best practices for API access (M33, ETH)
  • D3.4.1 Kramerius Research Workspace - software released as open source - Result Code - R (M21, KNAV)
  • D3.4.2 API documentation of Research Workspace (M21, KNAV)
  • D3.8 New corpora publicly available in CNC (M33, CUNI)
  • D3.9 New corpora publicly available at LCP (M33, UZH)
WP4

Data Analysis

Coordination: CUNI
Participation: UZH

AI assistant for translating natural-language queries into formal corpus queries, integrated into KonText and LCP.

WP4 builds directly upon the improved data of WP2 and access mechanisms of WP3 by advancing analytical tools for corpus research. It addresses both accessibility for non-technical researchers and advanced linguistic functionality. The work package’s main innovation is an AI-powered query assistant.

The AI-powered query assistant will help lower the barrier for many SSH researchers that still exists for advanced utilisation of the available query engines. Using the LLMs, it will transform user queries from natural language to formal corpus query language, with the possibility to edit the AI-generated query if needed, and run the query on a selected corpus. From a technical perspective, the plan is to utilize APIs to LLMs running at large academic non-commercial centres such as e-INFRA CZ.

On the Czech side, the assistant will be integrated into the KonText interface as another type of query option. As the CNC’s KonText query engine is deployed also by other CLARIN centres in Europe, this new functionality is likely to spread also beyond the Swiss-Czech cooperation. In LCP, the assistant will be integrated as an option in the CQL query mode, after bringing the current CQL implementation closer to the standard CQL specification.

Tasks

  1. T4.1 Development of the AI-powered query assistant for KonText and LCP with integration of LLM APIs.

Deliverables

  • D4.1.1 Scientific Article - Result Code - J (M33, CUNI)
  • D4.1.2 AI-powered query assistant deployed at CNC/KonText (M33, CUNI)
  • D4.1.3 AI-powered query assistant deployed at LCP (M33, UZH)
WP5

Knowledge Exchange, Training and Dissemination

Coordination: UZH, NKP
Participation: all partners

Project visibility, workshops, project website, public reports, presentation of results, and active engagement of the research community.

WP5 ensures that the project achieves visibility, uptake by target pan-european scientific communities, long-term sustainability and strategic cooperation. It focuses on knowledge transfer, dissemination and community building. Internal and public workshops, conference presentations, and shared open and FAIR-compliant resources all contribute to creating a strong network of users and contributors both within and outside the project framework.

The WP begins with an internal kick-off workshop in Switzerland, followed by two public workshops (one in Switzerland and one in the Czech Republic), as well as an international workshop for the cross-ERIC DARIAH/CLARIN Library Working Group. These activities are essential for successfully exchanging expertise, best practices and building a long-term strategic partnership between leading institutions in the areas of digital humanities and language research and in Czechia and Switzerland. Dissemination activities will include the project website, public reports and presentations at international conferences such as the CLARIN Annual Conference and the DARIAH Annual Event. By sharing knowledge and open resources with the scientific and infrastructure communities dealing with arts, humanities and language FAIR-compliant digital research objects, the project achieves high impact for its duration and will continue to support innovation after its completion.

Tasks

  1. T5.1 Knowledge exchange workshops and outreach events.
  2. T5.2 Dissemination of activities, outcomes, and publications.

Deliverables

  • D5.1.1 Workshop Series Plan (M2, UZH)
  • D5.1.2 First internal workshop in CH (M4, ETHZ)
  • D5.1.3 First public workshop in CH - Result Code - W (M21, UZH)
  • D5.1.4 Second public workshop in CZ - Result Code - W (M33, NKP)
  • D5.1.5 Final Knowledge Exchange Report and Recommendations (M33, UZH)
  • D5.2.1 Project website (M4, KNAV)
  • D5.2.2 DARIAH/CLARIN Library working group workshop - W (M12, KNAV)
  • D5.2.3 Presentation of project results - J (M33, UZH)

Outputs

Concrete outputs and expected results that the project will create for the research community.

Software

Tools and Applications

  • • DL-enricher released as open-source software (D2.1, M30)
  • • Kramerius Research Workspace released as open-source software (D3.4.1, M21)
  • • AI-powered query assistant deployed at CNC/KonText (D4.1.2, M33)
  • • AI-powered query assistant deployed at LCP (D4.1.3, M33)
Prototypes and Services

Search and Enrichment

  • • Segmented summaries for monographs (D2.4, M33)
  • • Functional backend, API, and semantic search prototype at the Moravian Library (D2.4, M33)
  • • Internal report on enrichment services and ground truth datasets (D2.2, M33)
  • • Prototype deployment at the Library of the Czech Academy of Sciences (D2.3, M36)
Data and Corpora

Research-Ready Data

  • • New corpora publicly available in CNC (D3.8, M33)
  • • New corpora publicly available at LCP (D3.9, M33)
  • • Data cleaning and conversion workflows for Czech library data (T3.8)
  • • Data cleaning and conversion workflows for ETH library data (T3.9)
Documents and Publications

Reports and Outputs

  • • Article on the AI-powered query assistant (D4.1.1, M33)
  • • Presentation of project results (D5.2.3, M33)
  • • DARIAH/CLARIN Library Working Group workshop (D5.2.2, M12)
  • • Public workshops in Switzerland and Czechia (D5.1.3, D5.1.4)

The overview follows the project revision and expected results table. Detailed links to software, documentation, datasets, publications, and workshop materials will be added progressively during the project.

Timeline

The project timeline is divided into three phases which correspond to the individual calendar years 2026-2028.

2026: Analysis and Preparation

  • Set-up of project management and communication
  • Kick-off meeting in Zurich
  • Establishment of the Steering Committee
  • DARIAH/CLARIN Library Groups in Prague
  • Preparation of the Data Management Plan
  • Analysis of WP2, WP3, and WP4 components

During the first phase, project management structures and communication channels will be established. An analytical phase of the project will be carried out, during which the aspects of individual components in WP2, WP3, and WP4 will be examined and discussed in greater detail. The analytical outputs will serve as a basis for the implementation of development work, which will commence in the second half of 2026.

As a part of the first phase, a kick-off meeting will be held in Zurich, the Project Steering Committee will be established, and a joint meeting of the DARIAH and CLARIN Library Groups will be organized in Prague, where the project objectives and outputs will be discussed. A project Data Management Plan will be prepared. A report on the progress and outcomes of the second phase will be prepared at the end of 2026.

2027: Development of Key Components

  • Development of the DL-enricher platform
  • Development of Kramerius Research Workspace
  • Implementation of LLM services
  • Development of the AI query assistant
  • Data imports into CNC and LCP
  • Public workshop in Zurich

The second phase will take place in 2027, and its main purpose will be the development of individual components of the technical solution, which together will form a complete ecosystem in the final phase of the project. Programming work will be carried out on the DL-enricher platform, the Research Workspace environment will be developed, LLMs will be implemented, AI assistants translating natural-language queries into formal query languages will be created, and data will be imported into language corpora. In parallel, external tools and services for transformation and enrichment will be explored and tested.

In the second half of 2027, a public workshop will be organized in Zurich, where current outputs and further planned objectives of the project will be presented and discussed. Towards the end of the second phase, the first version of a new user interface for data extraction from digital libraries – the Kramerius Research Workspace – will be released. A report on the progress and outcomes of the second phase will be prepared at the end of 2027.

2028: Completion and Dissemination

  • Completion of DL-enricher and publication of the source code
  • Test deployment in the Digital Library of the Czech Academy of Sciences
  • Summarisation and semantic search at the Moravian Library
  • New corpora in CNC and LCP
  • Completion of the AI assistant
  • Public workshop in Prague and final report

The third phase will take place in 2028, and its main purpose will be the completion of development work as well as increased dissemination and communication of the project outputs. Development of the DL-enricher platform will be finalized, including the release of the source code with documentation. In a test operation, DL-enricher will be implemented at the Library of the Czech Academy of Sciences together with selected transformation and enrichment tools recommended in a dedicated report prepared within the project. In a pilot operation, semantic search and summarization will be deployed on monographs in the digital library of the Moravian Library, including the developed backend, storage, and documented API. This solution will be transferable to other digital libraries as well. New CNC and LCP corpora will be made available based on the data obtained from the digital libraries, and the implementation of the AI assistants translating natural-language queries into formal query languages will be completed.

At the end of the third phase, the two research articles will be published and a public workshop will be organized in Prague, at which the project outputs will be presented to researchers, students, and administrators of research infrastructures. At the end of 2028, a final report on the progress and outcomes of the project for the entire 2026-2028 period will be prepared.

Partners

Project partners from Czech and Swiss CLARIN and DARIAH infrastructures.

Project Team

Overview of key partner representatives and project contacts.

Charles University

prof. RNDr. Jan Hajič, Dr.

LINDAT/CLARIAH-CZ

hajic@ufal.mff.cuni.cz

Charles University

Mgr. Michal Křen, Ph.D.

Czech National Corpus

michal.kren@ff.cuni.cz

Library of the Czech Academy of Sciences

Ing. Martin Lhoták

LINDAT/CLARIAH-CZ

lhotak@knav.cz

Moravian Library

Ing. Petr Žabička

LINDAT/CLARIAH-CZ

Petr.Zabicka@mzk.cz

National Library of the Czech Republic

Mgr. Bc. Michaela Bežová

LINDAT/CLARIAH-CZ

michaela.bezova@nkp.cz

University of Zurich

Dr. Cristina Grisot

CLARIN-CH

cristina.grisot@uzh.ch

University of Basel

Prof. Dr. Rita Gautschy

DARIAH-CH

rita.gautschy@unibas.ch

ETH Zurich

Michael Gasser

DARIAH-CH

michael.gasser@library.ethz.ch

News

Latest information about the project, workshops, and published outputs.

Development

Start of Development Work

Lorem ipsum dolor sit amet, consectetur adipisicing elit. Cupiditate ducimus facilis inventore iste laudantium quos similique voluptas! At consequatur consequuntur debitis eaque eos laborum…

Workshop

DARIAH/CLARIN Library Working Group

Lorem ipsum dolor sit amet, consectetur adipisicing elit. Cupiditate ducimus facilis inventore iste laudantium quos similique voluptas! At consequatur consequuntur debitis eaque eos laborum…

Development

Launch of the ELC-DHLR Project

The project will officially start on 1 April 2026. The first kick-off meeting will take place in Zurich. Lorem ipsum dolor sit amet, consectetur…

Funding

Call 8K2502 for proposals for bilateral cooperation between Czech and Swiss research infrastructures under the Czech-Swiss cooperation programme in research infrastructures.

Provider: Ministry of Education, Youth and Sports.