A Curated Corpus of Climate Finance Literature, 1990–2024

Multilingual Retrieval and Institutional Reports

Authors

DOI:

https://doi.org/10.52024/fn48qf66

Keywords:

climate finance, bibliometric corpus , multilingual, sentence-transformer embeddings, scientometrics, history of economic thought

Abstract

This data paper presents a curated, multilingual corpus of 33,344 works on climate finance published between 1990 and 2024. The dataset is assembled from 8 complementary sources that combine academic databases with a selected layer of institutional reports and key documents (1.4% of the corpus), adding institutional vocabulary and negotiation records absent from academic indexes. A multilingual retrieval strategy based on an eight-language keyword taxonomy is used to capture relevant works across linguistic contexts, while a reproducible pipeline integrates deduplication, metadata harmonisation, and quality filtering. English accounts for 93.8% of the refined works. The corpus includes a citation network derived from Crossref and OpenAlex, as well as embeddings pre-computed by a multilingual sentence-transformer.

Downloads

Published

2026-09-26

Issue

Section

Data Papers

How to Cite

A Curated Corpus of Climate Finance Literature, 1990–2024: Multilingual Retrieval and Institutional Reports. (2026). Research Data Journal for the Humanities and Social Sciences, 10. https://doi.org/10.52024/fn48qf66