A Curated Corpus of Climate Finance Literature, 1990–2024
Multilingual Retrieval and Institutional Reports
DOI:
https://doi.org/10.52024/fn48qf66Keywords:
climate finance, bibliometric corpus , multilingual, sentence-transformer embeddings, scientometrics, history of economic thoughtAbstract
This data paper presents a curated, multilingual corpus of 33,344 works on climate finance published between 1990 and 2024. The dataset is assembled from 8 complementary sources that combine academic databases with a selected layer of institutional reports and key documents (1.4% of the corpus), adding institutional vocabulary and negotiation records absent from academic indexes. A multilingual retrieval strategy based on an eight-language keyword taxonomy is used to capture relevant works across linguistic contexts, while a reproducible pipeline integrates deduplication, metadata harmonisation, and quality filtering. English accounts for 93.8% of the refined works. The corpus includes a citation network derived from Crossref and OpenAlex, as well as embeddings pre-computed by a multilingual sentence-transformer.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Minh Ha-Duong

This work is licensed under a Creative Commons Attribution 4.0 International License.
Licensing information can be found here.
