Back to Search Start Over

Wikipedia citations: A comprehensive data set of citations with identifiers extracted from English Wikipedia

Authors :
Harshdeep Singh
Robert West
Giovanni Colavizza
Source :
Quantitative Science Studies, Vol 2, Iss 1, Pp 1-19 (2021)
Publication Year :
2021
Publisher :
The MIT Press, 2021.

Abstract

AbstractWikipedia’s content is based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close this gap, we release Wikipedia Citations, a comprehensive data set of citations extracted from Wikipedia. We extracted29.3 million citations from 6.1 million English Wikipedia articles as of May 2020, and classified as being books, journal articles, or Web content. We were thus able to extract 4.0 million citations to scholarly publications with known identifiers—including DOI, PMC, PMID, and ISBN—and further equip an extra 261 thousand citations with DOIs from Crossref. As a result, we find that 6.7% of Wikipedia articles cite at least one journal article with an associated DOI, and that Wikipedia cites just 2% of all articles with a DOI currently indexed in the Web of Science. We release our code to allow the community to extend upon our work and update the data set in the future.

Subjects

Subjects :
Science (General)
Q1-390

Details

Language :
English
ISSN :
26413337
Volume :
2
Issue :
1
Database :
Directory of Open Access Journals
Journal :
Quantitative Science Studies
Publication Type :
Academic Journal
Accession number :
edsdoj.4230c7d80f54d11a8b402c06c6696b3
Document Type :
article
Full Text :
https://doi.org/10.1162/qss_a_00105