Back to Search Start Over

De-identification of Free Text Data containing Personal Health Information: A Scoping Review of Reviews

Authors :
Bekelu Negash
Alan Katz
Christine J. Neilson
Moniruzzaman Moni
Marcello Nesca
Alexander Singer
Jennifer E. Enns
Source :
International Journal of Population Data Science, Vol 8, Iss 1 (2023)
Publication Year :
2023
Publisher :
Swansea University, 2023.

Abstract

Introduction Using data in research often requires that the data first be de-identified, particularly in the case of health data, which often include Personal Identifiable Information (PII) and/or Personal Health Identifying Information (PHII). There are established procedures for de-identifying structured data, but de-identifying clinical notes, electronic health records, and other records that include free text data is more complex. Several different ways to achieve this are documented in the literature. This scoping review identifies categories of de-identification methods that can be used for free text data. Methods We adopted an established scoping review methodology to examine review articles published up to May 9, 2022, in Ovid MEDLINE; Ovid Embase; Scopus; the ACM Digital Library; IEEE Explore; and Compendex. Our research question was: What methods are used to de-identify free text data? Two independent reviewers conducted title and abstract screening and full-text article screening using the online review management tool Covidence. Results The initial literature search retrieved 3,312 articles, most of which focused primarily on structured data. Eighteen publications describing methods of de-identification of free text data met the inclusion criteria for our review. The majority of the included articles focused on removing categories of personal health information identified by the Health Insurance Portability and Accountability Act (HIPAA). The de-identification methods they described combined rule-based methods or machine learning with other strategies such as deep learning. Conclusion Our review identifies and categorises de-identification methods for free text data as rule-based methods, machine learning, deep learning and a combination of these and other approaches. Most of the articles we found in our search refer to de-identification methods that target some or all categories of PHII. Our review also highlights how de-identification systems for free text data have evolved over time and points to hybrid approaches as the most promising approach for the future.

Details

Language :
English
ISSN :
23994908
Volume :
8
Issue :
1
Database :
Directory of Open Access Journals
Journal :
International Journal of Population Data Science
Publication Type :
Academic Journal
Accession number :
edsdoj.0359c1ddd3d942d999628b386989fd0f
Document Type :
article
Full Text :
https://doi.org/10.23889/ijpds.v8i1.2153