Back to Search Start Over

CLARA-MeD corpus

Authors :
Ministerio de Ciencia e Innovación (España)
Agencia Estatal de Investigación (España)
Campillos-Llanos, Leonardo [0000-0003-3040-1756]
Valverde Mateos, Ana [0000-0003-1610-0770]
Capllonch Carrión, Adrián [0000-0001-9593-8621]
Campillos-Llanos, Leonardo [leonardo.campillos@csic.es]
Campillos-Llanos, Leonardo
Terroba Reinares, Ana Rosa
Zakhir Puig, Sofía
Valverde Mateos, Ana
Capllonch Carrión, Adrián
Ministerio de Ciencia e Innovación (España)
Agencia Estatal de Investigación (España)
Campillos-Llanos, Leonardo [0000-0003-3040-1756]
Valverde Mateos, Ana [0000-0003-1610-0770]
Capllonch Carrión, Adrián [0000-0001-9593-8621]
Campillos-Llanos, Leonardo [leonardo.campillos@csic.es]
Campillos-Llanos, Leonardo
Terroba Reinares, Ana Rosa
Zakhir Puig, Sofía
Valverde Mateos, Ana
Capllonch Carrión, Adrián
Publication Year :
2022

Abstract

A collection of 24 298 pairs of professional and simplified texts (>96 million tokens) for automatic medical text simplification in Spanish. A parallel corpus with a subset of 3800 sentence pairs of professional and laymen variants (149 862 tokens) is released as a benchmark for medical text simplification. This dataset was collected in the CLARA-MeD project, with the goal of simplifying medical texts in the Spanish language and reducing the language barrier to patient's informed decision making. In particular, the project aims at developing linguistic resources for automatic medical term simplification in Spanish; and conducting experiments in automatic text simplification.

Details

Database :
OAIster
Notes :
http://www.geonames.org/2510769/kingdom-of-spain.html, start=2022-01-04; end=2022-02-23, Spanish
Publication Type :
Electronic Resource
Accession number :
edsoai.on1333185859
Document Type :
Electronic Resource