Back to Search
Start Over
CLARA-MeD corpus
- Publication Year :
- 2022
-
Abstract
- A collection of 24 298 pairs of professional and simplified texts (>96 million tokens) for automatic medical text simplification in Spanish. A parallel corpus with a subset of 3800 sentence pairs of professional and laymen variants (149 862 tokens) is released as a benchmark for medical text simplification. This dataset was collected in the CLARA-MeD project, with the goal of simplifying medical texts in the Spanish language and reducing the language barrier to patient's informed decision making. In particular, the project aims at developing linguistic resources for automatic medical term simplification in Spanish; and conducting experiments in automatic text simplification.
Details
- Database :
- OAIster
- Notes :
- http://www.geonames.org/2510769/kingdom-of-spain.html, start=2022-01-04; end=2022-02-23, Spanish
- Publication Type :
- Electronic Resource
- Accession number :
- edsoai.on1333185859
- Document Type :
- Electronic Resource