Back to Search Start Over

Version [1.0]- [SAMbA-RaPis music to scientists’ ears: Adding provenance support to spark-based scientific workflows]

Authors :
Guedes, Thaylon
Mattoso, Marta
Bedo, Marcos
de Oliveira, Daniel
Source :
SoftwareX; December 2024, Vol. 28 Issue: 1
Publication Year :
2024

Abstract

While researchers benefit from Apache Spark for executing scientific workflows at scale, they often lack provenance support due to the framework’s design limitations. This paper presents SAMbA-RaP, a provenance extension for Apache Spark. It focuses on: (i)Executing external, black-box applications with intensive I/O operations within the workflow while leveraging Spark’s in-memory data structures, (ii)Extracting domain-specific data from in-memory data structures and (iii)Implementing data versioning and capturing the provenance graph in a workflow execution. SAMbA-RaPalso provides real-time reports via a web interface, enabling scientists to explore dataflow transformations and content evolution as they run workflows.

Details

Language :
English
ISSN :
23527110
Volume :
28
Issue :
1
Database :
Supplemental Index
Journal :
SoftwareX
Publication Type :
Periodical
Accession number :
ejs67685623
Full Text :
https://doi.org/10.1016/j.softx.2024.101927