Back to Search
Start Over
Version [1.0]- [SAMbA-RaPis music to scientists’ ears: Adding provenance support to spark-based scientific workflows]
- Source :
- SoftwareX; December 2024, Vol. 28 Issue: 1
- Publication Year :
- 2024
-
Abstract
- While researchers benefit from Apache Spark for executing scientific workflows at scale, they often lack provenance support due to the framework’s design limitations. This paper presents SAMbA-RaP, a provenance extension for Apache Spark. It focuses on: (i)Executing external, black-box applications with intensive I/O operations within the workflow while leveraging Spark’s in-memory data structures, (ii)Extracting domain-specific data from in-memory data structures and (iii)Implementing data versioning and capturing the provenance graph in a workflow execution. SAMbA-RaPalso provides real-time reports via a web interface, enabling scientists to explore dataflow transformations and content evolution as they run workflows.
Details
- Language :
- English
- ISSN :
- 23527110
- Volume :
- 28
- Issue :
- 1
- Database :
- Supplemental Index
- Journal :
- SoftwareX
- Publication Type :
- Periodical
- Accession number :
- ejs67685623
- Full Text :
- https://doi.org/10.1016/j.softx.2024.101927