Speculative Distributed CSV Data Parsing for Big Data Analytics

Authors :: Badrish Chandramouli
Chang Ge
Yinan Li
Donald Kossmann
Eric L. Eilebrecht
Source :: SIGMOD Conference
Publication Year :: 2019
Publisher :: ACM, 2019.
Abstract: There has been a recent flurry of interest in providing query capability on raw data in today's big data systems. These raw data must be parsed before processing or use in analytics. Thus, a fundamental challenge in distributed big data systems is that of efficient parallel parsing of raw data. The difficulties come from the inherent ambiguity while independently parsing chunks of raw data without knowing the context of these chunks. Specifically, it can be difficult to find the beginnings and ends of fields and records in these chunks of raw data. To parallelize parsing, this paper proposes a speculation-based approach for the CSV format, arguably the most commonly used raw data format. Due to the syntactic and statistical properties of the format, speculative parsing rarely fails and therefore parsing is efficiently parallelized in a distributed setting. Our speculative approach is also robust, meaning that it can reliably detect syntax errors in CSV data. We experimentally evaluate the speculative, distributed parsing approach in Apache Spark using more than 11,000 real-world datasets, and show that our parser produces significant performance benefits over existing methods.

Subjects :: Parsing
Information retrieval
business.industry
Computer science
media_common.quotation_subject
Big data
Context (language use)
02 engineering and technology
Ambiguity
computer.software_genre
Syntax
Analytics
020204 information systems
0202 electrical engineering, electronic engineering, information engineering
Syntax error
business
Raw data
computer
media_common

Database :: OpenAIRE
Journal :: Proceedings of the 2019 International Conference on Management of Data
Accession number :: edsair.doi...........a3e5dcd1bfb0f78db52d1ad4521a70d4
Full Text :: https://doi.org/10.1145/3299869.3319898