Start Over

ANALYSIS OF METHODS AND ALGORITHMS FOR PROCESSING UNSTRUCTURED TEXT DATA BASED ON JSON TECHNOLOGY.

Authors :: Kucherenko, Yehor
Kulakovska, Inessa
Source :: Technology Audit & Production Reserves. 2024, Vol. 3 Issue 2(77), p10-18. 9p.
Publication Year :: 2024
Abstract: The object of research is the process of automating systems for structuring data from several sources. The subject of the research is methods and algorithms for implementing a complete system for automated and parallel processing, validation and structuring of data. One of the most problematic areas is the merging of databases with different structures and several common fields into a generalized structure. The research was aimed at developing a system to increase the efficiency of automation of big data processing. As a result of the work, optimization methods were studied, the influence of their internal parameters on the operation of algorithms was analyzed, their main advantages and disadvantages were determined, and software was developed in which the corresponding methods were implemented. An algorithm for structuring data before processing has been obtained. Data structuring is achieved by performing the «mapping» operation. Mapping can take place by indexes of already cleaned data or using a defined dictionary with a given set of keys, which allows not to care about the sequence of storing values and their possible shift. The practical significance of the developed system lies in the improvement of methods of collecting and processing information for the purpose of its further validation, cleaning and accumulation in the following categories: geographic addresses and geo-coordinates, validation and automated addition of a mobile phone number to the international format, processing of car numbers (in modern and outdated format), VIN code of the engine and car brand, validation of urls of social networks, passport data and processing of personal data. Compared to similar methods for processing large volumes of data, the possibility of splitting the input file or stream into separate parts was used, the cleaned data from which is combined at the end of the system operation. Thanks to this, it is possible to process data whose size exceeds the available volume of the device’s RAM, and the method of working with loosely structured text files in CSV format has been improved. [ABSTRACT FROM AUTHOR]

Subjects :: *DATABASES
*ALGORITHMS
*PARALLEL processing
*TEXT files
*INFORMATION processing
*UNIFORM Resource Locators
*ELECTRONIC data processing
*TELEPHONE numbers

Details

Language :: English
ISSN :: 26649969
Volume :: 3
Issue :: 2(77)
Database :: Academic Search Index
Journal :: Technology Audit & Production Reserves
Publication Type :: Academic Journal
Accession number :: 178197765
Full Text :: https://doi.org/10.15587/2706-5448.2024.306435

Full Text Access

View/download PDF

Tools

Email
Cite

Printer

Authors Abstract Subjects Details

Searchworks

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources

ANALYSIS OF METHODS AND ALGORITHMS FOR PROCESSING UNSTRUCTURED TEXT DATA BASED ON JSON TECHNOLOGY.

Abstract

Subjects

Details

Tools

Searchworks

Select search scope, currently: Articles Catalog books, media & more in Jio Institute collections Articles journal articles & other e-resources

ANALYSIS OF METHODS AND ALGORITHMS FOR PROCESSING UNSTRUCTURED TEXT DATA BASED ON JSON TECHNOLOGY.

Abstract

Subjects

Details

Tools

Select search scope, currently: Articles

Catalog

books, media & more in Jio Institute collections

Articles

journal articles & other e-resources