Data storage format for analytical systems based on metadata and dependency graphs between CSV and JSON

Aleksey Nikolaevich Alpatov; Алпатов Алексей Николаевич; Anna Alekseevna Bogatireva; Богатырева Анна Алексеевна

doi:10.7256/2454-0714.2024.2.70229

Data storage format for analytical systems based on metadata and dependency graphs between CSV and JSON

Авторлар: Alpatov A.N.¹, Bogatireva A.A.¹
Мекемелер:
Шығарылым: № 2 (2024)
Беттер: 1-14
Бөлім: Articles
URL: https://journal-vniispk.ru/2454-0714/article/view/359412
DOI: https://doi.org/10.7256/2454-0714.2024.2.70229
EDN: https://elibrary.ru/TVEPRE
ID: 359412

Дәйексөз келтіру

Толық мәтін

Аннотация
Авторлар туралы
Әдебиет тізімі
Қосымша файлдар
Статистика

Аннотация

In the modern information society, the volume of data is constantly growing, and its effective processing is becoming key for enterprises. The transmission and storage of this data also plays a critical role. Big data used in analytics systems is most often transmitted in one of two popular formats: CSV for structured data and JSON for unstructured data. However, existing file formats may not be effective or flexible enough for certain data analysis tasks. For example, they may not support complex data structures or provide sufficient control over metadata. Alternatively, analytical tasks may require additional information about the data, such as metadata, data schema, etc. Based on the above, the subject of this study is a data format based on the combined use of CSV and JSON for processing and analyzing large amounts of information. The option of sharing the designated data types for the implementation of a new data format is proposed. For this purpose, designations have been introduced for the data structure, which includes CSV files, JSON files, metadata and a dependency graph. Various types of functions are described, such as aggregating, transforming, filtering, etc. Examples of the application of these functions to data are given. The proposed approach is a technique that can significantly facilitate the processes of information analysis and processing. It is based on a formalized approach that allows you to establish clear rules and procedures for working with data, which contributes to their more efficient processing. Another aspect of the proposed approach is to determine the criteria for choosing the most appropriate data storage format. This criterion is based on the mathematical principles of information theory and entropy. The introduction of a criterion for choosing a data format based on entropy makes it possible to evaluate the information content and compactness of the data. This approach is based on the calculation of entropy for selected formats and weights reflecting the importance of each data value. By comparing entropies, you can determine the required data transmission format. This approach takes into account not only the compactness of the data, but also the context of their use, as well as the possibility of including additional meta-information in the files themselves and supporting data ready for analysis.

Негізгі сөздер

Data storage formats, JSON, CSV, Analysis Ready Data, Metadata, Data processing, Data analysis, Integration of data formats, Apache Parquet, Big Data

Авторлар туралы

Aleksey Alpatov

Email: aleksej01-91@mail.ru
ORCID iD: 0000-0001-8624-1662

Anna Bogatireva

Email: pecherni@gmail.com

Әдебиет тізімі

Malcolm R., Morrison C., Grandison T., Thorpe S., Christie K., Wallace A., Green D., Jarrett J., Campbell A. Increasing the accessibility to big data systems via a common services api // IEEE International Conference on Big Data. 2014. Pp. 883-892.
Wu T. System of teaching quality analyzing and evaluating based on data warehouse // Computer Engineering and Design. 2009. No. 6(2). Pp. 1545-1547.
Vitagliano G. et al. Pollock: A Data Loading Benchmark // Proceedings of the VLDB Endowment. 2023. No. 8(16). Pp. 1870-1882.
Xiaojuan L., Yu Z. A data integration tool for the integrated modeling and analysis for east // Fusion Engineering and Design. 2023. No. 195. Pp. 113933. URL: https://doi.org/10.1016/j.fusengdes.2023.113933
Lemzin A. Streaming Data Processing // Asian Journal of Research in Computer Science. 2023. No. 1(15). Pp. 11-21.
Hughes LD, Tsueng G, DiGiovanna J, Horvath TD, Rasmussen LV, Savidge TC, Stoeger T, Turkarslan S, Wu Q, Wu C, Su AI, Pache L. Addressing barriers in FAIR data practices for biomedical data // Scientific Data. 2023. No. 1(10). P. 98. DOI: https://doi.org/10.1038/s41597-023-01969-8
Gohil A., Shroff A., Garg A., Kumar S. A Compendious Research on Big Data File Formats. "em"2022 6th International Conference on Intelligent Computing and Control Systems (ICICCS)."/em" IEEE Press, Madurai, India. 2022. Pp. 905-913. DOI: https://doi.org/10.1109/ICICCS53718.2022.9788141
Елсуков П. Ю. Информационная асимметрия и информационная неопределенность // ИТНОУ: Информационные технологии в науке, образовании и управлении. 2017. No. 4 (4). С. 69-76.
Bromiley P. A., Thacker N. A., Bouhova-Thacker E. Shannon entropy, Renyi entropy, and information // Statistics and Inf. Series (2004-004). 2004. No. 9. Pp. 2-8.
Dwyer, J. L. Roy, D. P., Sauer B., Jenkerson C. B., Zhang H. K., Lymburner L. Analysis ready data: enabling analysis of the Landsat archive // Remote Sensing. 2018. №. 9(10). 1363.

Қосымша файлдар

Әрекет

1. JATS XML

Жүктеу

Пайдаланушының аты
Құпиясөз
Мені есте сақтау

Құпия сөзді ұмыттыңыз ба?	Тіркеу

Пайдаланушының аты
Құпиясөз
Мені есте сақтау

Құпия сөзді ұмыттыңыз ба?	Тіркеу

№ 3 (2025)

№ 3 (2025)

Data storage format for analytical systems based on metadata and dependency graphs between CSV and JSON

Толық мәтін

Аннотация

Негізгі сөздер

Авторлар туралы

Aleksey Alpatov

Anna Bogatireva

Әдебиет тізімі

Қосымша файлдар