Benchmarking Konfigurasi Apache Spark Single-Node: Studi Format, Partitioning, dan Caching pada Dataset Log Sintetis

Authors

  • Syahrur Ro'uf Universitas Teknologi Digital Indonesia
  • Bambang Purnomosidi Dwi Putranto Universitas Teknologi Digital Indonesia

DOI:

https://doi.org/10.56427/jcbd.v5i3.1015

Keywords:

Apache Spark , benchmarking single-node, format Parquet , shuffle partitioning, in-memory caching

Abstract

Studi benchmarking Apache Spark umumnya dilakukan pada klaster multi-node berskala komersial, sehingga panduan konfigurasi untuk mode single-node (local mode) yang lazim dipakai peneliti dengan sumber daya terbatas masih jarang diverifikasi secara empiris dan terbuka. Penelitian ini mengisi kesenjangan tersebut dengan mengevaluasi pipeline data lima tahap (ingest, cleanse, transform, aggregate, dan sink) menggunakan Apache Spark 4.1 pada mesin single-node (4 core, memori driver 4 GB) untuk mengolah dataset log transaksional sintetis 5 juta rekaman (424 MB) yang meniru karakteristik data akses terbuka. Empat eksperimen terkendali dilakukan: (E1) format penyimpanan CSV versus Parquet, (E2) strategi shuffle partitioning, (E3) in-memory caching untuk beban kerja iteratif, dan (E4) profiling waktu eksekusi per tahap. Pada lingkungan uji tersebut, Parquet memberikan percepatan baca 4,03× dengan rasio kompresi 2,92× dibanding CSV; jumlah partisi optimal adalah 4, sesuai jumlah core CPU; in-memory caching menghasilkan percepatan kumulatif 2,22× pada empat kueri berulang; dan pipeline end-to-end mencapai throughput 26.130 rekaman per detik nilai yang spesifik terhadap mesin uji dengan tahap sink Parquet terpartisi sebagai bottleneck dominan (75,1% total waktu). Kontribusi bersifat konfirmatoris-empiris: memverifikasi secara terkendali panduan konfigurasi yang terdokumentasi namun jarang diukur pada konteks single-node, sekaligus menyediakan kode sumber dan generator dataset secara terbuka untuk mendukung reproducibility.

Downloads

Download data is not yet available.

References

[1] E. Papachristou and E. Gounopoulos, “A systematic literature review of open government data: Assessing concepts, trends, and current research gaps,” AIP Conf. Proc., vol. 3220, no. 1, Oct. 2024, doi: 10.1063/5.0234826.

[2] C. Alexopoulos, S. Saxena, M. Janssen, N. Rizun, M. Lnenicka, and R. Matheus, “Why do Open Government Data initiatives fail in developing countries? A root cause analysis of the most prevalent barriers and problems,” THE ELECTRONIC JOURNAL OF INFORMATION SYSTEMS IN DEVELOPING COUNTRIES, vol. 90, no. 2, p. e12297, Mar. 2024, doi: 10.1002/isd2.12297.

[3] M. Zaharia et al., “Apache Spark: a unified engine for big data processing,” Association for Computing Machinery, vol. 59, no. 11, pp. 56–65, Nov. 2016, doi: 10.1145/2934664.

[4] J. Dean and S. Ghemawat, “MapReduce: simplified data processing on large clusters,” Association for Computing Machinery, vol. 51, no. 1, pp. 107–113, Jan. 2008, doi: 10.1145/1327452.132749.

[5] M. Zaharia et al., “Resilient distributed datasets: a fault-tolerant abstraction for in-memory cluster computing,” Proceedings of the 9th USENIX Conference on Networked Systems Design and Implementation, Apr. 2012, doi: 10.5555/2228298.2228301.

[6] M. Armbrust et al., “Spark SQL: Relational Data Processing in Spark,” Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, pp. 1383–1394, 2015, doi: 10.1145/2723372.2742797.

[7] A. Antolínez García, “Introduction to Apache Spark for Large-Scale Data Analytics,” in Hands-on Guide to Apache Spark 3 : Build Scalable Computing Engines for Batch and Stream Data Processing, Apress, 2023, pp. 3–21. doi: 10.1007/978-1-4842-9380-5_1.

[8] T. Shwe and M. Aritsugi, “Optimizing Data Processing: A Comparative Study of Big Data Platforms in Edge, Fog, and Cloud Layers,” Applied Sciences, vol. 14, no. 1, p. 452, 2024, doi: 10.3390/app14010452.

[9] S. Puthenpariyarath, “Big Data Analytics: Performance Tuning in Apache Spark,” International Journal of Computer Engineering and Technology, vol. 16, no. 2, pp. 99–117, 2025, doi: 10.34218/IJCET_16_02_006.

[10] X. Zeng, Y. Hui, J. Shen, A. Pavlo, W. McKinney, and H. Zhang, “An Empirical Evaluation of Columnar Storage Formats,” Proceedings of the VLDB Endowment, vol. 17, no. 2, pp. 148–161, 2023, doi: 10.14778/3626292.3626298.

[11] O. Oloruntoba, D. O. Oyeyemi, and O. Omolayo, “Designing Scalable ETL Pipelines for Multi-Source Graph Database Ingestion,” Journal of Computer Analysis and Applications, vol. 34, no. 7, pp. 236–258, 2025, [Online]. Available: https://eudoxuspress.com/index.php/pub/article/view/3336

[12] L. Theodorakopoulos, A. Karras, A. Theodoropoulou, and G. Kampiotis, “Benchmarking Big Data Systems: Performance and Decision-Making Implications in Emerging Technologies,” Technologies (Basel)., vol. 12, no. 11, p. 217, 2024, doi: 10.3390/technologies12110217.

[13] G. R. Fim, G. Mencagli, and D. Griebler, “Benchmarking Batch and Stream Processing Execution Modes in Apache Flink,” Computing, vol. 108, no. 1, p. 21, 2026, doi: 10.1007/s00607-025-01608-7.

[14] L. Theodorakopoulos, A. Karras, and G. A. Krimpas, “Optimizing Apache Spark MLlib: Predictive Performance of Large-Scale Models for Big Data Analytics,” Algorithms, vol. 18, no. 2, p. 74, 2025, doi: 10.3390/a18020074.

[15] E. Dritsas and M. Trigka, “Applying Machine Learning on Big Data with Apache Spark,” IEEE Access, vol. 13, pp. 53377–53393, 2025, doi: 10.1109/ACCESS.2025.3552042.

Downloads

Published

15-08-2026

How to Cite

Syahrur Ro’uf, & Bambang Purnomosidi Dwi Putranto. (2026). Benchmarking Konfigurasi Apache Spark Single-Node: Studi Format, Partitioning, dan Caching pada Dataset Log Sintetis. Journal of Computers and Digital Business, 5(3), 234–241. https://doi.org/10.56427/jcbd.v5i3.1015