Evaluasi Multi-Dimensi Sebelas Model Pembelajaran Mesin untuk Klasifikasi Kekuatan Kata Sandi pada Lingkungan Sumber Daya Terbatas

Authors

  • Muhammad Ali Sofian Universitas Teknologi Digital Indonesia
  • Widyastuti Andriyani Universitas Teknologi Digital Indonesia

DOI:

https://doi.org/10.56427/jcbd.v5i2.998

Keywords:

Keamanan kata sandi, Pembelajaran mesin, Klasifikasi, keamanan siber, Pohon keputusan

Abstract

Kata sandi tetap menjadi mekanisme autentikasi paling banyak digunakan, namun mayoritas pelanggaran data tetap bersumber dari kredensial lemah. Studi ini mengevaluasi sebelas model pembelajaran mesin mencakup pendekatan klasik (Logistic Regression, Decision Tree, SVM, Random Forest) dan modern (XGBoost, LightGBM, CatBoost, Extra Trees, MLPClassifier, TabNet, Ensemble Stacking) pada dua dimensi evaluasi: kinerja prediktif dan efisiensi sistem. Dari korpus 1,6 juta kata sandi, diekstraksi sampel berimbang 30.000 entri (10.000 per kelas) dengan delapan fitur terinterpretasi. Temuan kritis adalah seluruh model mencapai akurasi mendekati 1,000 pada himpunan uji bukan bukti generalisasi, melainkan konsekuensi deterministik dari label berbasis aturan pada dataset Kaggle; sehingga sumbu pembeda model bergeser dari akurasi ke efisiensi. Distribusi latensi inferensi pada n=1.000 sampel uji menunjukkan Decision Tree memberikan keseimbangan terbaik (rerata 0,11 ms; persentil-95 0,18 ms; ukuran 1,06 KB), jauh di bawah ambang 100 ms yang direkomendasikan NIST SP 800-63B. Kontribusi utama studi ini adalah kerangka evaluasi multi-dimensi yang reproducible dengan kriteria efisiensi terkuantifikasi, bukan klaim akurasi yang trivial.

Downloads

Download data is not yet available.

References

[1] Verizon, “2024 Data Breach Investigations Report,” Verizon Business, 2024. [Online]. Available: https://www.verizon.com/business/resources/reports/2024-dbir-data-breach-investigations-report.pdf

[2] W. Melicher, B. Ur, S. M. Segreti, S. Komanduri, L. Bauer, N. Christin, et al., “Fast, Lean, and Accurate: Modeling Password Guessability Using Neural Networks,” in Proceedings of the 25th USENIX Security Symposium, Austin, TX, USA, 2016, pp. 175–191. [Online]. Available: https://www.usenix.org/conference/usenixsecurity16/technical-sessions/presentation/melicher

[3] S. Ji, S. Yang, X. Hu, W. Han, Z. Li, and R. Beyah, “Zero-Sum Password Cracking Game: A Large-Scale Empirical Study on the Crackability, Correlation, and Security of Passwords,” IEEE Transactions on Dependable and Secure Computing, vol. 14, no. 5, pp. 550–564, 2017, doi: 10.1109/TDSC.2015.2481884.

[4] P. G. Kelley, S. Komanduri, M. L. Mazurek, R. Shay, T. Vidas, L. Bauer, et al., “Guess Again (and Again and Again): Measuring Password Strength by Simulating Password-Cracking Algorithms,” in Proceedings of the 2012 IEEE Symposium on Security and Privacy, San Francisco, CA, USA, 2012, pp. 523–537, doi: 10.1109/SP.2012.38.

[5] M. Ozkan-Okay, E. Akin, Ö. Aslan, S. Kosunalp, T. Iliev, I. Stoyanov, et al., “A Comprehensive Survey: Evaluating the Efficiency of Artificial Intelligence and Machine Learning Techniques on Cyber Security Solutions,” IEEE Access, vol. 12, pp. 12229–12256, 2024, doi: 10.1109/ACCESS.2024.3355547.

[6] B. Bansal, “Password Strength Classifier Dataset,” Kaggle, 2019. [Online]. Available: https://www.kaggle.com/datasets/bhavikbb/password-strength-classifier-dataset

[7] J. Mo, H. Kuang, and X. Li, “Password Strength Detection via Machine Learning: Analysis, Modeling, and Evaluation,” arXiv preprint, arXiv:2505.16439, May 2025, doi: 10.48550/arXiv.2505.16439.

[8] M. E. M. Mazelan, N. H. A. Mutalib, and N. AlDahoul, “Enhancing Password Security Through a High-Accuracy Scoring Framework Using Random Forests,” arXiv preprint, arXiv:2511.09492, Nov. 2025, doi: 10.48550/arXiv.2511.09492.

[9] D. L. Wheeler, “zxcvbn: Low-Budget Password Strength Estimation,” in Proceedings of the 25th USENIX Security Symposium, Austin, TX, USA, 2016, pp. 157–173. [Online]. Available: https://www.usenix.org/conference/usenixsecurity16/technical-sessions/presentation/wheeler

[10] B. Ur, P. G. Kelley, S. Komanduri, J. Lee, M. Maass, M. L. Mazurek, et al., “How Does Your Password Measure Up? The Effect of Strength Meters on Password Creation,” in Proceedings of the 21st USENIX Security Symposium, Bellevue, WA, USA, 2012, pp. 65–80. [Online]. Available: https://www.usenix.org/conference/usenixsecurity12/technical-sessions/presentation/ur

[11] S. Komanduri, R. Shay, P. G. Kelley, M. L. Mazurek, L. Bauer, N. Christin, et al., “Of Passwords and People: Measuring the Effect of Password-Composition Policies,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI ’11), Vancouver, BC, Canada, 2011, pp. 2595–2604, doi: 10.1145/1978942.1979321.

[12] P. A. Grassi, J. L. Fenton, E. M. Newton, R. A. Perlner, A. R. Regenscheid, W. E. Burr, et al., “Digital Identity Guidelines: Authentication and Lifecycle Management,” NIST Special Publication 800-63B, National Institute of Standards and Technology, Gaithersburg, MD, USA, 2017, doi: 10.6028/NIST.SP.800-63b.

[13] C. E. Shannon, “A Mathematical Theory of Communication,” Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948, doi: 10.1002/j.1538-7305.1948.tb01338.x.

[14] L. Breiman, “Random Forests,” Machine Learning, vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.

[15] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ’16), San Francisco, CA, USA, 2016, pp. 785–794, doi: 10.1145/2939672.2939785.

[16] G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, et al., “LightGBM: A Highly Efficient Gradient Boosting Decision Tree,” in Advances in Neural Information Processing Systems 30 (NeurIPS 2017), Long Beach, CA, USA, 2017, pp. 3146–3154. [Online]. Available: https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree

[17] S. Ö. Arik and T. Pfister, “TabNet: Attentive Interpretable Tabular Learning,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 8, pp. 6679–6687, May 2021, doi: 10.1609/aaai.v35i8.16826.

[18] J. Beach, “Top 10 Million Passwords Dataset,” Kaggle, 2023. [Online]. Available: https://www.kaggle.com/datasets/joebeachcapital/top-10-million-passwords

[19] G. James, D. Witten, T. Hastie, and R. Tibshirani, An Introduction to Statistical Learning: with Applications in R, 2nd ed. New York, NY, USA: Springer, 2021, doi: 10.1007/978-1-0716-1418-1.

[20] C. M. Bishop, Pattern Recognition and Machine Learning. New York, NY, USA: Springer, 2006.

[21] P. Geurts, D. Ernst, and L. Wehenkel, “Extremely Randomized Trees,” Machine Learning, vol. 63, no. 1, pp. 3–42, 2006, doi: 10.1007/s10994-006-6226-1.

[22] L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, “CatBoost: Unbiased Boosting with Categorical Features,” in Advances in Neural Information Processing Systems 31 (NeurIPS 2018), Montréal, QC, Canada, 2018, pp. 6638–6648. [Online]. Available: https://papers.nips.cc/paper/7898-catboost-unbiased-boosting-with-categorical-features

[23] W. McKinney, “Data Structures for Statistical Computing in Python,” in Proceedings of the 9th Python in Science Conference, Austin, TX, USA, 2010, pp. 56–61, doi: 10.25080/Majora-92bf1922-00a.

[24] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011. [Online]. Available: https://www.jmlr.org/papers/v12/pedregosa11a.html

[25] F. Yu, “On Deep Learning in Password Guessing, a Survey,” arXiv preprint, arXiv:2208.10413, 2022, doi: 10.48550/arXiv.2208.10413.

[26] A. C. Nurcahyo, H. Y. Ting, and A. F. Atanda, “Network Log Implementation for GRU Based Bandwidth Classification,” Journal of Computers and Digital Business, vol. 4, no. 2, pp. 76–89, 2025, doi: 10.56427/jcbd.v4i2.763.

Downloads

Published

29-05-2026

How to Cite

Sofian, M. A., & Widyastuti Andriyani. (2026). Evaluasi Multi-Dimensi Sebelas Model Pembelajaran Mesin untuk Klasifikasi Kekuatan Kata Sandi pada Lingkungan Sumber Daya Terbatas. Journal of Computers and Digital Business, 5(2), 99–108. https://doi.org/10.56427/jcbd.v5i2.998

Issue

Section

Articles