YouTube Sentiment Analysis on Felt Earthquake News Using LSTM and IndoBERT

Authors

  • Oktifar Tri Bandono Universitas Pamulang
  • Agung Budi Susanto Informatics Engineering, Computer Science Faculty, Universitas Pamulang, Indonesia
  • Makhsun Makhsun Informatics Engineering, Computer Science Faculty, Universitas Pamulang, Indonesia

DOI:

https://doi.org/10.55227/ijhess.v6i1.2420

Keywords:

sentiment analysis, earthquake, YouTube, LSTM, IndoBERT

Abstract

This study analyzes public sentiment in YouTube comments on felt-earthquake news in Indonesia and compares a bidirectional Long Short-Term Memory implementation (LSTM) with IndoBERT. An experimental quantitative design was used. Comments were collected through the YouTube Data API v3 using official earthquake-event references from the Indonesian Agency for Meteorology, Climatology, and Geophysics (BMKG) and keyword-based video searches covering January 2021 to August 2025. The acquisition stage produced 51,870 comments from 1,626 unique videos. Data were processed through duplicate removal, text cleaning, case folding, slang normalization, tokenization, and stopword removal, resulting in 49,041 clean comments. Positive, negative, and neutral labels were assigned by aggregating word-polarity scores from the Indonesian Sentiment Lexicon (InSet), which served as weak supervision. Stratified sampling divided the dataset into 80% training data and 20% testing data. The LSTM model used a 100-dimensional embedding, a 64-unit bidirectional LSTM layer, global max pooling, and early stopping; IndoBERT was fine-tuned from indobenchmark/indobert-base-p2 for four epochs. Performance was assessed with accuracy, macro precision, macro recall, macro F1-score, and confusion matrices. Positive sentiment accounted for 41.2% of the corpus, negative sentiment for 35.5%, and neutral sentiment for 23.2%. IndoBERT achieved 91.50% accuracy and a 91.03% macro F1-score, outperforming LSTM at 90.91% accuracy and a 90.34% macro F1-score. IndoBERT provided the strongest contextual classification, while LSTM remained a competitive and substantially lighter option for resource-constrained monitoring

References

Bird, P. (2003). An updated digital model of plate boundaries. Geochemistry, Geophysics, Geosystems, 4(3), 1027. https://doi.org/10.1029/2001GC000252

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171-4186). https://doi.org/10.18653/v1/N19-1423

Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735-1780. https://doi.org/10.1162/neco.1997.9.8.1735

Koto, F., Rahmaningtyas, F. D., & Louvan, S. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of AACL-IJCNLP 2020 (pp. 843-857). https://arxiv.org/abs/2009.05387

Medhat, W., Hassan, A., & Korashy, H. (2014). Sentiment analysis algorithms and applications: A survey. Ain Shams Engineering Journal, 5(4), 1093-1113. https://doi.org/10.1016/j.asej.2014.04.011

Nursiyono, J. A., & Gibran, R. K. (2023). Natural language processing for unstructured data: Earthquakes spatial analysis in Indonesia using Twitter. Innovatics, 5(1), 22-29.

Safitri, Y. D., Faisal, M. R., Kartini, D., Saragih, T. H., Abadi, F., & Bachtiar, A. M. (2025). Automatic analysis of natural disaster messages on social media using IndoBERT and multilingual BERT. Telematika, 18(2), 105-120. https://doi.org/10.35671/telematika.v18i2.3140

Sokolova, M., & Lapalme, G. (2009). A systematic analysis of performance measures for classification tasks. Information Processing & Management, 45(4), 427-437. https://doi.org/10.1016/j.ipm.2009.03.002

Taboada, M., Brooke, J., Tofiloski, M., Voll, K., & Stede, M. (2011). Lexicon-based methods for sentiment analysis. Computational Linguistics, 37(2), 267-307. https://doi.org/10.1162/COLI_a_00049

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998-6008. https://arxiv.org/abs/1706.03762

Yefferson, D. Y., Lawijaya, V., & Girsang, A. S. (2024). Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news. IAES International Journal of Artificial Intelligence, 13(2), 1913-1924. https://doi.org/10.11591/ijai.v13.i2.pp1913-1924

Zhang, L., Wang, S., & Liu, B. (2022). Deep learning for sentiment analysis: A survey. WIREs Data Mining and Knowledge Discovery, 12(1), e1435. https://doi.org/10.1002/widm.1435

Downloads

Published

2026-08-12

How to Cite

Bandono, O. T., Agung Budi Susanto, & Makhsun Makhsun. (2026). YouTube Sentiment Analysis on Felt Earthquake News Using LSTM and IndoBERT. International Journal Of Humanities Education and Social Sciences (IJHESS), 6(1). https://doi.org/10.55227/ijhess.v6i1.2420

Issue

Section

Social Science