IJISA Vol. 18, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 825KB)
PDF (825KB), PP.226-238
Views: 0 Downloads: 0
Topic Modeling, Sentiment Analysis, News Headlines, BERTopic, IndoBERT
The high volume of information from online media in Indonesia poses challenges for manual analysis in identifying emerging themes and sentiments. News headlines, as the primary element seen by the public, play a crucial role in shaping opinion; however, their massive volume and diverse themes necessitate automated approaches to identify topics and underlying sentiments. To address this challenge, this study analyzed 30,329 news headlines from the online news portal detik.com for the entire year 2024. A quantitative Natural Language Processing (NLP) framework was applied, comprising data collection via web scraping, text preprocessing, transformer-based topic modeling using BERTopic, sentiment classification using IndoBERT, and a topic-sentiment intersection analysis. Preprocessing included case folding, text cleaning, normalization of informal words, and tokenization. For lexicon-based labeling, stopword removal and stemming were applied. In contrast, transformer-based models utilized text that underwent only case folding, cleaning, and normalization to preserve contextual information. Topic modeling was performed using BERTopic, while sentiment classification (positive, negative, and neutral) used the IndoBERT model. The main objective of this study was to evaluate the combined performance of the two models in mapping dominant issues and the sentiments contained in media reports. The results showed that BERTopic successfully identified 366 topics. An evaluation of the 10 most dominant topics yielded a coherence score of 0.5145, indicating a relevant topic clustering. IndoBERT demonstrated high agreement with lexicon-generated sentiment labels, with an accuracy of 94.78%, a precision of 95.04%, a recall of 94.79%, and an F1-score of 94.81%. These findings confirm that the combination of transformer-based models is effective for in-depth discourse analysis of political news headlines from detik.com, a major Indonesian online news portal.
Bagas Yana Prayoga, Qurrotul Aini, Fitroh Fitroh, "Topic Modeling and Sentiment Analysis on News Headlines Using BERTopic and IndoBERT Models", International Journal of Intelligent Systems and Applications(IJISA), Vol.18, No.4, pp.226-238, 2026. DOI:10.5815/ijisa.2026.04.12
[1]A. D. Riyanto, “Hootsuite (We Are Social): Data Report Digital Indonesia 2024.” Accessed: Apr. 24, 2025. [Online]. Available: https://andi.link/hootsuite-we-are-social-data-digital-indonesia-2024/
[2]A. Yaman, B. Sartono, and A. M. Soleh, “Topic modeling in fertilizer-related patent documents in Indonesia based on latent Dirichlet allocation,” (in Indonesian), Berk. Ilmu Perpust. dan Inf., vol. 17, no. 2, pp. 168–180, 2021, doi: 10.22146/bip.v17i2.2147.
[3]Z. Alhaq, A. Mustopa, S. Mulyatun, and J. D. Santoso, “Application of support vector machine method for Twitter user sentiment analysis,” (in Indonesian), J. Inf. Syst. Manag., vol. 3, no. 1, pp. 16–21, 2021, doi: 10.24076/joism.2021v3i2.558.
[4]R. Egger and J. Yu, “A topic modeling comparison between LDA, NMF, Top2Vec, and BERTopic to demystify Twitter posts,” Front. Sociol., vol. 7, p. 886498, 2022, doi: 10.3389/fsoc.2022.886498.
[5]D. Hendry et al., “Topic modeling for customer service chats,” in 2021 International Conference on Advanced Computer Science and Information Systems (ICACSIS), IEEE, 2021, pp. 1–6, doi: 10.1109/ICACSIS53237.2021.9631322.
[6]Y. Liu, “Comparison of LDA and BERTopic in news topic modeling: A case study of the New York Times’ reports on China,” Pacific Int. J., vol. 7, no. 3, pp. 47–51, 2024, doi: 10.55014/pij.v7i3.616.
[7]B. Wilie et al., “IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding,” arXiv preprint arXiv:2009.05387, 2020, [Online]. Available: http://arxiv.org/abs/2009.05387
[8]Fransiscus and A. S. Girsang, “Sentiment analysis of COVID-19 public activity restriction (PPKM) impact using BERT method,” Int. J. Eng. Trends Technol., vol. 70, no. 12, pp. 281–288, 2022, doi: 10.14445/22315381/IJETT-V70I12P226.
[9]W. Nurfitri and A. Chowanda, “Sentiment analysis on COVID-19 positive cases based on media coverage in Indonesia using IndoBERT,” (in Indonesian), Progresif J. Ilm. Komput., vol. 20, no. 1, pp. 580–593, 2024, doi: 10.35889/progresif.v20i1.1897.
[10]N. Mahfudiyah and A. Alamsyah, “Understanding user perception of ride-hailing services sentiment analysis and topic modelling using IndoBERT and BERTopic,” in 2023 International Conference on Digital Business and Technology Management (ICONDBTM), IEEE, 2023, pp. 1–6, doi: 10.1109/ICONDBTM59210.2023.10327320.
[11]E. Effendi, I. Sartika, N. L. T. Purba, and S. Ritonga, “Writing headlines and leads for news and features,” (in Indonesian), J. Pendidik. dan Konseling, vol. 5, no. 2, pp. 4680–4683, 2023, doi: 10.31004/jpdk.v5i2.14206.
[12]E. C. D. A. Nanda, “Apa Konten Berita yang Paling Banyak Diakses Warganet 2024?” (in Indonesian), Goodstats, 2024. [Online]. Available: https://goodstats.id/article/apa-konten-berita-yang-paling-banyak-diakses-warganet-pada-2024-B5hDY. Accessed: Jan. 16, 2025.
[13]N. Newman, R. Fletcher, C. T. Robertson, A. R. Arguedas, and R. K. Nielsen, “Reuters Institute Digital News Report 2024,” 2024. [Online]. Available: https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2024
[14]M. Han, I. Canli, J. Shah, X. Zhang, I. G. Dino, and S. Kalkan, “Perspectives of machine learning and natural language processing on characterizing positive energy districts,” Buildings, vol. 14, no. 2, pp. 1–28, 2024, doi: 10.3390/buildings14020371.
[15]G. S. Mahendra et al., Tren Teknologi AI: Pengantar, Teori, dan Contoh Penerapan Artificial Intelligence di Berbagai Bidang (in Indonesian). PT. Sonpedia Publishing Indonesia, 2024.
[16]D. L. C. Pardede and M. A. I. Waskita, “Topic modeling analysis for reviews about PeduliLindungi,” (in Indonesian), J. Ilm. Inform. Komput., vol. 28, no. 1, pp. 17–26, 2023, doi: 10.35760/ik.2023.v28i1.7925.
[17]S. Alaparthi and M. Mishra, “Bidirectional encoder representations from transformers (BERT): A sentiment analysis odyssey,” arXiv preprint arXiv:2007.01127, 2020, [Online]. Available: https://arxiv.org/abs/2007.01127
[18]M. Grootendorst, “BERTopic: Neural topic modeling with a class-based TF-IDF procedure,” arXiv preprint arXiv:2203.05794, 2022, [Online]. Available: https://arxiv.org/abs/2203.05794
[19]B. Liu, Sentiment Analysis and Opinion Mining. Chicago: Morgan & Claypool Publisher, 2020. doi: 10.1017/9781108639286.
[20]S. B. Panuntun, D. Krismawati, S. Pramana, and E. T. Astuti, “Analysis of telemedicine news coverage in Indonesia: Sentiment, NER, topic modeling, and social network approaches to understanding issues and perceptions,” (in Indonesian), Indones. Heal. Inf. Manag. J., vol. 11, no. 1, pp. 56–67, 2023, doi: 10.47007/inohim.v11i1.500.
[21]F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020), 2020, pp. 757–770, doi: 10.18653/v1/2020.coling-main.66.
[22]A. Simanjuntak et al., “Study and analysis of IndoBERT hyperparameter tuning in fake news detection,” (in Indonesian), J. Nas. Tek. Elektro dan Teknol. Inf., vol. 13, no. 1, pp. 60–67, 2024, doi: 10.22146/jnteti.v13i1.8532.
[23]S. Sazzed and S. Jayarathna, “SSentiA: A self-supervised sentiment analyzer for classification from unlabeled data,” Mach. Learn. with Appl., vol. 4, pp. 1–14, 2021, doi: 10.1016/j.mlwa.2021.100026.
[24]R. Mohammed, J. Rawashdeh, and M. Abdullah, “Machine learning with oversampling and undersampling techniques: Overview study and experimental results,” in 2020 11th International Conference on Information and Communication Systems (ICICS), IEEE, 2020, pp. 243–248, doi: 10.1109/ICICS49469.2020.239556.
[25]Y. Yunitasari and A. R. Putera, “Analysis of public sentiment on Twitter regarding the COVID-19 pandemic,” (in Indonesian), Smatika J., vol. 11, no. 01, pp. 22–26, 2021, doi: 10.32664/smatika.v11i01.520.
[26]B. P. E. Wijaya, D. P. Hostiadi, and P. D. W. Ayu, “Sentiment analysis of Instagram comments for cyberbullying detection using gradient boosting models,” (in Indonesian) in Seminar Hasil Penelitian Informatika dan Komputer (SPINTER), Institut Teknologi dan Bisnis STIKOM Bali, 2025, pp. 1093–1098.
[27]D. Aryani, I. Lucia Kharisma, A. Sujjada, and K. Kamdan, “Topic modeling of the 2024 election using the BERTopic method on detik.com news articles,” Inf. J. Ilm. Bid. Teknol. Inf. dan Komun., vol. 9, no. 2, pp. 171–180, 2024, doi: 10.25139/inform.v9i2.8429.
[28]I. N. Nugraha and E. Utami, “Evaluation of creative economy and tourism industry trends based on LDA analysis with BERTopic,” Digit. Zo. J. Teknol. Inf. dan Komun., vol. 15, no. 2, pp. 182–195, 2024, doi: 10.31849/digitalzone.v15i2.23796.
[29]A. Rahmawati, A. Alamsyah, and A. Romadhony, “Hoax news detection analysis using IndoBERT deep learning methodology,” in 2022 10th International Conference on Information and Communication Technology (ICoICT), 2022, pp. 368–373. doi: 10.1109/ICoICT55009.2022.9914902.
[30]P. Sayarizki, Hasmawati, and H. Nurrahmi, “Implementation of IndoBERT for sentiment analysis of indonesian presidential candidates,” J. Comput., vol. 9, no. 2, pp. 61–72, 2024, doi: 10.34818/indojc.2024.9.2.934.