IJISA Vol. 18, No. 4, 8 Aug. 2026
Cover page and Table of Contents: PDF (size: 832KB)
PDF (832KB), PP.226-238
Views: 0 Downloads: 0
Topic Modeling, Sentiment Analysis, News Headline, BERTopic, IndoBERT
The high flow of information from online media in Indonesia makes it difficult for manual analysis to identify emerging themes and sentiments. News headlines, as the first element seen by the public, play a crucial role in shaping opinion, but their massive volume and diverse themes make it difficult for manual analysis to identify topics and their underlying sentiments. To address this challenge, this study analyzed 30,329 news headlines from the online news portal detik.com for the entire year 2024. A quantitative Natural Language Processing (NLP) framework was applied, consisting of data collection through web scraping, text preprocessing, transformer-based topic modeling using BERTopic, sentiment classification using IndoBERT, and a topic sentiment intersection analysis. Preprocessing included case folding, text cleaning, normalization of informal words, and tokenization. For lexicon-based labeling, stopword removal and stemming were applied, while transformer-based models utilized minimally processed text to preserve contextual information. Topic modeling was performed using BERTopic, while sentiment classification (positive, negative, and neutral) used the IndoBERT model. The main objective of this study was to evaluate the combined performance of the two models in mapping dominant issues and the sentiments contained in media reports. The results showed that BERTopic successfully identified 366 topics. An evaluation of the 10 most dominant topics yielded a coherence score of 0.5145, indicating a relevant topic clustering. The IndoBERT demonstrated high agreement with lexicon-generated sentiment labels, with an accuracy of 94.78%, a precision of 95.04%, a recall of 94.79%, and an F1-score of 94.81%. These findings confirm that the combination of transformer-based models is effective for in-depth analysis of discourse in Indonesian-language political news headlines from a major Indonesian online news portal (detik.com).
Bagas Yana Prayoga, Qurrotul Aini, Fitroh Fitroh, "Topic Modeling and Sentiment Analysis on News Headlines Using BERTopic and IndoBERT Models", International Journal of Intelligent Systems and Applications(IJISA), Vol.18, No.4, pp.226-238, 2026. DOI:10.5815/ijisa.2026.04.12
[1]A. D. Riyanto, “Hootsuite (We Are Social): Data Report Digital Indonesia 2024.” Accessed: Apr. 24, 2025. [Online]. Available: https://andi.link/hootsuite-we-are-social-data-digital-indonesia-2024/
[2]A. Yaman, B. Sartono, and A. M. Soleh, “Topic modeling in fertilizer-related patent documents in indonesia based on latent dirichlet allocation,” (in Indonesian), Berk. Ilmu Perpust. dan Inf., vol. 17, no. 2, pp. 168–180, 2021, doi: 10.22146/bip.v17i2.2147.
[3]Z. Alhaq, A. Mustopa, S. Mulyatun, and J. D. Santoso, “Application of support vector machine method for twitter user sentiment analysis,” (in Indonesian), J. Inf. Syst. Manag., vol. 3, no. 1, pp. 16–21, 2021, doi: 10.24076/joism.2021v3i2.558.
[4]R. Egger and J. Yu, “A topic modeling comparison between LDA, NMF, Top2Vec, and BERTopic to demystify twitter posts,” Front. Sociol., vol. 7, no. May, pp. 1–16, 2022, doi: 10.3389/fsoc.2022.886498.
[5]D. Hendry et al., “Topic modeling for customer service chats,” in 2021 International Conference on Advanced Computer Science and Information Systems (ICACSIS), IEEE, 2021, pp. 1–6. doi: 10.1109/ICACSIS53237.2021.9631322.
[6]Y. Liu, “Comparison of LDA and BERTopic in news topic modeling: A case study of the new york times’ reports on china,” Pacific Int. J., vol. 7, no. 3, pp. 47–51, 2024, doi: 10.55014/pij.v7i3.616.
[7]B. Wilie et al., “IndoNLU: Benchmark and resources for evaluating indonesian natural language Understanding,” arXiv Prepr. arXiv2009.05387, 2020, [Online]. Available: http://arxiv.org/abs/2009.05387
[8]Fransiscus and A. S. Girsang, “Sentiment analysis of COVID-19 public activity restriction (PPKM) impact using BERT method,” Int. J. Eng. Trends Technol., vol. 70, no. 12, pp. 281–288, 2022, doi: 10.14445/22315381/IJETT-V70I12P226.
[9]W. Nurfitri and A. Chowanda, “Sentiment analysis on Covid-19 positive cases based on media coverage in indonesia using IndoBERT,” (in Indonesian), Progresif J. Ilm. Komput., vol. 20, no. 1, pp. 580–593, 2024, doi: 10.35889/progresif.v20i1.1897.
[10]N. Mahfudiyah and A. Alamsyah, “Understanding user perception of ride-hailing services sentiment analysis and topic modelling using IndoBERT and BERTopic,” in 2023 International Conference on Digital Business and Technology Management, ICONDBTM 2023, IEEE, 2023, pp. 1–6. doi: 10.1109/ICONDBTM59210.2023.10327320.
[11]E. Effendi, I. Sartika, N. L. T. Br.Purba, and S. Ritonga, “Writing headlines and leads for news and features,” (in Indonesian), J. Pendidik. dan Konseling, vol. 5, no. 2, pp. 4680–4683, 2023, doi: 10.31004/jpdk.v5i2.14206.
[12]E. C. D. A. Nanda, “Apa Konten Berita yang Paling Banyak Diakses Warganet 2024?,” 2024, Goodstats. Accessed: Jan. 16, 2025. [Online]. Available: https://goodstats.id/article/apa-konten-berita-yang-paling-banyak-diakses-warganet-pada-2024-B5hDY
[13]N. Newman, R. Flecther, C. T. Robertson, A. R. Arguedas, and R. K. Nielsen, “Reuters Institute Digital News Report 2024,” 2024. [Online]. Available: https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2024
[14]M. Han, I. Canli, J. Shah, X. Zhang, I. G. Dino, and S. Kalkan, “Perspectives of machine learning and natural language processing on characterizing positive energy districts,” Buildings, vol. 14, no. 2, pp. 1–28, 2024, doi: 10.3390/buildings14020371.
[15]G. S. Mahendra et al., Tren Teknologi AI: Pengantar, Teori, dan Contoh Penerapan Artificial Intelligence di Berbagai Bidang. PT. Sonpedia Publishing Indonesia, 2024.
[16]D. L. C. Pardede and M. A. I. Waskita, “Topic modeling analysis for reviews about peduli lindungi,” (in Indonesian), J. Ilm. Inform. Komput., vol. 28, no. 1, pp. 17–26, 2023, doi: 10.35760/ik.2023.v28i1.7925.
[17]S. Alaparthi and M. Mishra, “Bidirectional encoder representations from transformers (BERT): A sentiment analysis odyssey,” arXiv Prepr. arXiv2007.01127, 2020, [Online]. Available: https://arxiv.org/abs/2007.01127
[18]M. Grootendorst, “BERTopic: Neural topic modeling with a class-based TF-IDF procedure,” arXiv Prepr. arXiv2203.05794, 2022, [Online]. Available: https://arxiv.org/abs/2203.05794
[19]B. Liu, Sentiment Analysis and Opinion Mining. Chicago: Morgan & Claypool Publisher, 2020. doi: 10.1017/9781108639286.
[20]S. B. Panuntun, D. Krismawati, S. Pramana, and E. T. Astuti, “Analysis of telemedicine news coverage in indonesia: Sentiment, NER, topic modeling, and social network approaches to understanding issues and perceptions,” (in Indonesian), Indones. Heal. Inf. Manag. J., vol. 11, no. 1, pp. 56–67, 2023, doi: 10.47007/inohim.v11i1.500.
[21]F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.
[22]A. Simanjuntak et al., “Study and analysis of IndoBERT hyperparameter tuning in fake news detection,” (in Indonesian), J. Nas. Tek. Elektro dan Teknol. Inf., vol. 13, no. 1, pp. 60–67, 2024, doi: 10.22146/jnteti.v13i1.8532.
[23]S. Sazzed and S. Jayarathna, “SSentiA: A self-supervised sentiment analyzer for classification from unlabeled data,” Mach. Learn. with Appl., vol. 4, pp. 1–14, 2021, doi: 10.1016/j.mlwa.2021.100026.
[24]R. Mohammed, J. Rawashdeh, and M. Abdullah, “Machine learning with oversampling and undersampling techniques: overview study and experimental results,” in 2020 11th International Conference on Information and Communication Systems (ICICS), 2020, pp. 243–248. doi: 10.1109/ICICS49469.2020.239556.
[25]Y. Yunitasari and A. R. Putera, “Analysis of public sentiment on twitter regarding the Covid-19 pandemic,” (in Indonesian), Smatika J., vol. 11, no. 01, pp. 22–26, 2021, doi: 10.32664/smatika.v11i01.520.
[26]B. P. E. Wijaya, D. P. Hostiadi, and P. D. W. Ayu, “Sentiment analysis of instagram comments for cyberbullying detection using gradient boosting models,” (in Indonesian) in Seminar Hasil Penelitian Informatika dan Komputer (SPINTER)| Institut Teknologi dan Bisnis STIKOM Bali, 2025, pp. 1093–1098.
[27]D. Aryani, I. Lucia Kharisma, A. Sujjada, and K. Kamdan, “Topic modeling of the 2024 election using the BERTopic method on detik.com news articles,” Inf. J. Ilm. Bid. Teknol. Inf. dan Komun., vol. 9, no. 2, pp. 171–180, 2024, doi: 10.25139/inform.v9i2.8429.
[28]I. N. Nugraha and E. Utami, “Evaluation of creative economy and tourism industry trends based on LDA analysis with BERTopic,” Digit. Zo. J. Teknol. Inf. dan Komun., vol. 15, no. 2, pp. 182–195, 2024, doi: 10.31849/digitalzone.v15i2.23796.
[29]A. Rahmawati, A. Alamsyah, and A. Romadhony, “Hoax news detection analysis using IndoBERT deep learning methodology,” in 2022 10th International Conference on Information and Communication Technology (ICoICT), 2022, pp. 368–373. doi: 10.1109/ICoICT55009.2022.9914902.
[30]P. Sayarizki, Hasmawati, and H. Nurrahmi, “Implementation of IndoBERT for sentiment analysis of indonesian presidential candidates,” J. Comput., vol. 9, no. 2, pp. 61–72, 2024, doi: 10.34818/indojc.2024.9.2.934.