Bagas Yana Prayoga

Work place: Department of Information System, Universitas Islam Negeri Syarif Hidayatullah Tangerang Selatan, 15412, Indonesia

E-mail: bagasyana21@mhs.uinjkt.ac.id

Website:

Research Interests:

Biography

Bagas Yana Prayoga was born in South Jakarta, Indonesia. He is currently pursuing the Bachelor of Science degree in Information Systems at UIN Syarif Hidayatullah Jakarta, Indonesia. His areas of interest include software quality assurance, user interface and user experience (UI/UX) principles, usability testing, and the application of data science and natural language processing (NLP) for text analytics.

Author Articles
Topic Modeling and Sentiment Analysis on News Headlines Using BERTopic and IndoBERT Models

By Bagas Yana Prayoga Qurrotul Aini Fitroh Fitroh

DOI: https://doi.org/10.5815/ijisa.2026.04.12, Pub. Date: 8 Aug. 2026

The high flow of information from online media in Indonesia makes it difficult for manual analysis to identify emerging themes and sentiments. News headlines, as the first element seen by the public, play a crucial role in shaping opinion, but their massive volume and diverse themes make it difficult for manual analysis to identify topics and their underlying sentiments. To address this challenge, this study analyzed 30,329 news headlines from the online news portal detik.com for the entire year 2024. A quantitative Natural Language Processing (NLP) framework was applied, consisting of data collection through web scraping, text preprocessing, transformer-based topic modeling using BERTopic, sentiment classification using IndoBERT, and a topic sentiment intersection analysis. Preprocessing included case folding, text cleaning, normalization of informal words, and tokenization. For lexicon-based labeling, stopword removal and stemming were applied, while transformer-based models utilized minimally processed text to preserve contextual information. Topic modeling was performed using BERTopic, while sentiment classification (positive, negative, and neutral) used the IndoBERT model. The main objective of this study was to evaluate the combined performance of the two models in mapping dominant issues and the sentiments contained in media reports. The results showed that BERTopic successfully identified 366 topics. An evaluation of the 10 most dominant topics yielded a coherence score of 0.5145, indicating a relevant topic clustering. The IndoBERT demonstrated high agreement with lexicon-generated sentiment labels, with an accuracy of 94.78%, a precision of 95.04%, a recall of 94.79%, and an F1-score of 94.81%. These findings confirm that the combination of transformer-based models is effective for in-depth analysis of discourse in Indonesian-language political news headlines from a major Indonesian online news portal (detik.com).

[...] Read more.
Other Articles