Bagas Yana Prayoga

Work place: Department of Information System, Universitas Islam Negeri Syarif Hidayatullah Tangerang Selatan, 15412, Indonesia

E-mail: bagasyana21@mhs.uinjkt.ac.id

Website:

Research Interests:

Biography

Bagas Yana Prayoga was born in South Jakarta, Indonesia. He is currently pursuing the Bachelor of Science degree in Information Systems at UIN Syarif Hidayatullah Jakarta, Indonesia. His areas of interest include software quality assurance, user interface and user experience (UI/UX) principles, usability testing, and the application of data science and natural language processing (NLP) for text analytics.

Author Articles
Topic Modeling and Sentiment Analysis on News Headlines Using BERTopic and IndoBERT Models

By Bagas Yana Prayoga Qurrotul Aini Fitroh Fitroh

DOI: https://doi.org/10.5815/ijisa.2026.04.12, Pub. Date: 8 Aug. 2026

The high volume of information from online media in Indonesia poses challenges for manual analysis in identifying emerging themes and sentiments. News headlines, as the primary element seen by the public, play a crucial role in shaping opinion; however, their massive volume and diverse themes necessitate automated approaches to identify topics and underlying sentiments. To address this challenge, this study analyzed 30,329 news headlines from the online news portal detik.com for the entire year 2024. A quantitative Natural Language Processing (NLP) framework was applied, comprising data collection via web scraping, text preprocessing, transformer-based topic modeling using BERTopic, sentiment classification using IndoBERT, and a topic-sentiment intersection analysis. Preprocessing included case folding, text cleaning, normalization of informal words, and tokenization. For lexicon-based labeling, stopword removal and stemming were applied. In contrast, transformer-based models utilized text that underwent only case folding, cleaning, and normalization to preserve contextual information. Topic modeling was performed using BERTopic, while sentiment classification (positive, negative, and neutral) used the IndoBERT model. The main objective of this study was to evaluate the combined performance of the two models in mapping dominant issues and the sentiments contained in media reports. The results showed that BERTopic successfully identified 366 topics. An evaluation of the 10 most dominant topics yielded a coherence score of 0.5145, indicating a relevant topic clustering. IndoBERT demonstrated high agreement with lexicon-generated sentiment labels, with an accuracy of 94.78%, a precision of 95.04%, a recall of 94.79%, and an F1-score of 94.81%. These findings confirm that the combination of transformer-based models is effective for in-depth discourse analysis of political news headlines from detik.com, a major Indonesian online news portal.

[...] Read more.
Other Articles