Publication Search

76,969 articles from 728 journals · 2,111 citations tracked

Showing 1-20 of 104

Analytics

Aqiilah, Inge Najwa; Saptono, Ristu; Syaifuddin, Akhmad

Journal of Computing Theories and Applications 2026 Universitas Dian Nuswantoro

Document-level sentiment analysis assigns a single polarity label to an entire review, often obscuring opinion diversity within multi-sentence submissions. This limitation is particularly evident in reviews of multi-service platforms, where users frequently express heterogeneous opinions toward different aspects of the platform in the same review. To address this challenge, this study proposes a sentence-level sentiment analysis framework for Indonesian Gojek app reviews collected from the Google Play Store. The proposed framework introduces a two-stage segmentation strategy that combines punctuation-aware rules with conjunction-aware splitting based on coordinating and adversative conjunctions (e.g., tapi [but], padahal [even though]) to identify opinion boundaries and decompose mixed-sentiment reviews into independently classifiable sentence units. A total of 14,730 raw reviews collected between May and July 2025 were subjected to data cleaning and quality filtering, resulting in 7,187 valid reviews that were further segmented into 14,187 sentence-level instances. Each instance was manually annotated by three annotators using a four-class labeling scheme consisting of app-positive, app-negative, app-neutral, and service categories. Sentiment-level inter-annotator agreement, computed on the subset of instances unanimously categorized as app-related by all three annotators (n = 4,384), achieved substantial agreement (Fleiss'  = 0.636). Hyperparameter optimization was conducted using Optuna with the Tree-structured Parzen Estimator (TPE) sampler across four experimental scenarios. The best performance was achieved by IndoBERTweet under Stratified K-Fold evaluation, attaining an accuracy of 0.751 and a macro F1-score of 0.729, outperforming all IndoBERT configurations. The results demonstrate the effectiveness of domain-adaptive pre-training on informal Indonesian text and highlight the value of conjunction-aware segmentation for preserving fine-grained opinion structures in mixed-sentiment reviews. These findings suggest that domain-aligned language representations provide a practical and effective solution for sentence-level sentiment analysis of Indonesian app reviews.

Vania Vipassana; Mela Karlina; Melati Syaftia; Nindi Juliani; Sakila Salsa Pratiwi +3 more

Bhinneka: Jurnal Bintang Pendidikan dan Bahasa 2026 Universitas Palan

This study aims to map the trajectory of syntactic acquisition in three-year-old children through syntactic patterns and communicative functions in naturalistic interaction. Using a mixed-methods approach, data from native Indonesian-speaking children were collected over a period of 1.5 months through the involve-conversation technique. Analysis of 80 utterances using frequency distribution, Mean Length of Utterance (MLU), and functional grammar revealed a dominant Subject–Verb–Object (S–V–O) structure (30%) and an MLU of 5.82 morphemes. These findings indicate a developmental transition from telegraphic speech to early multi-clause constructions, reflecting increasing linguistic complexity. Cognitive compensation is marked by the use of pragmatic particles and non-canonical sentence patterns driven by ideational, interpersonal, and textual functions. The results support the usage-based hypothesis, suggesting that early syntactic development is functional, sequential, and non-linear in nature. Furthermore, the study highlights the role of interactional experience in shaping emerging grammatical competence. This classification serves as a micro-longitudinal assessment tool and provides a pedagogical basis for scaffolding interventions aimed at stabilizing complex linguistic patterns and enhancing language development in early childhood education settings.

Priyambodo, Aji; Isnanto, R. Rizal; Sanjaya, Ridwan

Journal of Computing Theories and Applications 2026 Universitas Dian Nuswantoro

Batik motif classification has attracted growing attention in visual computing due to its role in cultural heritage preservation, textile informatics, museum documentation, and automated cataloging. Although many studies report high classification accuracy, robustness under real-world acquisition conditions remains insufficiently understood. Batik images are frequently affected by illumination variation, blur, folds, watermark overlays, wearable deformation, scale inconsistency, and background clutter, creating challenges that extend beyond conventional image-noise assumptions. Existing studies largely focus on improving classification performance, while the interactions among acquisition variability, feature representation, evaluation practice, and deployment constraints remain fragmented. This systematic literature review addresses this gap by synthesizing batik classification research through a robustness-aware perspective. Using query expansion, backward and forward citation chaining, relevance screening, and thematic coding, 116 candidate records were identified, resulting in 50 highly relevant studies for detailed analysis. The review reveals that robustness is shaped less by denoising alone than by the combined effects of acquisition conditions, representation design, evaluation realism, and deployment context. Handcrafted descriptors remain competitive for small datasets and structured motifs due to their data efficiency and interpretability, whereas deep learning models achieve the highest reported accuracy when supported by sufficient data diversity and realistic augmentation. Hybrid representations emerge as the most consistently balanced approach, combining local texture stability with higher-level abstraction across heterogeneous acquisition settings. The review further identifies recurring robustness failure patterns, including background dependency, illumination instability, motif-scale inconsistency, wearable deformation, and source-shift vulnerability. Based on these findings, a robustness-oriented research agenda is proposed, emphasizing cross-acquisition evaluation, representation-stability analysis, batik-specific robustness benchmarks, acquisition-aware augmentation, and deployable lightweight or hybrid architectures. The study contributes a domain-specific synthesis that reframes batik motif classification from an accuracy-centric task toward a robustness-aware visual recognition problem.

Rifna, Iza; Nurdin, Nurdin

IT-Explore: Jurnal Penerapan Teknologi Informasi dan Komunikasi 2026 Fakultas Teknologi Informasi, Universitas Kristen Satya Wacana

The Free Nutritional Meal Program (MBG) is a government policy that is widely discussed by the public through social media, especially TikTok. Various comments that have emerged indicate differences in public opinion towards the program, so an analysis is needed to determine the tendency of public sentiment. This study aims to analyze TikTok user sentiment towards the Free Nutritional Meal Program using the Naive Bayes method. The research method is carried out through several steps, namely collecting TikTok comment data, preprocessing text, labeling sentiment data into positive, negative, and neutral, feature transformation using TF-IDF, and classification using the Naive Bayes algorithm. Based on the analysis of 500 comment data, the results show that positive sentiment dominates public opinion by 42% (210 data), followed by negative sentiment by 36% (180 data), and neutral sentiment by 22% (110 data). Testing the classification model using Naive Bayes produces excellent performance with an accuracy rate of 86%, precision of 84%, recall of 85%, and F1-score of 84%. The conclusion of this study shows that the Naive Bayes method is effective as an approach in social media sentiment analysis to map public responses to government policies.

Damayanti, Nadia; Puspasari, Shinta; Suhandi, Nazori

Teknik: Jurnal Ilmu Teknik dan Informatika 2026 LPPM Sekolah Tinggi Ilmu Ekonomi - Studi Ekonomi Modern

Nature tourism is one of the sectors that plays an important role in supporting the development of regional tourism, including in Lahat Regency, which has significant waterfall tourism potential. Currently, many visitors share their reviews and experiences through digital platforms such as Google Maps. This review can be used as a source of information to understand the public's evaluation of the quality of tourist attractions. This study aims to examine public perception of tourist attractions in Lahat Regency using the Support Vector Machine (SVM) method. Research data were collected through scraping from Google Maps, totaling 500 reviews from five tourist attractions, namely Curup Maung, Curup Buluh, Senyawe Waterfall, Panjang Waterfall, and Green Canyon. The research stages include data preprocessing, consisting of cleaning, case folding, normalization, tokenization, stopword removal, and stemming. After that, feature extraction was carried out using the TF-IDF method and the classification process using the SVM algorithm. Based on the research results, the Support Vector Machine (SVM) method is able to perform sentiment classification quite well, although the accuracy level varies for each tourist attraction. Curup Maung and Panjang Waterfall achieved the highest accuracy level of 90%. Nevertheless, most visitor reviews were dominated by negative sentiments. This indicates that there are still several aspects that need to be improved, particularly related to tourist facilities and services. This research is expected to serve as a consideration for tourism managers and local governments in efforts to improve management quality as well as the development of tourism in Lahat Regency.

Veri Arinal; Satria Wira Yudha; Muhammad Joko Umbaran Kharis Bahrudin; Dessyanti Ryantina

International Journal of Information Engineering and Science 2026 Asosiasi Riset Teknik Elektro dan Infomatika Indonesia

QRIS (Quick Response Code Indonesian Standard) has become a widely used national digital payment standard. User satisfaction with this service needs to be monitored continuously to ensure its sustainability. This study aims to predict the level of QRIS user satisfaction based on their experiences and perceptions expressed organically on the Twitter social media platform. The method used is sentiment analysis with the Naive Bayes classification algorithm implemented using RapidMiner software. The research data was obtained from Twitter user comments collected through web scraping techniques. The text data then went through a preprocessing stage that included cleansing, stopword filtering, stemming, and tokenizing to be prepared as features ready to be processed by the model. The data was divided into training (80%) and testing (20%) subsets for model training and validation. The results showed that the Naive Bayes model was able to predict user satisfaction sentiment with an accuracy of 80.99%. These findings indicate that the model is highly accurate in identifying satisfied comments and sufficiently sensitive in detecting dissatisfaction. This study concludes that sentiment analysis of Twitter UGC data using Naive Bayes is an effective and efficient approach for predicting QRIS user satisfaction in real time. The practical implication of this study is to provide an automatic feedback system for service providers to monitor public sentiment and take targeted corrective actions.

Mesra Betty Yel; Sopan Adrianto; Rasiban Rasiban; Eva Widiyanti

International Journal of Information Engineering and Science 2026 Asosiasi Riset Teknik Elektro dan Infomatika Indonesia

The growth of information technology has driven changes in consumer behavior, one of which is through e-commerce platforms such as Shopee. This phenomenon has generated a large number of customer reviews, including those for local cosmetic products such as Wardah. These reviews serve as an important source of information for understanding customer perceptions and satisfaction levels. However, manual analysis of large and linguistically diverse datasets is inefficient and potentially subjective. This study aims to implement the multi-category Naive Bayes algorithm to classify the sentiment of Wardah product reviews on Shopee into three categories: positive, negative, and neutral. The data were collected using a web scraping technique and processed through a series of preprocessing stages including case folding, tokenization, stopword removal, stemming, and text cleaning. Subsequently, term weighting was performed using the TF-IDF method prior to classification. Model performance was evaluated using a confusion matrix as well as accuracy, precision, and recall metrics. The results indicate that the multi-category Naive Bayes algorithm achieved an accuracy of 86.00%, a precision of 86.63%, and a recall of 98.24%. This approach can assist business practitioners in objectively understanding customer opinions and support decision-making in business strategy and product development.

Yuma Akbar; Frencis Matheos Sarimolle; Dwi Swasono Rachmad; Muhammad Derry Oktaviandi

International Journal of Applied Mathematics and Computing 2026 Asosiasi Riset Ilmu Matematika dan Sains Indonesia

This study aims to analyze public sentiment toward the hashtag #KaburAjaDulu, which has circulated widely on the social media platform X (formerly Twitter). The hashtag reflects the growing anxiety among the public, especially younger generations, regarding socio-political issues in Indonesia. The data were collected using web scraping techniques, focusing on user-generated tweets that contain the hashtag. A comprehensive text preprocessing phase was conducted to clean the raw data by removing irrelevant elements such as URLs, emojis, numbers, and punctuation. The research applies a hybrid classification approach using a combination of Support Vector Machine (SVM) and Random Forest algorithms to categorize sentiment into three classes: positive, negative, and neutral. The performance of the model was evaluated using metrics such as accuracy, precision, recall, and F1-score to determine the effectiveness of the classification. The study aims to demonstrate that combining algorithms can improve classification performance compared to using a single algorithm. This research contributes to the field of sentiment analysis and provides valuable insights for researchers, policymakers, and social observers in understanding public opinion trends in digital media.

Jamila Tun Nabilah Hasanuddin; Marwiah Marwiah; Aco K

Bhinneka: Jurnal Bintang Pendidikan dan Bahasa 2026 Universitas Palan

This research aims to analyze how gender and culture are represented in the Grade VIII Indonesian language textbook of the Merdeka Curriculum. The focus is on various aspects of gender representation, such as the depiction of characters, their roles, activities, attributes, social status, gender equality, and stereotypes related to gender. Additionally, the study explores cultural representation, which encompasses cultural forms, diversity, local traditions, context, ways of presentation, and the cultural values expressed in the textbook's texts and illustrations. The methodology employed is descriptive qualitative research with a content analysis framework. Data were gathered through documenting and note-taking methods on the content of the textbook, followed by an analysis process that includes identification, classification, interpretation, and drawing conclusions. Findings indicate that the gender representation in the textbook predominantly portrays men as leading figures in public roles, leadership, and decision-making, whereas women are mainly shown in nurturing, domestic, and supportive capacities. On the other hand, the cultural representation illustrates the variety of Indonesian culture by showcasing regional customs, languages, art forms, traditional cuisine, practices, and societal norms. The study concludes that although the textbook presents cultural diversity adequately, there is a need for improvement in gender representation balance to better reflect equality values in the educational experience.

Untung Surapati; Veri Arinal; Tri Wahyudi; Ahmad Fauzan

International Journal of Applied Mathematics and Computing 2026 Asosiasi Riset Ilmu Matematika dan Sains Indonesia

The rise of social media has created a digital public sphere that enables users to express their opinions on social and political issues openly and in real-time. One of the most discussed topics on social media platform X is the trending hashtag #IndonesiaGelap, which reflects public concern and criticism regarding various governmental and societal conditions. This study aims to conduct sentiment analysis on tweets containing the hashtag to determine the overall sentiment trend among users. The method employed in this research is the Naive Bayes classification algorithm, known for its simplicity and effectiveness in text classification. To enhance the model’s performance, Particle Swarm Optimization (PSO) is applied to optimize feature selection and parameter tuning. The dataset consists of public tweets collected via the Twitter API, followed by preprocessing, feature extraction using TF-IDF, and sentiment classification into three categories: positive, negative, and neutral. The results indicate that the integration of PSO significantly improves the classification accuracy of the Naive Bayes model compared to the baseline. The majority of tweets related to #IndonesiaGelap exhibit a negative sentiment, indicating widespread public dissatisfaction and criticism. This research is expected to contribute to a better understanding of public perception and serve as valuable input for stakeholders in addressing social issues in the digital age.

Mukhlisin Nata Hudin; Radit Septa Wijaya; Muhammad Daffa Pratama; Hudaidah Hudaidah; Risa Marta Yati

Jurnal Pendidikan Dirgantara 2026 Asosiasi Riset Ilmu Pendidikan Indonesia

This research is based on the importance of studying Malay-Jawi religious manuscripts as a source of transmission of Islamic teachings in the archipelago, particularly in the field of monotheism. The study aims to examine the textual content of Jawi manuscripts containing the treatise of monotheism, especially the concept of the sentence of monotheism and the attributes of twenty, and to explain their position in the intellectual tradition of Malay Islam. The research employs This research is based on the importance of studying Malay-Jawi religious manuscripts as a source of transmission of Islamic teachings in the archipelago, particularly in the field of monotheism. The study aims to examine the textual content of Jawi manuscripts containing the treatise of monotheism, especially the concept of the sentence of monotheism and the attributes of twenty, and to explain their position in the intellectual tradition of Malay Islam. The research employs a qualitative method with a philological approach and content analysis. Primary data consist of Jawi manuscripts, while secondary data are obtained through library research. Data were collected through documentation and literature review and analyzed descriptively. The findings reveal that the manuscripts contain systematically arranged monotheistic teachings, including the meaning of lā ilāha illa Allāh through the principles of negation and affirmation, as well as the concept of faith involving the heart, speech, and actions. The manuscripts also explain the twenty attributes within the classifications of nafsiyah, salbiyah, ma‘ani, and ma‘nawiyah, reflecting the theological framework of Ahlussunnah wal Jama‘ah. These manuscripts function as both religious texts and pedagogical media, highlighting the importance of preserving Nusantara Islamic manuscripts as part of the region’s intellectual heritage.

Aura Rahayu Aksa Radiana; Fathoni Mahardika; Dani Indra Junaedi

Merkurius : Jurnal Riset Sistem Informasi dan Teknik Informatika 2026 Asosiasi Riset Teknik Elektro dan Informatika Indonesia

This study aims to develop a sentiment classification method for YouTube user comments related to the game Love and Deepspace using the Naïve Bayes algorithm, focusing on improving the text data processing and understanding user perceptions. Comment data were collected through scraping from YouTube videos, followed by preprocessing including text cleaning, normalization, stopword removal, stemming, and translation into English. Initial labeling was conducted using TextBlob, then the data were randomly sampled for training the Naïve Bayes model. Evaluation involved comparing sentiment distributions and visualization using Word Cloud and bar charts. The Naïve Bayes model achieved an accuracy of 77.36% in sentiment classification. The sentiment distribution shows differences between TextBlob (positive: 1,011, neutral: 1,312, negative: 575) and Naïve Bayes (positive: 901, neutral: 1,627, negative: 370), with Naïve Bayes being more conservative. The Word Cloud visualization identifies dominant words such as "bang," "game," and "main," while the bar chart shows the largest proportion of neutral sentiment. Naïve Bayes is effective for sentiment classification on informal comment data, with significant differences from rule-based methods like TextBlob. This research contributes to the development of text data processing techniques and user perception analysis, as well as opening up optimization opportunities with other algorithms like SVM for better accuracy.

Ayu Astuti Siregar; Al-Khowarizmi

Merkurius : Jurnal Riset Sistem Informasi dan Teknik Informatika 2026 Asosiasi Riset Teknik Elektro dan Informatika Indonesia

Social media has evolved into a significant platform where consumers freely express their opinions, experiences, and levels of satisfaction regarding various products, including those offered by Micro, Small, and Medium Enterprises (MSMEs). The comments and reviews shared by customers on these platforms contain diverse sentiments that can serve as valuable indicators of how consumers perceive product quality. Understanding these sentiments is crucial for MSME owners, as it allows them to evaluate their products and adapt to market expectations more effectively. This study aims to analyze customer sentiment toward MSME products on social media by utilizing the Naïve Bayes algorithm, a widely used classification method in text mining. The data used in this research consist of customer comments collected from various social media platforms. The research process involves several stages, including data collection, manual labeling of sentiments, text preprocessing (such as tokenization, case folding, and stopword removal), and splitting the dataset into training and testing subsets. Subsequently, the classification process is carried out using the Naïve Bayes algorithm to categorize sentiments into positive, negative, and neutral classes. The results of this study demonstrate that the Naïve Bayes method is effective in classifying customer sentiments with a satisfactory level of accuracy. These findings provide a comprehensive overview of consumer perceptions regarding the quality of MSME products. Furthermore, this research is expected to assist MSME business owners in understanding customer feedback more systematically and using it as a basis for improving product quality and enhancing customer satisfaction in a competitive digital marketplace.

Winarno, Edy; Nur, Indah Manfaati; Karim, Abdul; Amri, Saeful; Wirdati, Ismi Elya +1 more

Journal of Computing Theories and Applications 2026 Universitas Dian Nuswantoro

Artificial intelligence has the potential to support radiology workflows by assisting in the identification of cases that may require additional clinical attention. However, alert-oriented medical AI systems should provide not only classification outputs but also interpretable evidence that can be reviewed and audited by clinicians. This study develops and evaluates an explainable multimodal framework for binary chest X-ray alert classification using paired radiology reports and chest X-ray images. The text branch employs TF-IDF n-gram features with a class-balanced Logistic Regression classifier, while the image branch fine-tunes a pretrained ResNet18 model. The two branches are integrated through probability-level late fusion using a validation-selected fusion weight. Explainability is implemented in a modality-specific manner: global coefficient analysis is used to identify influential textual cues, while Grad-CAM heatmaps are used to visualize salient image regions. Experiments were conducted on paired samples from the Open-i/IU X-Ray dataset using text-only, image-only, and fusion-based evaluation settings. Additional analyses include case-level complementarity analysis, bootstrap confidence intervals for ROC-AUC, shortcut-feature inspection, and qualitative Grad-CAM auditing. The results indicate that the text modality provides the dominant predictive signal under the current proxy-label setting. Late fusion produced a small descriptive improvement on the test set, increasing accuracy from 0.8533 to 0.8667, F1-score from 0.8817 to 0.8936, and ROC-AUC from 0.8936 to 0.9025 compared with the text-only baseline. However, the observed ROC-AUC improvement was not statistically conclusive based on bootstrap analysis. These findings suggest that the proposed framework is useful as a reproducible and auditable multimodal prototype, while also highlighting important limitations, including proxy-label ambiguity, potential label leakage from radiology reports, limited image-branch contribution, lack of external validation, and the need for stronger explanation and calibration assessment.

Sulaeni, Dini; Purnamasari, Ade Irma; Ali, Irfan; Kurniawan, Rudi; Nurdiawan, Odi +5 more

JUISI : Jurnal Ilmiah Sistem Informasi 2026 LPPM Universitas Sains dan Teknologi Komputer

The increasing use of mobile applications in the retail industry has generated a large volume of user reviews that contain valuable insights regarding customer experience and service quality. However, the unstructured nature of these reviews requires an automated approach to extract meaningful patterns efficiently. This study aims to perform sentiment analysis on user reviews of the Indomaret Poinku application by integrating lexicon-based labeling with machine learning classification. A total of 10,000 reviews were collected from Google Play Store and processed through a series of text preprocessing steps, including cleaning, case folding, normalization, tokenization, stopword removal, and stemming. Sentiment labeling was performed using the Indonesian Sentiment Lexicon (InSet), producing three sentiment classes: positive, negative, and neutral. The labeled data were vectorized using CountVectorizer and classified using two algorithms: K-Nearest Neighbors (KNN) and Random Forest (RF). Evaluation results show that Random Forest outperforms KNN, achieving an accuracy of 82.5%, compared to 69% for KNN. Random Forest demonstrates superior performance in handling high-dimensional sparse text features and yields more stable predictions across sentiment classes. This study contributes to the growing body of research on Indonesian sentiment analysis by demonstrating the effectiveness of combining lexicon-based labeling with ensemble learning methods, offering practical implications for developers seeking to improve the quality and user satisfaction of digital retail applications.

Trianto, Nafil Rizq; Wijaya, Alfarizi; Pardede, Arion; Pandiangan, Daniel; Syahputra, Hermawan

Teknik: Jurnal Ilmu Teknik dan Informatika 2026 LPPM Sekolah Tinggi Ilmu Ekonomi - Studi Ekonomi Modern

Communication is an essential human right, yet a significant communication gap persists between individuals with sensory disabilities, specifically the deaf and speech-impaired, and the general public. While many technological solutions have been proposed to translate sign language, existing models primarily rely on heavy deep learning architectures such as Convolutional Neural Networks (CNN) or Recurrent Neural Networks (RNN/LSTM). These models often demand high computational power, leading to latency and limiting real-time application on standard devices. This study proposes a lightweight, fast, and highly responsive sign language translation system specifically designed to recognize static alphabets (A-Z) and single-character air writing. The system utilizes MediaPipe for hand tracking, where feature extraction is intelligently processed by calculating the relative spatial coordinates of fingertips to the wrist, reducing dependency on raw camera coordinates. Classification is performed using a Support Vector Machine (SVM) with a Radial Basis Function (RBF) kernel, prioritizing computational efficiency without sacrificing accuracy. To enhance user experience, the system introduces three key novelties: smart relative feature extraction, an anti-duplication hold system with a 1-second timer to prevent input spamming, and a non-blocking multithreaded audio execution (Daemon Thread) utilizing Google Text-to-Speech (gTTS), ensuring the webcam feed remains fluid during audio playback. Additionally, an alternative air-writing mode is integrated, utilizing geometric heuristics and PyTesseract OCR to read single drawn letters in the air. The results indicate that the proposed system operates swiftly and efficiently, bridging the communication barrier with a hardware-friendly approach.

Budianoor, Rahmat; Saputro, Setyo Wahyu; Abadi, Friska; Nugroho, Radityo Adi; Farmadi, Andi

Journal of Computing Theories and Applications 2026 Universitas Dian Nuswantoro

Indonesian culinary comments on social media platforms such as Instagram are characterized by informal spelling, regional language mixing, slang expressions, and emojis, posing substantial challenges for automated sentiment classification. While IndoBERT has demonstrated strong performance across Indonesian natural language processing tasks, the contribution of individual preprocessing components to fine-tuning performance on informal text remains underexplored, particularly in the culinary domain. This study addresses this gap by conducting a systematic preprocessing ablation study on IndoBERT-Base fine-tuning for Indonesian culinary sentiment classification, accompanied by a comparative evaluation against Naive Bayes with TF-IDF, SVM with TF-IDF, and BiLSTM as representative baselines. A dataset of 3,500 manually labeled Instagram culinary comments across three sentiment classes was used, with a stratified 80/10/10 split. Six preprocessing variants were evaluated under identical experimental conditions to isolate the contribution of each component. The results show that slang normalization is the most impactful single preprocessing step, yielding a macro F1-score gain of +0.0609 over the no-preprocessing baseline, while the full pipeline achieves an accuracy of 0.8800 and a macro F1-score of 0.8465. IndoBERT-Base with the full pipeline outperforms all baselines across all evaluation metrics. Per-class analysis reveals that the negative class achieves the lowest F1-score of 0.7600, with sarcastic expressions and Banjar regional vocabulary identified as primary sources of misclassification. These findings indicate that preprocessing decisions have a measurable and non-uniform effect on IndoBERT fine-tuning performance. In this study, slang normalization provides the most substantial individual contribution in bridging the vocabulary gap between informal user-generated text and the model’s pre-training distribution.

Nabeel Fazle Mawla Buntaran; Safrizal Safrizal

JURNAL PENELITIAN SISTEM INFORMASI 2026 Institut Teknologi dan Bisnis (ITB) Semarang

This study aims to examine user opinion tendencies toward Gojek services by integrating Random Forest and K-Means Clustering approaches. The dataset consists of 15,000 user reviews collected throughout 2025 using web scraping techniques. The initial stage focuses on data preprocessing, including text cleaning, case normalization, tokenization, removal of non-informative stop words, and lemmatization to restore words to their base forms. Subsequently, sentiment labels are assigned using a lexicon-based approach. The next phase involves classification modeling through Random Forest to identify sentiment tendencies, while K-Means Clustering is employed to uncover latent patterns within the opinion data. The findings indicate that the Random Forest model achieves an accuracy level of 0.878, demonstrating strong performance in distinguishing positive and negative sentiments, as reflected by f1-scores of 0.932 and 0.818, respectively. However, the model shows limitations in consistently identifying neutral sentiment. In contrast, the implementation of K-Means Clustering successfully categorizes the data into three primary clusters, providing a more structured representation of user opinion characteristics. Overall, these results offer empirical insights that can serve as a strategic reference for enhancing the quality of Gojek’s service delivery.

Fifi Amelia Sitinjak; Putrizal Nada Yasmin; Rahmi Anggita Lubis; Zahira Salsabila

Jurnal Riset Rumpun Ilmu Bahasa 2026 Pusat riset dan Inovasi Nasional

This research is driven by the significance of examining language meaning, particularly connotative meaning which is often used to convey implicit messages in literary works. Folktales, as one type of oral literature, often utilize character names that carry specific meanings, such as Malin Kundang from Minangkabau. The aim of this research is to uncover the connotative meaning of the character name “Malin Kundang”, as well as its relationship with the moral and cultural values of the society. The methodology involves qualitative research using semantic analysis and a descriptive approach. Data were collected from the Malin Kundang folktale text thru library study and documentation techniques. Next, the processes of identification, classification, and interpretation of meaning were done in order to examine the data. The research result show that denotatively, the name Malin Kundang only functions as the identity of the character, but connotatively, it functions as a symbol of a disobedient child who does not respect their parents. The meaning is formed from the plot of the story and reinforced by the Minangkabau cultural values that uphold respect for parents. The implications of this research indicate that the naming of characters in folklore is not only linguistic but also reflects the moral and cultural values inherited by society.

Ilham Saputra; Anita Qoiriah

Merkurius : Jurnal Riset Sistem Informasi dan Teknik Informatika 2026 Asosiasi Riset Teknik Elektro dan Informatika Indonesia

The proliferation of online gambling promotional comments on Indonesian social media has become a serious issue requiring fast and accurate automated handling. This study aims to implement a Hybrid Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) method to classify online gambling comments and compare its performance with standalone RNN and LSTM models. The research utilized a dataset of 10,230 comments subjected to comprehensive preprocessing stages, including the normalization of non-standard language using a slang dictionary. Testing was conducted across three data-splitting scenarios: 90:10, 80:20, and 70:30. Experimental results demonstrate that the standalone LSTM model achieved the highest average accuracy of 97.45%. However, the Hybrid RNN–LSTM model showed significant superiority in terms of performance stability, yielding the lowest standard deviation (0.0027) and the smallest Coefficient of Variation (0.28%) across all scenarios. These findings indicate that while the LSTM architecture is highly effective at capturing short-text context, the Hybrid approach provides better robustness against fluctuations in data proportions, making it highly relevant for implementation as an automated detection system on social media.