Publication Search

79,575 articles from 739 journals · 2,111 citations tracked

Showing 21-40 of 94

Analytics

Nanda Mediya Sari; Jasmir Jasmir; Elvi Yanti

Prosiding Seminar Nasional Ilmu Teknik 2025 Asosiasi Riset Ilmu Teknik Indonesia

Sentiment analysis is a technique in Natural Language Processing (NLP) used to identify user opinion tendencies based on textual reviews. This study analyzer user reviews of the Maxim application on the Google Play Store and compares three Machine Learning algoritmhs-Naïve Bayes, Support Vector Machine (SVM), and CatBoost-in classifying sentiment. The research stages include data collection, text preprocessing, feature extraction using TF-IDF and Chi-Square, class balancing using SMOTE, and performance evaluation through Accuracy, Precision, Recall, and F1-Score. ANOVA is used to examine the influence of feature selection on model performance. The results show that each model exhibits different performance level across the tested feature combinations. The CatBoost achieved the highest accuracy of 99,26% and demonstrating the most stable performance. Meanwhile, the Naïve Bayes and SVM models experienced performance decreases experiments, especially after applying SMOTE. These findings indicate that the choise of algorithm, feature extraction method, and class balancing technique significantly affects classification outcomes. Overall, CatBoost is identified as the best-performing model, providing more consistenst classification result in accordance with the characteristics of the user reviews.

Ryzal Nur Alvandy; Ryzal Nur Alvandy; Arita Witianti

Jurnal Elektronika dan Komputer 2025 STEKOM PRESS

The rapid expansion of e-commerce in Indonesia has resulted in a significant rise in the number of customer reviews, which serve as a valuable source of insight for understanding consumer satisfaction. This study aims to classify or identify sentiments from product reviews on the Tokopedia platform into three categories, using the Support Vector Machine algorithm. The classification method data were ethically collected through web scraping and include review text, ratings, and the number of “likes.”  The preprocessing stage involved several NLP techniques such as pre-procesesing data representation was generated using the Term Frequency–Inverse Document Frequency method, while the issue of class imbalance was addressed using the Synthetic Minority Over-sampling Technique.  Based on the test results, the SVM model achieved an accuracy of 79.48% on the test data using a linear kernel, showing the best performance in classifying positive sentiments. However, the classification of neutral and negative sentiments still requires improvement. This study demonstrates that the combination of the TF-IDF method, additional numerical features, and data balancing techniques can produce an an efficient sentiment analysis model within the e-commerce domain.

Annisa Fathia Aziza; Hayati Noor; Rina Alfah

Jurnal Sistem Informasi dan Ilmu Komputer 2025 International Forum of Researchers and Lecturers

The world of work is an environment related to the work we are currently in. In other words, it is a place where various individuals perform an activity. The quality of college graduates is not only seen in terms of high or good grades / GPA. There are many other considerations, where large companies see a potential possessed by the person concerned. The dataset in this study was taken from student respondents about the world of work. One way to classify the influence of competence on the world of work in machine learning is to use datasets as training data so that performance testing can be carried out with the right classification method. From the results of the tests carried out, it is concluded that the results of the comparison are different, which shows the accuracy value of KNN which is around 96%, while the results of the SVM accuracy tested are 98%, so that the accuracy of SVM is better than KNN.

Hamza, Ali; Hussain, Wahid; Iftikhar, Hassan; Ahmad, Aziz; Shamim, Alamgir Md

Journal of Computing Theories and Applications 2025 Universitas Dian Nuswantoro

The rapid growth of open-source software (OSS) in machine learning (ML) has intensified the need for reliable, automated methods to assess project quality, particularly as OSS increasingly underpins critical applications in science, industry, and public infrastructure. This study evaluates the effectiveness of a diverse set of machine learning and deep learning (ML/DL) algorithms for classifying GitHub OSS ML projects as engineered or non-engineered using a SMOTE-enhanced and explainable modeling pipeline. The dataset used in this research includes both numerical and categorical attributes representing documentation, testing, architecture, community engagement, popularity, and repository activity. After handling missing values, standardizing numerical features, encoding categorical variables, and addressing the inherent class imbalance using the Synthetic Minority Oversampling Technique (SMOTE), seven different classifiers—K-Nearest Neighbors (KNN), Decision Tree (DT), Random Forest (RF), XGBoost (XGB), Logistic Regression (LR), Support Vector Machine (SVM), and a Deep Neural Network (DNN)—were trained and evaluated. Results show that LR (84%) and DNN (85%) outperform all other models, indicating that both linear and moderately deep non-linear architectures can effectively capture key quality indicators in OSS ML projects. Additional explainability analysis using SHAP reveals consistent feature importance across models, with documentation quality, unit testing practices, architectural clarity, and repository dynamics emerging as the strongest predictors. These findings demonstrate that automated, explainable ML/DL-based quality assessment is both feasible and effective, offering a practical pathway for improving OSS sustainability, guiding contributor decisions, and enhancing trust in ML-based systems that depend on open-source components.

Rahmeisi, Nazli; Gani, Eksa Umar; Arfriandi, Arief; Rahmeisi, Nazli; Gani, Eksa +1 more

JUISI : Jurnal Ilmiah Sistem Informasi 2025 LPPM Universitas Sains dan Teknologi Komputer

The rapid growth of web technologies and online services has increased the exposure of web applications to cyber threats such as Cross-Site Scripting (XSS) and SQL Injection (SQLi). Conventional rule-based mechanisms, such as Web Application Firewalls (WAFs), often fail to detect emerging attack patterns. To address this, Machine Learning (ML) and Deep Learning (DL) have emerged as adaptive approaches for enhancing web attack detection. This study performs a Systematic Literature Review (SLR) following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines to analyze recent ML/DL-based detection methods. Of the 263 retrieved studies, 15 met the inclusion criteria for detailed review. The findings reveal that Random Forest (RF), Support Vector Machine (SVM), Convolutional Neural Network (CNN), and Long Short-Term Memory (LSTM) are the most applied algorithms. At the same time, recent works emphasize Transformer-based and hybrid ML–DL models. These approaches achieved robust performance (accuracy 85–97%, F1-score >90%) but still face challenges in dataset representativeness, class imbalance, and computational cost. This review highlights future research directions in Explainable Artificial Intelligence (XAI), Federated Learning (FL), and adversarial robustness to develop more efficient and trustworthy web attack detection systems.

Sipasulta, Angelica Mailen; Bayu, Teguh Indra

IT-Explore: Jurnal Penerapan Teknologi Informasi dan Komunikasi 2025 Fakultas Teknologi Informasi, Universitas Kristen Satya Wacana

Bea Cukai has recently been in the public spotlight, especially regarding the supervision of goods from abroad. News and public responses regarding Bea Cukai's supervision create pros and cons, thus triggering a variety of responses from the public. This study aims to analyze the sentiment of Indonesian people towards the performance of Bea Cukai in monitoring goods from abroad by utilizing Twitter social media. In this research, the Support Vector Machine (SVM) algorithm is applied to classify public comments on Twitter into positive or negative sentiments. Through the crawling process carried out from June 1, 2023, to May 12, 2024, 9,051 entries of data were collected. The analysis results showed an accuracy of 93.87%, precision 94%, recall 93%, and F1-score 94%. These results show that the SVM method is effective in analyzing public sentiment, especially related to Bea Cukai's supervision.

Gunawan, Ricardho; Hendry, Hendry

IT-Explore: Jurnal Penerapan Teknologi Informasi dan Komunikasi 2025 Fakultas Teknologi Informasi, Universitas Kristen Satya Wacana

Sentiment analysis of guest reviews is a crucial aspect in improving the quality of hotel services. This study aims to analyze the sentiment of guest reviews regarding the services of Grand Diamond Hotel Yogyakarta using a machine learning approach with the Support Vector Machine (SVM) algorithm. SVM was chosen because it can handle high-dimensional data such as text and is capable of forming an optimal separating hyperplane between sentiment classes. The research data was obtained through web scraping from Traveloka, yielding 1,119 reviews, which were processed through preprocessing, translation, and sentiment labeling using the TextBlob library. After TF-IDF weighting, the data was divided into 80% for training and 20% for testing. The linear kernel SVM model achieved 80% accuracy in classifying the reviews into positive, negative, and neutral categories. The results of this study were implemented in a web-based application equipped with data visualization and model evaluation features, allowing hotel management to efficiently monitor and analyze guest sentiment and support data-driven service quality improvement.

Bambang Irwansyah; Novica Jolyarni Dornik; Riswan Syahputra Damanik

Sevaka : Hasil Kegiatan Layanan Masyarakat 2025 STIKES Columbia Asia Medan

Hair loss is one of the common health problems experienced by many people and often causes psychological impacts, particularly on self-confidence. The factors contributing to hair loss are diverse, ranging from genetics, diet, and stress to lifestyle. The lack of public knowledge about these risk factors, as well as the low level of digital literacy in the use of predictive technology, makes it difficult for people to take early preventive measures. This community service activity aims to provide education and simple training on predicting hair loss risk using the Support Vector Machine (SVM) algorithm for residents of Rantau Prapat Village. The implementation methods include a pre-test to measure initial understanding, interactive counseling on hair loss risk factors, practical simulation of risk prediction using SVM based on a simple dataset, and evaluation through a post-test. The results of the activity showed a significant increase in participants’ understanding, from an average of 45.2% in the pre-test to 81.6% in the post-test, with a participant satisfaction level reaching 92%. This counseling not only improved health literacy but also introduced the practical application of artificial intelligence in the health sector.

Wahyu Saputro

Mars: Jurnal Teknik Mesin, Industri, Elektro Dan Ilmu Komputer 2025 Asosiasi Riset Teknik Elektro dan Informatika Indonesia

Human Resource Management (HRM) plays a strategic role in improving organizational competitiveness through proper management of employee placement, training, and performance evaluation. To support the achievement of these goals, a predictive model is needed that can provide an accurate picture of employee performance. This study utilizes a Human Resource Management (HRM) dataset of 1,200 data and applies several classification algorithms to compare their effectiveness, namely J48 or C4.5, Random Forest, Naive Bayes, K-Nearest Neighbor (KNN), Logistic Regression, and Support Vector Machine (SVM). To obtain more optimal results, this study uses resampling techniques and attribute selection methods with a correlation attribute eval approach, so that class distribution can be more balanced and model accuracy increases. From the test results, the Decision Tree J48 algorithm showed the best performance with an accuracy level reaching 95.41%, a kappa value of 0.8925, a mean absolute error (MAE) of 0.0432, a precision of 0.955, a recall of 0.954, and an area under the ROC curve of 0.964. These findings indicate that J48 has excellent predictive capabilities compared to other algorithms. Furthermore, this study also found that the most influential variables in determining employee performance include the percentage of the last salary increase (EmpLast Salary Hike Percent), the level of work environment satisfaction (Emp Environment Satisfaction), the length of time since the last promotion (Years Since Last Promotion), and experience in the current role (Experience Years in Current Role). Overall, the results of the study indicate that the C4.5 algorithm with the application of the resampling technique can be an optimal solution in building an employee performance prediction system. Thus, this model has the potential to be a strong basis for managerial decision-making, particularly in designing HR development strategies and policies to improve organizational performance.

Prashanthan, Amirthanathan

Journal of Computing Theories and Applications 2025 Universitas Dian Nuswantoro

The study presents a comprehensive framework for optimizing customer retention budget by integrating clustering, classification, and mathematical optimization techniques. The study begins with the IBM Telco dataset, which is prepared through data cleansing, encoding, and scaling.  In the preliminary phase, customer segmentation is performed using K-Means clustering, with k = 3 and k = 4 identified as optimal based on the elbow method and Silhouette score. The configurations produced three (Premium, Standard, Low) and four (Premium, Standard Plus, Standard, Low) customer segments based on purchase preferences, which served as input features for churn prediction. In the second phase, the dataset was divided into training and test sets in an 80:20 ratio, followed by data balancing using the Synthetic Minority Over-sampling Technique (SMOTE) and Edited Nearest Neighbors (ENN). Multiple classification algorithms were evaluated, including Naive Bayes (NB), Random Forest (RF), Categorical Boosting (CatBoost), Light Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), Gradient Boosting (GB), Support Vector Machine (SVM), Logistic Regression (LR), K-Nearest Neighbors (KNN), and Multi-Layer Perceptron (MLP) using F1-score as the performance metric. CatBoost and LightGBM, with k values of 3 and 4, respectively, were the highest-performing classification models, with only minimal differences in performance.    Ultimately, customer segmentation established customer prioritization, whereas churn prediction assessed customer churn likelihood. Four distinct configurations were assessed utilizing mixed-integer linear programming (MILP) to optimise retention budget allocation within uniform budget constraints, discount amounts, and churn thresholds. In both the k=3 and k=4 scenarios, CatBoost surpassed LightGBM, with CatBoost at K=3 effectively discounting 66% of at-risk consumers across all three segments, hence improving the intervention's efficacy and budget allocation, making it the ideal choice for maximizing customer retention. The results demonstrate the importance of segmentation in enhancing retention budgeting and budget optimization, particularly concerning parameter sensitivity.

Tambunan, Fiktor Januari; Tarigan, Perwira; Hulu, Yakin Rianto; Halawa, Hendi Jaya; Prabowo, Agung

Dinamik 2025 Universitas Stikubank

Abstrak Penelitian ini mengembangkan sistem deteksi aritmia pada lansia menggunakan sinyal elektrokardiogram (EKG) 5-lead dan algoritma Support Vector Machine (SVM). Data EKG yang diperoleh melalui perangkat Smart Holter direkam secara kontinu dan diproses melalui tahapan praproses, meliputi koreksi baseline, filtering dengan metode Butterworth, ekstraksi fitur, normalisasi, serta pelabelan manual oleh dokter spesialis jantung untuk validitas klinis. Model SVM kemudian dilatih dan diuji dengan hasil akurasi sebesar 95,80% pada data pelatihan dan 94,57% pada data pengujian. Evaluasi performa model menggunakan confusion matrix, nilai presisi, recall, dan kurva ROC menunjukkan kemampuan klasifikasi empat kategori aritmia secara akurat dan seimbang dengan nilai AUC antara 0,98 hingga 1,00. Hasil ini menunjukkan potensi sistem sebagai alat bantu diagnosis dini aritmia khususnya pada pasien lansia. Untuk penelitian selanjutnya, disarankan peningkatan variasi data, perbandingan dengan metode lain seperti CNN atau LSTM, peningkatan kualitas sinyal dan fitur, serta pengujian di lingkungan klinis guna mengoptimalkan penerapan sistem dalam praktik medis.    Kata Kunci: Elektrokardiogram (EKG), Aritmia, Support Vector Machine (SVM), Lansia

Nainggolan, Johannes Kristian; Sinaga, Ferdinand; Sitorus, Andriani M.; Khairia, Anisa; Wijaya, Bayu Angga

Dinamik 2025 Universitas Stikubank

Tingkat keberhasilan deteksi penyakit jantung sangat bergantung pada akurasi model klasifikasi yang digunakan. Penelitian ini bertujuan membandingkan kinerja dua algoritma klasifikasi, yaitu K-Nearest Neighbor (KNN) dan Support Vector Machine (SVM), dalam mendeteksi penyakit jantung menggunakan dataset berjumlah 1025 sampel dengan dua kelas target, yakni sehat dan penyakit jantung. Proses pra-pemrosesan data meliputi pembersihan dan normalisasi fitur medis seperti usia, tekanan darah, serta kadar kolesterol. Evaluasi performa model dilakukan menggunakan metode Confusion Matrix, K-Fold Cross Validation, kurva Receiver Operating Characteristic (ROC), dan kurva Precision-Recall untuk mengukur akurasi, presisi, recall, serta keseimbangan antara presisi dan recall. Hasil pengujian menunjukkan bahwa algoritma KNN unggul dalam menghasilkan akurasi tinggi yaitu 99% dengan AUC ROC sempurna 1.00 dan presisi yang hampir konsisten sepanjang recall, sementara SVM menunjukkan performa stabil dengan akurasi 91%, AUC ROC 0.97, dan AP Precision-Recall sebesar 0.96. Penelitian ini menegaskan efektivitas KNN dalam menghasilkan prediksi penyakit jantung yang sangat akurat dengan potensi risiko overfitting pada parameter k kecil, sedangkan SVM memberikan kestabilan model dengan kemampuan generalisasi yang lebih baik. Temuan ini diharapkan dapat menjadi referensi dalam pemilihan algoritma klasifikasi yang sesuai untuk mendukung diagnosis penyakit jantung secara klinis.

Gayatri Dwi Santika; Valiant Shabri Rabbani

Proceeding International Conference Of Innovation Science, Technology, Education, Children And Health 2025 Program Studi DIII Rekam Medis dan Informasi Kesehatan

Stroke is one of the leading causes of death globally and is particularly prevalent in Indonesia. Early prediction of stroke is critical to reducing the risk of long-term disability and mortality. This study aims to build a stroke prediction model using the Support Vector Machine (SVM) classification method. The dataset used is sourced from Kaggle, containing 5,110 records with class imbalance. To address the imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was applied during preprocessing. The study evaluates model performance across multiple data splits (70:30, 80:20, 90:10) and k-fold cross-validation values (k=5, 7, 10). The SVM was tested with various kernel types—linear, polynomial, and radial basis function (RBF)—along with parameter tuning for C, gamma, and degree. The results show that the polynomial kernel yielded the highest prediction accuracy of 92%. The model performance was evaluated using accuracy, precision, recall, and F1-score metrics.

Muzzakin, Muhamad; Pramono, Basworo Ardi; ., Susanto

Dinamik 2025 Universitas Stikubank

Bitcoin sebagai salah satu cryptocurrency paling populer yang menawarkan peluang investasi besar namun disertai dengan volatilitas harga yang tinggi. Penelitian ini bertujuan untuk memprediksi harga Bitcoin menggunakan model Support Vector Machine (SVM) dan menganalisis risiko pasar yang melekat. Dataset historis Bitcoin digunakan untuk melatih model dengan fitur seperti harga pembukaan, harga tertinggi, harga terendah, dan volume perdagangan. Penelitian menggunakan model SVM yang dioptimalkan melalui tuning parameter untuk meningkatkan akurasi prediksi. Evaluasi model dilakukan menggunakan metrik Mean Absolute Error (MAE) dan Root Mean Squared Error (RMSE). Hasil evaluasi menunjukkan performa model yang baik dengan MAE sebesar 0,0036 dan RMSE sebesar 0,0050. Korelasi fitur menunjukkan hubungan yang kuat antara harga penutupan dengan variabel harga lainnya, sementara volume memiliki hubungan moderat. Analisis risiko menggunakan pengembalian harian mengidentifikasi volatilitas signifikan, yang menjadi tantangan dalam pengambilan keputusan investasi. Penelitian ini menyimpulkan bahwa model SVM efektif dalam memprediksi tren harga Bitcoin. Namun, analisis risiko tetap penting untuk mendukung strategi investasi yang lebih bijaksana.    

Gayatri Dwi Santika; Valiant Shabri Rabbani

Proceeding International Conference Of Innovation Science, Technology, Education, Children And Health 2025 Program Studi DIII Rekam Medis dan Informasi Kesehatan

Stroke is one of the leading causes of death globally and is particularly prevalent in Indonesia. Early prediction of stroke is critical to reducing the risk of long-term disability and mortality. This study aims to build a stroke prediction model using the Support Vector Machine (SVM) classification method. The dataset used is sourced from Kaggle, containing 5,110 records with class imbalance. To address the imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was applied during preprocessing. The study evaluates model performance across multiple data splits (70:30, 80:20, 90:10) and k-fold cross-validation values (k=5, 7, 10). The SVM was tested with various kernel types—linear, polynomial, and radial basis function (RBF)—along with parameter tuning for C, gamma, and degree. The results show that the polynomial kernel yielded the highest prediction accuracy of 92%. The model performance was evaluated using accuracy, precision, recall, and F1-score metrics.

Rosa Ratri Kusuma Hariningsih; Diwahana Mutiara Candrasari; Endang Setyawati; Syamsu Wahidin; Jevon Nataniel Putra

International Journal of Computer Technology and Science 2025 Asosiasi Riset Teknik Elektro dan Infomatika Indonesia

Dengue Fever (DF) continues to be a major public health threat in Indonesia, especially in urban areas with high population density, such as Purwokerto City. This study aims to develop a predictive model to identify high-risk areas for DF outbreaks by integrating Machine Learning (ML) algorithms and Geographic Information Systems (GIS). The research utilizes historical dengue case data, meteorological parameters (rainfall, temperature, humidity), and population density as predictive variables. Three ML classification algorithms—Naïve Bayes, Logistic Regression, and Support Vector Machine (SVM)—were implemented to develop risk prediction models. Extensive data preprocessing, feature selection, and spatial integration were applied to ensure model robustness. The results show that the SVM model outperformed other methods, achieving the highest accuracy, precision, recall, and F1-score in classifying dengue risk zones. Risk maps generated through GIS visualization successfully identify priority areas for targeted interventions. The novelty of this research lies in the combination of local epidemiological data, multi-algorithm comparison, and geospatial mapping to improve early warning systems for DF in Purwokerto. This integrated approach is expected to support more effective prevention strategies and enhance public health preparedness.

Eugenea Chiquita Zahrani Assyarif; I Kadek Dwi Nuryana

Modem : Jurnal Informatika dan Sains Teknologi 2025 Asosiasi Profesi Telekomunikasi Dan Informatika Indonesia

This study aims to conduct customer segmentation and develop a classification model to predict the clusters of new customers at Monex Toys Abadi Bekasi, a micro, small, and medium enterprise (MSME). Segmentation was performed using the K-Means Clustering algorithm, incorporating parameters such as Recency, Frequency, Monetary (RFM), purchased products, payment methods, shipping cost discounts, and the total number of products purchased by customers. The segmentation results revealed two clusters: (1) Discount Hunters and (2) Loyal Customers. Subsequently, a classification process was conducted to predict customer clusters using the K-Nearest Neighbor (KNN) and Support Vector Machine (SVM) algorithms. Evaluation results indicated that all models achieved high accuracy exceeding 98%. The best-performing model was obtained with SVM using a 70:30 data split, achieving an accuracy of 98.81%. This classification model was then implemented into a Streamlit-based cluster prediction application, enabling users to identify customer segments in real-time. The findings of this research are expected to assist MSMEs in understanding customer behavior, enhancing service quality, and supporting more effective marketing strategies.

Sarassati, Dwi Sinta; Joko Prasetyo , Sri Yulianto

IT-Explore: Jurnal Penerapan Teknologi Informasi dan Komunikasi 2025 Fakultas Teknologi Informasi, Universitas Kristen Satya Wacana

Tidal flooding is an event of a natural phenomenon when sea water rises to land due to the influence of changes in sea tides, which causes waterlogging around the coastal area. This tidal flood hit the Demak-Semarang area, especially in the Sayung District area, which hampers and impacts community life. The purpose of this analysis is to analyze public sentiment regarding the impact of tidal flooding in Demak Regency using data obtained from social media, and the results of the analysis can be used as an evaluation for the government and related parties to formulate more responsive and effective policies to overcome the problem of tidal flooding. The SVM (Support Vector Machine) method is used to classify sentiment from each data into positive, negative, or neutral categories. The results of the analysis using SVM showed 3580 initial data, after preprocessing, 3147 data were obtained, with sentiment results of 1581 neutral opinions, 1257 negative, and 309 positive. Most opinions are neutral, indicating that people consider tidal flooding as a natural phenomenon and are used to dealing with it. However, significant negative opinions indicate dissatisfaction with the government's handling, while positive opinions are very minimal. SVM showed 84.44 percent accuracy, 86.7 percent precision, and 97.8 percent recall. The study recommends improvements in flood mitigation, assistance for affected communities, and infrastructure improvements.

Fitri Dwianasari; Rohmah Diah Yani; Karlina Novianto Laksono; Nurhafillah Mujaliza; Riza Fahlapi

Kajian Ekonomi dan Akuntansi Terapan 2025 Asosiasi Riset Ekonomi dan Akuntansi Indonesia

Mining activities in the Raja Ampat area have sparked various public reactions, both supportive and critical, particularly on social media platforms such as Twitter. This study aims to analyze public sentiment regarding the mining operations by employing two classification algorithms. A total of 500 tweets related to Raja Ampat were collected from the X platform, and after data cleaning, 168 were identified as positive sentiments and 303 as negative. Sentiment analysis was conducted using text mining techniques by comparing two algorithms: Support Vector Machine (SVM) and Naïve Bayes. To address the issue of data imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was applied. The analysis results showed that SVM achieved an accuracy of 80%, outperforming Naïve Bayes, which reached only 68%. This indicates that SVM performed better in classifying sentiment. Additionally, the application of SMOTE effectively enhanced both algorithms’ abilities to detect positive sentiment, as reflected in the precision, recall, and F1-score metrics. For SVM, precision reached 85%, recall 80%, and F1-score 80%, while Naïve Bayes recorded a precision and recall of 69%, and an F1-score of 68%.

Yayang Tika Robiatush Sholiha; Lubna Asjad Muhda Nabilah; Imron Imron

Saturnus: Jurnal Teknologi dan Sistem Informasi 2025 Asosiasi Riset Teknik Elektro dan Informatika Indonesia

This study aims to evaluate user sentiment toward the Liputan6.com application available on the Google Play Store. In the digital era, user reviews serve as a significant indicator in assessing the quality of an application. However, the inconsistency between rating scores and review content renders manual analysis less objective. To address this issue, a machine learning approach was adopted by comparing two algorithms, namely Support Vector Machine (SVM) and Naïve Bayes (NB). A total of 2,500 reviews were collected through a web scraping process and automatically labeled based on the rating (positive if ≥ 3, negative if < 3). The data preprocessing stages included cleaning, case folding, tokenizing, stopword removal, and token filtering. Subsequently, word weighting was carried out using the TF-IDF method, followed by classification using 10-Fold Cross Validation in RapidMiner. The evaluation results indicate that, in the positive class, NB demonstrated superior precision (89.47%), whereas SVM achieved higher recall (98.94%) and F1-score (90.96%). In the negative class, SVM performed better in terms of precision (66.15%), while NB attained higher recall (65.65%) and F1-score (36.34%). Further evaluation based on AUC and accuracy positioned SVM in the good category (AUC 0.842; accuracy 83.82%), while NB was categorized as fail (AUC 0.505; accuracy 60.87%). Overall, SVM is considered to be more effective than NB.