Deteksi Komentar Toxic terhadap Isu BBM Etanol menggunakan Naive Bayes dengan TF-IDF dan Lexicon Abusive Words.
Loading...
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Fakultas Ilmu Komputer
Abstract
Social media comments regarding the ethanol-blended fuel policy often contain high levels of toxicity, which are difficult to detect due to the characteristics of informal language, slang, and a wide variety of abusive regional expressions. To address these challenges, this study proposes the integration of an additional Lexicon Abusive Words feature to increase model sensitivity. Traditional Machine Learning using the Naive Bayes method serves as the classification basis, supported by TF-IDF term weighting. Based on testing conducted on 2,959 comment entries, the model achieved optimal performance with an accuracy of 84% at a 90:10 data split ratio. The analysis results indicate that the inclusion of the lexicon feature significantly reduced the number of False Positives, demonstrating improved accuracy in distinguishing between toxic and non-toxic contexts. However, error analysis reveals that False Negatives still occur due to limited dictionary coverage, the model's inability to interpret sarcasm, and cultural discrepancies between the annotators' perspectives and the actual context of the comments. This study concludes that the effectiveness of toxicity detection does not depend solely on algorithm sophistication, but also on the comprehensiveness of the supporting database.
Description
:: Finalisasi file repositori 24 September 2026_Kurnadi
