Document Similarity Application Using Multiple Algorithms at UNDIPA Makassar

Authors

DOI:

https://doi.org/10.22303/csrid-.18.2.2026.320-332

Keywords:

Document Similarity, Bag-of-Words (BoW), TF-IDF, Jaccard Similarity, Word Embeddings, Doc2Vec, Sentence-BERT

Abstract

Measuring similarity between text documents is a fundamental task in Natural Language Processing (NLP) with broad applications, such as plagiarism detection, document clustering, and recommendation systems. This study aims to analyze and compare the performance of various document similarity algorithms, ranging from traditional lexical approaches to modern semantic methods. The algorithms reviewed include Bag-of-Words (BoW), TF-IDF, and Jaccard Similarity, as well as semantic representation-based methods such as Word Embeddings, Doc2Vec, and Sentence-BERT. An interactive web application was developed using the Gradio library to visualize comparison results in real-time and allow users to upload their own documents. The results indicate that lexical methods are effective at detecting keyword-based similarity but fail to capture semantic similarity when synonyms are used. Conversely, semantic methods—particularly Sentence-BERT—significantly outperform others in identifying contextual and semantic similarity, yielding more accurate scores for documents that differ in vocabulary yet share similar meanings. The study concludes that selecting the appropriate algorithm requires considering document characteristics and analysis objectives, and that the developed interactive tool can serve as an educational and experimental platform for such evaluations

References

Rajaraman, A., & Ullman, J. D. (2011). Mining of Massive Datasets. Cambridge University Press.

Cho, J., & Kim, H. (2016). A Survey on Document Similarity Measures. Journal of Information Science, 42(3), 323-339.

Luhn, H. P. (1957). A Statistical Approach to the Auto-encoding of Documents. IBM Journal of Research and Development.

Salton, G., & Buckley, C. (1988). Term-weighting approaches in automatic text retrieval. Information Processing & Management, 24(5), 513-523.

Jaccard, P. (1901). Étude comparative de la distribution florale dans une portion des Alpes et des Jura. Bulletin de la Société vaudoise des sciences naturelles.

Mikolov, T., et al. (2013). Efficient Estimation of Word Representations in Vector Space. ICLR Workshop.

Le, Q., & Mikolov, T. (2014). Distributed Representations of Sentences and Documents. International Conference on Machine Learning (ICML).

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing

Feldman, R., & Sanger, J. (2007). The Text Mining Handbook. Cambridge University Press.

Han, J., Kamber, M., & Pei, J. (2011). Data Mining: Concepts and Techniques. Morgan Kaufmann.

Young, T., Hazarika, D., Poria, S., & Cambria, E. (2018). Recent trends in deep learning based natural language processing. IEEE Computational Intelligence Magazine, 13(3), 55-71.

Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press.

Pedregosa, F., et al. (2011). Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, 2825-2830. [14] Gensim. (2023). Gensim Documentation. Diakses dari https://radimrehurek.com/gensim/

Honnibal, M., & Montani, I. (2017). spaCy: Industrial-strength Natural Language Processing in Python.

Gradio. (2023). Gradio Documentation. Diakses dari https://www.gradio.app/

Published

2026-06-01

Issue

Section

Articles

How to Cite

Document Similarity Application Using Multiple Algorithms at UNDIPA Makassar. (2026). CSRID (Computer Science Research and Its Development Journal), 18(2), 320-332. https://doi.org/10.22303/csrid-.18.2.2026.320-332

Similar Articles

1-10 of 25

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)