Implementation and Evaluation of Small Language Models for CRM Chatbots Using LoRA and RAG

Authors

DOI:

https://doi.org/10.22303/csrid-.18.2.2026.351-365

Keywords:

Small Language Model, LoRA, CRM

Abstract

The adoption of large language models (LLMs) in customer relationship management (CRM) has increased rapidly, but their use in micro, small, and medium enterprises (MSMEs) remains limited due to computational and cost constraints. This study explores the use of small language models (SLMs) as an efficient alternative for customer service chatbots, using parameter-efficient fine-tuning (PEFT) with Low-Rank Adaptation (LoRA) and comparing it with retrieval-augmented generation (RAG). This research follows the Design Science Research (DSR) approach with a case study on a local garment production business in Pontianak. The Qwen2-1.5B-Instruct model is adapted using LoRA and deployed on a 6GB GPU. Evaluation is conducted through quantitative and qualitative methods. Results show that the base model performs poorly without adaptation. The LoRA approach achieves the most stable performance, with intent accuracy up to 95%–100% and 0% hallucination rate, while RAG improves contextual understanding but lacks output consistency. The study concludes that domain-specific efficient fine-tuning is crucial to enabling SLM-based CRM solutions for SMEs.

References

Kurnia, E. (2025). AI makin terjangkau, investasi di Indonesia bisa lebih optimal. Diambil dari https://www.kompas.id/artikel/ai-makin-terjangkau-investasi-di-indonesia-bisa-lebih-optimal

Mohiuddin, K., et al. (2023). Attention is all you need. Proceedings of the 31st Conference on Neural Information Processing Systems, hal. 1–15.

Guizani, S., Mazhar, T., Shahzad, T., Ahmad, W., Bibi, A., & Hamam, H. (2025). A systematic literature review to implement large language model in higher education: Issues and solutions. Discover Education, 4(1).

Moenks, N., Penava, P., & Buettner, R. (2025). A systematic literature review of large language model applications in industry. IEEE Access, 13, hal. 160010–160033.

Mustafa, H., et al. (2023). The impact of using WhatsApp business (API) in marketing for small business. Proceedings of the 2023 10th International Conference on Social Networks Analysis, Management and Security (SNAMS), hal. 1–9.

Belcak, P., et al. (2025). Small language models are the future of agentic AI. Stanford University Press, hal. 1–17.

Xu, L., Xie, H., Qin, S. Z. J., Tao, X., & Wang, F. L. (2023). Parameter-efficient fine-tuning methods for pretrained language models: A critical review and assessment. IEEE Transactions on Pattern Analysis and Machine Intelligence, hal. 1–20.

Hu, E. J., et al. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations (ICLR), hal. 1–13.

Wang, F., et al. (2025). A survey on small language models in the era of large language models: Architecture, capabilities, and trustworthiness. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2, hal. 6173–6183.

Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems.

Guainazzo, N., Delzanno, G., Ancona, D., & D’Agostino, D. (2026). Navigating the seas of AI: Effectiveness of small language models on edge devices for maritime applications. Sensors, 26(5).

Li, X., Lu, Z., Cai, D., Ma, X., & Xu, M. (2024). Large language models on mobile devices: Measurements, analysis, and insights. EdgeFM 2024 - Proceedings of the 2024 Workshop on Edge and Mobile Foundation Models, hal. 1–6.

Wang, J., Zeng, Y., Guo, J., Ma, Y., Liu, A., & Liu, X. (2025). SLMQuant: Benchmarking Small Language Model Quantization for Practical Deployment (Vol. 1, No. 1). Association for Computing Machinery.

Zheng, Y., Chen, Y., Qian, B., Shi, X., Shu, Y., & Chen, J. (2025). A review on edge large language models: Design, execution, and applications. ACM Computing Surveys, 57(8).

Dutta, A., Ghosh, N., & Chatterjee, A. (2025). CARE: A QLoRA-fine tuned multi-domain chatbot with fast learning on minimal hardware.

Fan, J., Zhang, Y., Li, X., & Nikolopoulos, D. S. (2025). Parallel CPU-GPU execution for LLM inference on constrained GPUs.

Hugging Face. (2026). Qwen2-1.5B-Instruct. Diambil dari https://huggingface.co/Qwen/Qwen2-1.5B-Instruct

Blessing, L. T. M., & Chakrabarti, A. (2009). DRM: A design research methodology. Springer.

Creswell, J. W. (2014). Research design: Qualitative, quantitative, and mixed methods approaches (4th ed.). Sage Publishing.

Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+ questions for machine comprehension of text. Proceedings of EMNLP, hal. 1–10.

Jurafsky, D., & Martin, J. H. (2026). Speech and language processing. Stanford University.

Dettmers, T., Lewis, M., Shleifer, S., & Zettlemoyer, L. (2022). 8-bit optimizers via block-wise quantization. International Conference on Learning Representations (ICLR), hal. 1–20.

Reimers, N., & Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese BERT-networks. EMNLP-IJCNLP, hal. 3982–3992.

Manning, C. D., Raghavan, P., & Schütze, H. (2009). Introduction to information retrieval. Cambridge University Press.

Liu, N. F., et al. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, hal. 157–173.

Pedregosa, F., et al. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, hal. 633–642.

Howard, J., & Ruder, S. (2018). Universal language model fine-tuning

Published

2026-06-30

Issue

Section

Articles

How to Cite

Implementation and Evaluation of Small Language Models for CRM Chatbots Using LoRA and RAG. (2026). CSRID (Computer Science Research and Its Development Journal), 18(2), 351-365. https://doi.org/10.22303/csrid-.18.2.2026.351-365

Similar Articles

41-47 of 47

You may also start an advanced similarity search for this article.