Acta Informatica Pragensia X:X | DOI: 10.18267/j.aip.32077

Enterprise Conversational AI with Retrieval-Augmented Generation: A CRM Case Study

Martin Sasinka ORCID..., Martin Kotyrba ORCID..., Eva Volna ORCID..., Martin Pavlicek ORCID...
Department of Informatics and Computers, Faculty of Science, University of Ostrava, Czech Republic

Background: The integration of Large Language Models (LLMs) into enterprise information systems represents a major challenge in Artificial Intelligence (AI) and digital transformation. Retrieval-Augmented Generation (RAG) has emerged as a promising approach to improve knowledge management and reduce information overload in enterprise environments.

Objective: This study aimed to design, implement, and evaluate a conversational assistant that combines RAG with fine-tuned LLMs in a Customer Relationship Management (CRM) environment. The research further sought to identify an optimal balance between performance, operational efficiency, and data governance.

Methods: Guided by Task-Technology Fit theory, four open-source models (LLaMA 3.2 3B, Mistral 7B, Qwen2.5 3B, and Gemma 3 4B) were fine-tuned using QLoRA on domain-specific CRM data and compared with the commercial baseline GPT-4o-mini. Evaluation included semantic metrics (BERTScore and embedding cosine similarity), traditional n-gram metrics (BLEU and ROUGE-L), and operational indicators such as inference latency and memory footprint.

Results: GPT-4o-mini achieved the highest semantic performance (BERTScore: 0.731; embedding similarity: 0.781) while maintaining low inference latency (1.89 s). Among the open-source alternatives, Mistral 7B achieved the best results (BERTScore: 0.711; similarity: 0.699). The findings demonstrate that open-source models can provide competitive performance while offering advantages in deployment flexibility and governance.

Conclusion: The study provides a practical framework for organizations deploying conversational AI in enterprise information systems, supporting decisions related to Total Cost of Ownership, operational efficiency, and data governance. Future research should further validate retrieval-aware metrics against human expert assessments and user-based task evaluations to better assess the business value of enterprise RAG systems.

Keywords: RAG; Large language model; LLM; Fine-tuning; Customer relationship management; Enterprise information systems; Task-technology fit.

Received: May 27, 2026; Revised: July 7, 2026; Accepted: July 7, 2026; Prepublished online: September 24, 2026 

Download citation

References

  1. Amugongo, L. M., Mascheroni, P., Brooks, S., Doering, S., & Seidel, J. (2025). Retrieval augmented generation for large language models in healthcare: A systematic review. PLOS Digital Health, 4(6), e0000877. https://doi.org/10.1371/journal.pdig.0000877 Go to original source...
  2. Balaguer, A., Benara, V., Cunha, R. L. D. F., Hendry, T., Holstein, D., Marsman, J., & Chandra, R. (2024). RAG vs fine-tuning: Pipelines, tradeoffs, and a case study on agriculture. arXiv preprint. arXiv:2401.08406. https://doi.org/10.48550/arXiv.2401.08406 Go to original source...
  3. Bazine, I., & Salahddine, M. A. (2025). Understanding the Task-Technology Fit (TTF) model: Genesis and conceptual developments. International Journal of Latest Technology in Engineering, Management & Applied Science, 14(8), 241-250. https://doi.org/10.51583/IJLTEMAS.2025.1408000030 Go to original source...
  4. Cohen, D., Burg, L., Pykhnivskyi, S., Gur, H., Kovynov, S., Atzmon, O., & Barkan, G. (2025). WixQA: A multi-dataset benchmark for enterprise retrieval-augmented generation. arXiv preprint. arXiv:2505.08643. https://doi.org/10.48550/arXiv.2505.08643 Go to original source...
  5. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. In 37th Conference on Neural Information Processing Systems (NeurIPS 2023), (pp. 10088-10115). NeurIPS. Go to original source...
  6. Es, S., James, J., Anke, L. E., & Schockaert, S. (2024). RAGAS: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, (pp. 150-158). ACL. https://doi.org/10.18653/v1/2024.eacl-demo.16 Go to original source...
  7. Filip, T., Pavlíček, M., & Sosík, P. (2024). Fine-tuning multilingual language models in Twitter/X sentiment analysis: A study on Eastern-European V4 languages. arXiv preprint. arXiv:2408.02044. https://doi.org/10.48550/arXiv.2408.02044 Go to original source...
  8. Goodhue, D. L., & Thompson, R. L. (1995). Task-technology fit and individual performance. MIS Quarterly, 19(2), 213-236. https://doi.org/10.2307/249689 Go to original source...
  9. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In Proceedings of ICLR 2022. ICLR.
  10. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1-38. https://doi.org/10.1145/3571730 Go to original source...
  11. Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., & Neubig, G. (2023). Active retrieval augmented generation. In Proceedings of EMNLP 2023 (pp. 7969-7992). https://doi.org/10.18653/v1/2023.emnlp-main.495 Go to original source...
  12. Karpukhin, V., Oguz, B., Min, S., Lewis, P. S., Wu, L., Edunov, S., & Yih, W. T. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, (pp. 6769-6781). ACL. https://doi.org/10.18653/v1/2020.emnlp-main.550 Go to original source...
  13. Landolsi, H., Letaief, K., Taghouti, N., & Abdeljaoued-Tej, I. (2025). CAPRAG: A large language model solution for customer service and automatic reporting using vector and graph retrieval-augmented generation. arXiv preprint. arXiv:2501.13993. https://doi.org/10.48550/arXiv.2501.13993 Go to original source...
  14. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459-9474.
  15. Li, F., Fang, P., Shi, Z., Khan, A., Wang, F., Feng, D., & Cui, Y. (2025). CoT-RAG: Integrating chain of thought and retrieval-augmented generation to enhance reasoning in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, (pp. 3119-3171). ACL. https://doi.org/10.18653/v1/2025.findings-emnlp.168 Go to original source...
  16. Lin, C. Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out (pp. 74-81). ACL.
  17. Liu, S., McCoy, A. B., & Wright, A. (2025). Improving large language model applications in biomedicine with retrieval-augmented generation: A systematic review, meta-analysis, and clinical development guidelines. Journal of the American Medical Informatics Association, 32(4), 605-615. https://doi.org/10.1093/jamia/ocaf009 Go to original source...
  18. Lu, Y. (2024). How much VRAM do you need to fine-tune LLMs? Modal Labs Blog. https://modal.com/blog/how-much-vram-need-fine-tuning
  19. Miao, J., Thongprayoon, C., Suppadungsuk, S., Garcia Valencia, O. A., & Cheungpasitporn, W. (2024). Integrating retrieval-augmented generation with large language models in nephrology: Advancing practical applications. Medicina, 60(3), 445. https://doi.org/10.3390/medicina60030445 Go to original source...
  20. Papineni, K., Roukos, S., Ward, T., & Zhu, W. J. (2002). BLEU: A method for automatic evaluation of machine translation. In ACL '02: Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, (pp. 311-318). https://doi.org/10.3115/1073083.1073135 Go to original source...
  21. Peng, B., Zhu, Y., Liu, Y., Bo, X., Shi, H., Hong, C., & Tang, S. (2025). Graph retrieval-augmented generation: A survey. ACM Transactions on Information Systems, 44(2), Article 35. https://doi.org/10.1145/3777378 Go to original source...
  22. Rangan, K., & Yin, Y. (2024). A fine-tuning enhanced RAG system with quantized influence measure as AI judge. Scientific Reports, 14(1), 27446. https://doi.org/10.1038/s41598-024-79110-x Go to original source...
  23. Randolph, C., Michaleas, A. M., & Ricke, D. O. (2025). Large language models for closed-library multi-document query, test generation, and evaluation. Frontiers in Artificial Intelligence, 8, 1592013. https://doi.org/10.3389/frai.2025.1592013 Go to original source...
  24. Ryu, C., Lee, S., Pang, S., Choi, C., Choi, H., Min, M., & Sohn, J. Y. (2023). Retrieval-based evaluation for LLMs: A case study in Korean legal QA. In Proceedings of the Natural Legal Language Processing Workshop 2023 (pp. 132-137). ACL. https://doi.org/10.18653/v1/2023.nllp-1.13 Go to original source...
  25. Solano, M. C., & Cruz, J. C. (2024). Integrating analytics in enterprise systems: A systematic literature review of impacts and innovations. Administrative Sciences, 14(7), Article 138. https://doi.org/10.3390/admsci14070138 Go to original source...
  26. Unlu, O., Shin, J., Mailly, C. J., Oates, M. F., Tucci, M. R., Varugheese, M., & Aronson, S. J. (2024). Retrieval augmented generation enabled GPT-4 performance for clinical trial screening. medRxiv. https://doi.org/10.1101/2024.02.08.24302376 Go to original source...
  27. Xu, Z., Cruz, M. J., Guevara, M., Wang, T., Deshpande, M., Wang, X., & Li, Z. (2024). Retrieval-augmented generation with knowledge graphs for customer service question answering. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2905-2909). https://doi.org/10.1145/3626772.3661370 Go to original source...
  28. Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In Proceedings of ICLR 2020. ICLR.

This is an open access article distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits use, distribution, and reproduction in any medium, provided the original publication is properly cited. No use, distribution or reproduction is permitted which does not comply with these terms.