Skip to main navigation Skip to search Skip to main content

Measuring the Effectiveness of Large Language Model for Answer Generation: A Case of Quantum Software Engineering

  • Nek Dil Khan
  • , Safia Kanwal
  • , Javed Ali Khan
  • , Arif Ali Khan
  • , Muhammad Azeem Akbar

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Quantum Software Engineering (QSE) is an emerging field garnering interest among quantum researchers, developers, and tech giants for developing software to realize quantum computing's potential. Quantum developers frequently refer to Question and Answer (Q&A) platforms to address QSE challenges. However, the average response time to developers' questions on Q&A platforms is not immediate and can take days, leading to frustration. To address this delay, this study proposes an alternative approach for generating accurate answers for QSE-related topics using Large Language Models (LLMs). The automated LLM-based pipeline uses a few-shot prompting technique to examine whether a medium-sized LLM can be guided to generate domain-specific answers aligned with expert expectations. To evaluate the quality of these responses, we measure their semantic similarity against validated, human-authored answers. The experiments were conducted on a dataset of 263 quantum computing-related Q&A pairs, showing that the LLM achieved an average semantic similarity score of 64% compared to developers' answers. Moreover, a detailed analysis of low similarity scores was conducted to find potential reasons using the LLMs-as-Jury approach. The results demonstrate that only 4% cases are possibly hallucinated or provided incorrect information. These results highlight the promise of LLMs in assisting developers with contextually relevant information while underscoring their current limitations in addressing highly technical and domain-specific challenges. The proposed approach can serve as a stepping stone for enhanced developer experiences by integrating LLM with the Q&A platforms, providing immediate responses. Additionally, the approach can provide software developers with opportunities to refine their responses by analysing the LLM-generated output.
Original languageEnglish
Title of host publicationSAC '26: Proceedings of the 41st ACM/SIGAPP Symposium on Applied Computing
PublisherACM Press
Pages1343-1350
Number of pages8
ISBN (Electronic)9798400722943
DOIs
Publication statusPublished - 9 Jun 2026
EventSAC '26: 41st ACM/SIGAPP Symposium on Applied Computing - Grand Hotel Palace, Thessaloniki, Greece
Duration: 23 Mar 202627 Mar 2026

Conference

ConferenceSAC '26: 41st ACM/SIGAPP Symposium on Applied Computing
Country/TerritoryGreece
CityThessaloniki
Period23/03/2627/03/26

Keywords

  • DeepSeek
  • large language model
  • LLMs-as-jury
  • quantum software engineering
  • question answering

Fingerprint

Dive into the research topics of 'Measuring the Effectiveness of Large Language Model for Answer Generation: A Case of Quantum Software Engineering'. Together they form a unique fingerprint.

Cite this