{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T19:58:54Z","timestamp":1787947134461,"version":"build-2784847793"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2024,6,27]],"date-time":"2024-06-27T00:00:00Z","timestamp":1719446400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Italian Ministry of University and Research, Projects PRIN 2022","award":["2022S49T4W"],"award-info":[{"award-number":["2022S49T4W"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Softw. Eng. Methodol."],"published-print":{"date-parts":[[2024,7,31]]},"abstract":"<jats:p>\n                    Voice-based virtual assistants are becoming increasingly popular. Such systems provide frameworks to developers for building custom apps. End-users can interact with such apps through a Voice User Interface (VUI), which allows the user to use natural language commands to perform actions. Testing such apps is not trivial: The same command can be expressed in different semantically equivalent ways. In this article, we introduce VUI-UPSET, an approach that adapts chatbot-testing approaches to VUI-testing. We conducted an empirical study to understand how VUI-UPSET compares to two state-of-the-art approaches (i.e., a chatbot testing technique and ChatGPT) in terms of (i) correctness of the generated paraphrases, and (ii) capability of revealing bugs. To this aim, we analyzed 14,898 generated paraphrases for 40 Alexa Skills. Our results show that VUI-UPSET generates more bug-revealing paraphrases than the two baselines with, however, ChatGPT being the approach generating the highest percentage of correct paraphrases. We also tried to use the generated paraphrases to improve the skills. We tried to include in the\n                    <jats:italic>voice interaction models<\/jats:italic>\n                    of the skills (i) only the bug-revealing paraphrases, (ii) all the valid paraphrases. We observed that including only bug-revealing paraphrases is sometimes not sufficient to make all the tests pass.\n                  <\/jats:p>","DOI":"10.1145\/3654438","type":"journal-article","created":{"date-parts":[[2024,4,5]],"date-time":"2024-04-05T08:00:37Z","timestamp":1712304037000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Help Them Understand: Testing and Improving Voice User Interfaces"],"prefix":"10.1145","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5443-1303","authenticated-orcid":false,"given":"Emanuela","family":"Guglielmi","sequence":"first","affiliation":[{"name":"University of Molise, Pesche, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5241-1608","authenticated-orcid":false,"given":"Giovanni","family":"Rosa","sequence":"additional","affiliation":[{"name":"University of Molise, Pesche, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1764-9685","authenticated-orcid":false,"given":"Simone","family":"Scalabrino","sequence":"additional","affiliation":[{"name":"University of Molise, Pesche, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2216-3148","authenticated-orcid":false,"given":"Gabriele","family":"Bavota","sequence":"additional","affiliation":[{"name":"Universita della Svizzera Italiana, Lugano, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7995-8582","authenticated-orcid":false,"given":"Rocco","family":"Oliveto","sequence":"additional","affiliation":[{"name":"University of Molise, Pesche, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,6,27]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2022. Stop Word List. Retrieved 2018 from https:\/\/countwordsfree.com\/stopwords"},{"key":"e_1_3_1_3_2","unstructured":"\u201cAmazon\u201d. 2018. Alexa. Retrieved 2018 from https:\/\/developer.amazon.com\/en-US\/alexa"},{"key":"e_1_3_1_4_2","unstructured":"\u201cAmazon\u201d. 2018. Alexa Slots. Retrieved 2018 from https:\/\/developer.amazon.com\/en-US\/docs\/alexa\/custom-skills\/slot-type-reference.html"},{"key":"e_1_3_1_5_2","unstructured":"\u201cAmazon\u201d. 2018. Amazon Developer. Retrieved 2018 from https:\/\/developer.amazon.com\/en\/"},{"key":"e_1_3_1_6_2","unstructured":"\u201cAmazon\u201d. 2018. Amazon Official Documentation. Retrieved 2018 from https:\/\/developer.amazon.com\/en-US\/docs\/alexa\/custom-skills\/get-utterance-recommendations.html"},{"key":"e_1_3_1_7_2","unstructured":"\u201cAmazon\u201d. 2018. NLU-Evaluation Tool. Retrieved 2018 from https:\/\/developer.amazon.com\/it-IT\/docs\/alexa\/smapi\/nlu-evaluation-tool-api.html"},{"key":"e_1_3_1_8_2","doi-asserted-by":"crossref","unstructured":"Jordan J. Bird Anik\u00f3 Ek\u00e1rt and Diego R. Faria. 2023. Chatbot Interaction with Artificial Intelligence: human data augmentation with T5 and language transformer ensemble for text classification. Journal of Ambient Intelligence and Humanized Computing 14 4 (2023) 3129\u20133144.","DOI":"10.1007\/s12652-021-03439-8"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s12652-021-03439-8"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/AITest.2019.00-10"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-31280-0_3"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/BotSE52550.2021.00014"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S17-2001"},{"key":"e_1_3_1_14_2","unstructured":"\u201cChatGPT\u201d. 2023. ChatGpt. Retrieved from https:\/\/chat.openai.com. Access date: 2023."},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","unstructured":"Alexandru Coca Bo-Hsiang Tseng Weizhe Lin and Bill Byrne. 2023. More robust schema-guided dialogue state tracking via tree-based paraphrase ranking. In Findings of the Association for Computational Linguistics: (EACL\u201923). 1443\u20131454.","DOI":"10.18653\/v1\/2023.findings-eacl.106"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.5555\/975074"},{"issue":"1","key":"e_1_3_1_17_2","first-page":"2063","article-title":"Pattern for python","volume":"13","author":"Smedt Tom De","year":"2012","unstructured":"Tom De Smedt and Walter Daelemans. 2012. Pattern for python. The Journal of Machine Learning Research 13, 1 (2012), 2063\u20132067.","journal-title":"The Journal of Machine Learning Research"},{"key":"e_1_3_1_18_2","unstructured":"Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT. 4171\u20134186."},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","unstructured":"Adrian Egli. 2023. ChatGPT GPT-4 and other large language models: The next revolution for clinical microbiology? Clinical Infectious Diseases 77 9 (2023) 1322\u20131328.","DOI":"10.1093\/cid\/ciad407"},{"key":"e_1_3_1_20_2","unstructured":"\u201cHugging Face\u201d. Hugging Face squad_v2. Retrieved 2022 from https:\/\/huggingface.co\/datasets\/squad_v2\/viewer\/squad_v2\/train?p=4&row=440"},{"key":"e_1_3_1_21_2","unstructured":"\u201cHugging Face\u201d. 2022. Hugging Face. Retrieved 2022 from https:\/\/huggingface.co\/cross-encoder\/stsb-roberta-large"},{"key":"e_1_3_1_22_2","unstructured":"\u201cHugging Face\u201d. 2022. Hugging Face ambig_qa. Retrieved 2022 from https:\/\/huggingface.co\/datasets\/ambig_qa\/viewer\/full\/train. Access date: 2022."},{"key":"e_1_3_1_23_2","unstructured":"\u201cHugging Face\u201d. 2022. Hugging Face break_data. Retrieved 2022 from https:\/\/huggingface.co\/datasets\/break_data\/viewer\/logical-forms\/test?row=1"},{"key":"e_1_3_1_24_2","unstructured":"\u201cHugging Face\u201d. 2022. Hugging Face conv_ai_3. Retrieved 2022 from https:\/\/huggingface.co\/datasets\/conv_ai_3\/viewer\/conv_ai_3\/train?row=36"},{"key":"e_1_3_1_25_2","unstructured":"Emanuela Guglielmi Giovanni Rosa Simone Scalabrino Gabriele Bavota and Rocco Oliveto. 2022. Replication Package of \u201cHelp Them Understand: Testing and Improving Voice User Interfaces\u201d. Retrieved 2022 from https:\/\/figshare.com\/s\/36c3475659710714175d"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556934"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/AITest.2019.000-7"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.3115\/1621474.1621565"},{"key":"e_1_3_1_29_2","unstructured":"Chaitra Hegde and Shrikumar Patil. 2020. Unsupervised paraphrase generation using pre-trained language models. arXiv:2006.05477. Retrieved from https:\/\/arxiv.org\/abs\/2006.05477"},{"key":"e_1_3_1_30_2","doi-asserted-by":"crossref","unstructured":"Kuan-Hao Huang and Kai-Wei Chang. 2021. Generating syntactically controlled paraphrases without using annotated parallel pairs. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 1022\u20131033.","DOI":"10.18653\/v1\/2021.eacl-main.88"},{"key":"e_1_3_1_31_2","unstructured":"\u201cKayLearch\u201d. 2018. KayLearch. Retrieved 2018 from https:\/\/github.com\/KayLerch\/alexa-utterance-generator\/"},{"key":"e_1_3_1_32_2","unstructured":"Federica Laricchia. 2022. Number of Digital Voice Assistants in use Worldwide From 2019 to 2024. Retrieved 2022 from https:\/\/www.statista.com\/statistics\/973815\/worldwide-digital-voice-assistant-in-use\/"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPCC.2006.320364"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3551349.3556957"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1080\/10862967609547193"},{"key":"e_1_3_1_36_2","unstructured":"Yinhan Liu Myle Ott Naman Goyal Jingfei Du Mandar Joshi Danqi Chen Omer Levy Mike Lewis Luke Zettlemoyer and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. (2019)."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/1276933.1276934"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.11144\/Javeriana.upsy10-2.cdcp"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/2931037.2931054"},{"key":"e_1_3_1_41_2","first-page":"746","volume-title":"Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Mikolov Tom\u00e1\u0161","year":"2013","unstructured":"Tom\u00e1\u0161 Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013. Linguistic regularities in continuous space word representations. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 746\u2013751."},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/219717.219748"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICSE-C.2017.166"},{"key":"e_1_3_1_44_2","first-page":"1","volume-title":"Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS)","author":"Nicolich-Henkin Leah","year":"2021","unstructured":"Leah Nicolich-Henkin, Taichi Nakatani, Zach Trozenski, Joel Whiteman, and Nathan Susanj. 2021. Comparing data augmentation and annotation standardization to improve end-to-end spoken language understanding models. In Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS). 1\u20136."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.12928\/telkomnika.v18i3.14791"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jjimei.2021.100025"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","unstructured":"Ranci Ren Mireya Zapata John W. Castro Oscar Dieste and Silvia T. Acu\u00f1a. 2022. Experimentation for chatbot usability evaluation: A secondary study. IEEE Access 10 (2022) 12430\u201312464. DOI:10.1109\/ACCESS.2022.3145323","DOI":"10.1109\/ACCESS.2022.3145323"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.3390\/fi15060192"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3457913.3457931"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSE.2016.2532875"},{"key":"e_1_3_1_51_2","unstructured":"Siamak Shakeri and Abhinav Sethy. 2019. Label dependent deep variational paraphrase generation. arXiv:1911.11952. Retrieved from https:\/\/arxiv.org\/abs\/1911.11952"},{"key":"e_1_3_1_52_2","unstructured":"Alex Sokolov and Denis Filimonov. 2018. Neural machine translation for paraphrase generation. (2018)."},{"key":"e_1_3_1_53_2","unstructured":"Liling Tan. 2014. Pywsd: Python Implementations of Word Sense Disambiguation (wsd) Technologies [Software]. Retrieved 2014 from https:\/\/github.com\/alvations\/pywsd"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1129"},{"key":"e_1_3_1_55_2","unstructured":"Jason Wei Yi Tay Rishi Bommasani Colin Raffel Barret Zoph Sebastian Borgeaud Dani Yogatama Maarten Bosma Denny Zhou Donald Metzler Ed H. Chi Tatsunori Hashimoto Oriol Vinyals Percy Liang Jeff Dean and William Fedus. 2022. Emergent abilities of large language models. arXiv:2206.07682. Retrieved from https:\/\/arxiv.org\/abs\/2206.07682"},{"key":"e_1_3_1_56_2","doi-asserted-by":"crossref","unstructured":"Sam Witteveen and Martin Andrews. 2019. Paraphrasing with large language models. In Proceedings of the 3rd Workshop on Neural Generation and Translation. 215\u2013220.","DOI":"10.18653\/v1\/D19-5623"},{"key":"e_1_3_1_57_2","doi-asserted-by":"crossref","unstructured":"Robert F. Woolson. 2007. Wilcoxon signed-rank test. Wiley Encyclopedia of Clinical Trials (2007) 1\u20133.","DOI":"10.1002\/9780471462422.eoct979"},{"key":"e_1_3_1_58_2","doi-asserted-by":"crossref","unstructured":"Chen Zhang Luis Fernando D\u2019Haro Qiquan Zhang Thomas Friedrichs and Haizhou Li. 2023. Poe: A panel of experts for generalized automatic dialogue assessment. IEEE\/ACM Transactions on Audio Speech and Language Processing 31 (2023) 1234\u20131250.","DOI":"10.1109\/TASLP.2023.3250825"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.414"}],"container-title":["ACM Transactions on Software Engineering and Methodology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654438","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3654438","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:57:14Z","timestamp":1750276634000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3654438"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,27]]},"references-count":58,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2024,7,31]]}},"alternative-id":["10.1145\/3654438"],"URL":"https:\/\/doi.org\/10.1145\/3654438","relation":{},"ISSN":["1049-331X","1557-7392"],"issn-type":[{"value":"1049-331X","type":"print"},{"value":"1557-7392","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,6,27]]},"assertion":[{"value":"2023-06-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-27","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}