ANALYSIS OF AUTOMATIC SPEECH RECOGNITION TECHNOLOGIES FOR ENTERING TEXTUAL INFORMATION INTO COMPUTER-ASSISTED PUBLISHING SYSTEMS
DOI:
https://doi.org/10.46299/j.isjea.20260505.08Keywords:
automatic speech recognition (ASR), desktop publishing system, text information, pre-press preparation, neural networks, Transformer, Whisper, WER, CERAbstract
This paper analyses modern ASR technologies for publishing systems. It proposes a methodology for assessing the quality of Ukrainian speech recognition and a model for integrating ASR into the pre-press preparation of publicationsReferences
Adobe. (2023). Tag content for XML in InDesign. Adobe Help Center. Available at: https://helpx.adobe.com/indesign/using/tagging-content-xml.html
Babu, A., Wang, C., Tjandra, A., Lakhotia, K., Xu, Q., Goyal, N., Singh, K., von Platen, P., Saraf, Y., Pino, J., Baevski, A., Conneau, A., & Auli, M. (2022). XLS-R: Self-supervised cross-lingual speech representation learning at scale. Proceedings of Interspeech 2022, 2278–2282. doi:10.21437/Interspeech.2022-143
Baevski, A., Zhou, Y., Mohamed, A., & Auli, M. (2020). wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in Neural Information Processing Systems, 33, 12449–12460. Available at: https://papers.neurips.cc/paper_files/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html
Conneau, A., Ma, M., Khanuja, S., Zhang, Y., Axelrod, V., Dalmia, S., Riesa, J., Rivera, C., & Bapna, A. (2023). FLEURS: Few-shot learning evaluation of universal representations of speech. 2022 IEEE Spoken Language Technology Workshop (SLT). doi:10.1109/SLT54892.2023.10023141
Graves, A., Fernández, S., Gomez, F., & Schmidhuber, J. (2006). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd International Conference on Machine Learning, 369–376. doi:10.1145/1143844.1143891
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., & Pang, R. (2020). Conformer: Convolution-augmented Transformer for speech recognition. Proceedings of Interspeech 2020, 5036–5040. doi:10.21437/Interspeech.2020-3015
KSE Research Group. (n.d.). Whisper-large-v3-dido-yvanchyk-v2 [Machine learning model]. Hugging Face. Available at: https://huggingface.co/KSE-RESEARCH-Group/whisper-large-v3-dido-yvanchyk-v2
Kyslyi, R., Orlovskyi, A., Khomenko, P., Onyshchenko, B., & Guzii, Z. (2026). Building ASR resources for the Hutsul dialect of Ukrainian. In Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects (pp. 186–195). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.vardial-1.15
Lipianina-Honcharenko, K., Bohuta, H., Ivaniush, A., & Soia, M. (2025). A dataset of real and synthetic speech in Ukrainian. Scientific Data, 12, Article 745. https://doi.org/10.1038/s41597-025-05084-8
Morris, A. C., Maier, V., & Green, P. (2004). From WER and RIL to MER and WIL: Improved evaluation measures for connected speech recognition. Proceedings of Interspeech 2004, 2765–2768. doi:10.21437/Interspeech.2004-668
Pratap, V., Tjandra, A., Shi, B., Tomasello, P., Babu, A., Kundu, S., Elkahky, A., Ni, Z., Vyas, A., Fazel-Zarandi, M., Baevski, A., Adi, Y., Zhang, X., Hsu, W.-N., Conneau, A., & Auli, M. (2024). Scaling speech technology to 1,000+ languages. Journal of Machine Learning Research, 25(97), 1–52. Available at: https://www.jmlr.org/beta/papers/v25/23-1318.html
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of the 40th International Conference on Machine Learning, 202, 28492–28518. Available at: https://proceedings.mlr.press/v202/radford23a
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., & Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. Proceedings of the 40th International Conference on Machine Learning, 202, 28492–28518. https://proceedings.mlr.press/v202/radford23a.html
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Yurii Syniavskyi, Zoryana Selmenska

This work is licensed under a Creative Commons Attribution 4.0 International License.




