World and domestic experience in the development of spoken corpora
DOI:
https://doi.org/10.31489/2026phi3(123)/150-161Abstract
Modern linguistics places significant emphasis on the development of national language corpora. In the era of rapid technological advancement, particularly the rise of artificial intelligence, the study of spoken subcorpora – linguistic databases containing samples of spoken language—has become increasingly relevant and requires focused scholarly attention. These subcorpora are essential tools for analyzing phonetic, prosodic, syntactic, and pragmatic aspects of language use. This article explores the history of the creation of the earliest spoken corpora, both internationally and within domestic linguistic traditions. It provides an overview
of the current state of corpus linguistic research, with particular attention to contemporary methodologies for compiling and annotating spoken corpora. The article compares and describes the construction principles of several major spoken corpora, including the Spoken Subcorpus of the Russian National Corpus, the Corpus of Russian Everyday Speech (developed by scholars at St. Petersburg State University), the Spoken Component of the British National Corpus, and the Arabic Speech Corpus. Special emphasis is placed on the analytical review of metadata annotation practices, as well as the use of orthoepic and orthographic models in the annotation of linguistic data. In addition, the article provides a comprehensive description of the methods of collecting lexical material and the content of the Oral Subcorpus—one of the components of the National Corpus of the Kazakh Language, prepared by the researchers of the A. Baitursynuly Institute of Linguistics. The Oral Subcorpus was initially enriched with public speeches, interviews, and both formal and informal utterances. Alongside the spoken language of ordinary workers and local residents, audio and video recordings of poets, writers, and public figures began to be collected. Since the 1960s, materials from speeches by writers such as M. Auezov, G. Musrepov, S. Mukanov, and other prominent figures have been gathered. Their texts under went orthographic transcription, orthoepic annotation, and prosodic analysis. This, on the one hand, made it possible to build a corpus of Kazakh oral speech from different periods, and on the other, to obtain written versions of speeches by prominent figures that had previously existed only in oral form on magnetic tapes. Since 2024, samples of contemporary colloquial speech and regional varieties have been added to the Oral Subcorpus. According to the project’s concept, in 2025 the developers of the Oral Corpus at the A. Baitursynuly Institute of Linguistics launched expeditions across various regions of Kazakhstan to collect stories from local residents familiar with traditional norms of the Kazakh language and oral culture. These stories were documented, recorded, and included in the National Corpus of the Kazakh Language for further research. During these expeditions, special attention was given to capturing traditional word usage, unpre pared, spontaneous, and natural speech. Audio and video recordings were made of master speakers— individuals with clear thinking, expressive language, and skill in improvisation. All collected materials have
been incorporated into the «National Corpus of the Kazakh Language».








