BS ISO 24624-2016 PDF
Name in English:
STB BS ISO 24624-2016
Name in Russian:
СТБ BS ISO 24624-2016
Original standard BS ISO 24624-2016 in PDF full version. Additional info + preview on request
Full title and description
STB BS ISO 24624-2016 — Language resource management — Transcription of spoken language. This international standard specifies rules for representing transcriptions of audio- and video-recorded spoken interactions in XML documents based on the guidelines of the TEI, and relates transcribed data to standards for annotated corpora.
Abstract
ISO 24624:2016 defines a TEI-based XML encoding model and a set of rules for encoding spoken-language transcriptions so they can be exchanged, searched and linked with other language-resource annotations. It is intended for sociolinguistics, conversation analysis, dialectology, corpus linguistics, lexicography, language technology and qualitative social research; it explicitly excludes transcription formats for handwritten manuscripts. Annex A offers a fully encoded example; Annex B supplies element and attribute indexes.
General information
- Status: Published / confirmed (international standard).
- Publication date: 2016-08 (first published mid‑2016; confirmed in 2022 during routine 5‑year review).
- Publisher: International Organization for Standardization (ISO); adopted nationally in identical form by several bodies (e.g., BS ISO).
- ICS / categories: 01.140.10 (Writing and transliteration / language resource management).
- Edition / version: Edition 1 (2016).
- Number of pages: 32 pages (ISO edition; some national adoptions include additional national foreword material and may show a larger page count).
Scope
ISO 24624:2016 applies to the representation of transcriptions of recorded spoken language (audio and video). It specifies an XML encoding model based on TEI guidelines to ensure interoperable, machine-processable transcriptions and to facilitate linking transcribed data with standards for annotated corpora. The standard is not intended for other transcription types, notably transcriptions of handwritten manuscripts. Annexes provide a full example and indexes for elements and attributes.
Key topics and requirements
- TEI-based XML encoding model and recommended elements/attributes for spoken‑language transcription (tiers for orthography, overlap, timing, non‑verbal events, etc.).
- Guidelines for linking transcription segments to media (time alignment) and to external annotations/corpora.
- Conventions to represent disfluencies, repairs, overlaps, pauses, and non‑verbal behaviour in a structured form.
- Compatibility notes and references to related ISO standards and data formats (date/time, character sets, morpho-syntactic frameworks) to promote interoperability.
- Example encoding (Annex A) and element/attribute indexes (Annex B) to support implementation and tool development.
Typical use and users
Researchers and engineers in corpus linguistics, conversation analysis, sociolinguistics, dialectology, lexicography, and language technology; developers of annotation tools, transcription workflows and language-resource repositories; archives and data centers that store, exchange or publish transcribed spoken data. The standard is used when interoperable, machine-readable transcriptions and linkage to annotated corpora are required.
Related standards
ISO 24624 is part of the ISO 246xx family for language resource management and cross-references/relies on other standards such as ISO 24611 (MAF — morpho-syntactic annotation framework), ISO 8601 (date/time formats), and character-set standards for interoperable encoding. National adoptions (e.g., BS ISO 24624:2016) are identical or formally adopted versions.
Keywords
transcription; spoken language; TEI; XML; language resources; corpus annotation; time alignment; conversation analysis; sociolinguistics; ISO 24624.
FAQ
Q: What is this standard?
A: ISO 24624:2016 is an international standard that prescribes a TEI-based XML model and rules for representing transcriptions of recorded spoken interactions so they are interoperable and machine‑processable.
Q: What does it cover?
A: It covers encoding conventions for orthography, timing, overlaps, disfluencies, non‑verbal events, media linking and how transcriptions can be related to annotated corpora; it includes example encodings and indexes of elements/attributes. It does not cover manuscript transcription.
Q: Who typically uses it?
A: Linguists, corpus builders, tool developers, archives, and language‑technology practitioners who need standardised, exchangeable spoken‑language transcriptions.
Q: Is it current or superseded?
A: The 2016 edition remains current; the ISO record shows the standard was confirmed in the routine 5‑year review (confirmation recorded in 2022). National adoptions (e.g., BS ISO 24624:2016) reflect the same content.
Q: Is it part of a series?
A: Yes. ISO 24624 sits within the ISO 246xx family addressing language resource management (corpus models, annotation frameworks and related metadata standards); it is intended to interoperate with other parts of that family.
Q: What are the key keywords?
A: TEI, XML, transcription, spoken language, corpus, annotation, time alignment, disfluency, conversation analysis.