home edit page issue tracker

This page pertains to UD version 2.

Guidelines for Spoken Language UD Treebanks

Current spoken treebanks

This page is a slim recap generated from per-treebank data files in workgroups/spoken-data/treebanks/. Each treebank has its own page with full metadata review details and a manual-check list. Advice references the standardized naming conventions in Metadata harmonisation. Use the search box, filters, or click a column header to sort. Contributor names are listed on each treebank’s page; contact emails are kept separately on the contributors and contacts page.

Note: this comparison was carried out semi-automatically with the help of Claude (Anthropic); errors or misunderstandings are possible, so please double-check anything unclear before acting on it.

Data version: the analysis below is based on the latest available version of each treebank (the master/dev branch HEAD as of May 2026); please check the treebank’s own repository for any updates published afterwards.

treebank type sentences tokens spoken identifiable? how identified document_id translations other text fields speaker metadata sound alignment general notes items to check issue draft
Abaza ATB only spoken 98 652 n/a - spoken data only `text_name`: this is a document identifier but is wrongly repeated on every… `text_orth` → `text_morphemic`; `text_transcription` → `text_transliteration` no speaker metadata exists at all. The 6 `text_name` filenames each seem to encode one… 3 draft ↗
Alemannic DIVITAL mixed 977 19334 yes via the `form` field: all 97 documents carry a `# form = ...` value `document_id`: OK `author` → `speaker_id` 1 draft ↗
Beja Autogramm only spoken 763 11948 n/a - spoken data only derive from recording basename `text_en` → `text_eng` `sent_timecode` → `sound_alignment_begin` 2 draft ↗
Bokota ChibErgIS only spoken 406 2713 n/a - spoken data only derive from recording basename; `sound_url` → move to document level `text_en` → `text_eng` `text_ortho` → `text_orthographic` `sent_timecode` → `sound_alignment_begin` 1 draft ↗
Bororo BDT mixed 21384 160356 no data-quality issue flagged (see page) 1 draft ↗
Cantonese HK only spoken 1004 13918 n/a - spoken data only add `# document_id` at sentences 1, 411, 548, and 651, using the proposed ids above… no speaker metadata exists; the legislative-council portion (sentences 651-1004) clearly… 2 draft ↗
Central_Romani Selice only spoken 0 0 no corpus appears empty (0 sentences) 0
Chinese HK only spoken 1004 9874 n/a - spoken data only add `# document_id` at sentences 1, 411, 548, and 651, using the proposed ids above… no speaker metadata exists; the legislative-council portion (sentences 651-1004) clearly… 5 draft ↗
Chukchi HSE only spoken 1004 5389 n/a - spoken data only derive from `sent_id` prefix (everything before `:`) `text[eng]` → `text_eng`; `text[rus]` → `text_rus` `text[eng']` → `text_eng_literal`; `text[phon]` → `text_phonetic` `timestamp` → `sound_alignment_begin` 0 draft ↗
Classical_Nahuatl FloCo mixed 0 0 no corpus appears empty (0 sentences) 0
Czech PDTC mixed 213897 3432078 yes via the `document_id` prefix `pdtsc` `global.Entity`: corpus-specific (coreference/entity annotation, project-wide) -… 1 draft ↗
Danish DDT mixed 5512 100733 no 1 draft ↗
Dargwa Mehweb only spoken 0 0 no corpus appears empty (0 sentences) 0
English CHILDES only spoken 48183 289817 n/a - spoken data only `corpus_name`: recompose: sort sentences by `original_sent_id` within each… `child_name` → `speaker_id`; `child_age` → `speaker_age`; `child_gender` → `speaker_gender`; `chi l d`: **data bug**, not a real field: a single malformed line (`# chi l d… 2 draft ↗
English ESLSpok only spoken 2320 21312 n/a - spoken data only derive from `sent_id` prefix (everything before `_<number>` no speaker metadata exists; each document is one L2 English speaker's interview session… 1 draft ↗
English GENTLE mixed 1334 17619 yes via `meta::genre = esports` `document_id`: OK `speaker` → `speaker_id` 3 draft ↗
English GUM mixed 14353 252284 yes via `meta::genre` `document_id`: OK `speaker` → `speaker_id` 3 draft ↗
French ParisStories only spoken 2776 42257 n/a - spoken data only derive from `sent_id` prefix (everything before the trailing `_<number>`); `sound_url` → move to document level `speaker` → `speaker_id` 1 draft ↗
French Rhapsodie only spoken 3209 43691 n/a - spoken data only derive from `sent_id` prefix (everything before the trailing `-<number>`); `sound_url` → move to document level `speaker`: corpus-specific turn-position label (`L1`, `L2`, ...), distinct from… 1 draft ↗
Frisian_Dutch Fame only spoken 400 3729 n/a - spoken data only `document_id`: OK `text_switch`: OK `speaker` (3rd segment, e.g. `sp0013f`) → `speaker_id`; `speaker` (2nd segment: `male`/`female`/`child`, 285/114/2 sentences) → `speaker_gender`; `speaker` (1st segment: `fr`/`nl`): split out; not a standard… 2 draft ↗
Gheg GPS only spoken 966 15990 n/a - spoken data only derive from `sent_id` prefix (everything before the trailing `_<number>`) consider adding `speaker_residence` (`Prishtina`/`Zurich`), derived from the `P`/`Z`…; consider a corpus-specific `speaker_generation` field (`G1`/`G2`/`G3`) - the…; add `speaker_age` if per-speaker ages are available… 1 draft ↗
Greek GDT mixed 2521 61773 yes via the source-outlet component embedded in `document_id` `document_id`: make tags: doc_id 2 draft ↗
Greek Lesbian mixed 625 6624 yes via the `oral_corpus` field, which marks sentences drawn from audio recordings as opposed to published… `text_el` → `text_ell` `text__el`: corpus-specific (sentence-level) - verify against metadata.html `source` (`Gender:...`): split out as `speaker_gender`; `source` (`Location:...`): split out as `speaker_residence`; `source` (`Date:...`): split out as a corpus-specific date field (no standard field covers…; `source`… 1 draft ↗
Hausa NorthernAutogramm only spoken 1305 15324 n/a - spoken data only `text_en` → `text_eng` `text_ortho` → `text_orthographic` `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) 1 draft ↗
Hausa SouthernAutogramm only spoken 1927 14398 n/a - spoken data only `text_en` → `text_eng` `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) 1 draft ↗
Hausa WesternAutogramm mixed 775 13862 no `text_en` → `text_eng` `text_ortho` → `text_orthographic` `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) 2 draft ↗
Hebrew IAHLTknesset mixed 2883 50499 no `document_id`: OK `speaker` → `speaker_id` 1 draft ↗
Highland_Puebla_Nahuatl ITML mixed 1260 10018 yes via `sent_id`: sentences from spoken material carry a `.eaf` `text[spa]` → `text_spa` `text[orig]` → `text_transcription`; `text[gloss]` → `text_glossing`; `text[glosa]`: typo 2 draft ↗
Ika ChibErgIS only spoken 628 5307 n/a - spoken data only `text_en` → `text_eng` `text_phrase-gls-es` → `text_esp`; `text_phrase-gls-tl`: not sure what this is `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) 1 draft ↗
Italian KIParlaForest only spoken 2221 18050 n/a - spoken data only 0 draft ↗
Japanese JDD only spoken 0 0 no corpus appears empty (0 sentences) 0
Khoekhoe KDT mixed 3589 27611 yes via the `document_id` prefix, which names the source type: `book` 2 draft ↗
Khunsari AHA mixed 10 74 n/a `text_en` → `text_eng`; `text_fa` → `text_fas` 0
Komi_Zyrian IKDP only spoken 214 2304 n/a - spoken data only derive from `sent_id` prefix identifying the source recording (please confirm… `text_en`: make tags: text_en; `text_ru`: text_rus; `text_end`: text_en 1 draft ↗
Latvian LVTB mixed 19580 330318 no 1 draft ↗
Ligurian GLT mixed 316 6568 no 2 draft ↗
Naija NSC only spoken 9241 140837 n/a - spoken data only derive from `sent_id` prefix identifying the source recording `text_en` → `text_eng` `text_ortho` → `text_orthographic` 1 draft ↗
Nayini AHA mixed 10 78 n/a `text_en` → `text_eng`; `text_fa` → `text_fas` 0
Nenets Tundra only spoken 170 1272 n/a - spoken data only `doc_title_`: use as `# document_id` (rename/repurpose the field) `text_en` → `text_eng`; `text_ru` → `text_rus` `text_p`: unclear 2 draft ↗
Nheengatu CompLin mixed 2839 26444 no `title` → `document_id` `text_eng`, `text_por`, `text_rus`: OK `speaker` → `speaker_id`; `speaker_gender`: OK 2 draft ↗
Northwest_Gbaya Autogramm only spoken 403 2692 n/a - spoken data only `text_fr` → `text_fra` `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) 1 draft ↗
Norwegian NynorskLIA only spoken 5250 55410 n/a - spoken data only `dialect`: OK 0 draft ↗
Persian Seraji mixed 5997 151627 no `text_en` → `text_eng` 2 draft ↗
Pesh ChibErgIS only spoken 524 4275 n/a - spoken data only derive from `sent_id` prefix identifying the source recording `text_en` → `text_eng` `text_phrase-gls-es` → `text_spa`; `text_phrase-gls-it` → `text_phonetic`; `text_phrase-gls-pro` → `text_prosodic`; `text_phrase-gls-tl`: corpus-specific - verify against metadata.html; `text_phrase-gls-de`:… `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) 0 draft ↗
Polish LFG mixed 17246 130967 yes via the sentence-level `genre` field 3 draft ↗
Scottish_Gaelic ARCOSG mixed 4748 86139 yes via the letter prefix of `document_id` `speaker` → `speaker_id` 0 draft ↗
Skolt_Sami Giellagas mixed 261 3049 yes per README description derive from `sent_id` source-identifier prefix `text_fi` → `text_fin`; `text_en` → `text_eng` 3 draft ↗
Slovenian SST only spoken 6121 98393 n/a - spoken data only 1 draft ↗
Soi AHA mixed 8 55 yes per README description `text_en` → `text_eng`; `text_fa` → `text_fas` declared type/genre may not match actual content - see modality note 0 draft ↗
South_Levantine_Arabic MADAR mixed 100 789 no declared type/genre may not match actual content - see modality note 1 draft ↗
Spanish COSER only spoken 539 7987 n/a - spoken data only `turn_time` → split (`sound_alignment_begin` and `sound_alignment_end`; derive `duration`); `time` → split (`sound_alignment_begin` and `sound_alignment_end` (same as…); `turn_time` → split (`sound_alignment_begin` and… 1 draft ↗
Swedish_Sign_Language SSLC only spoken 203 1610 n/a - spoken data only derive from `sent_id` prefix (everything before the first `:`) 0 draft ↗
Telugu_English TECT only spoken 97 456 no declared type/genre may not match actual content - see modality note 1 draft ↗
Turkish_English BUTR only spoken 58 441 no `text_en` → `text_eng` declared type/genre may not match actual content - see modality note 2 draft ↗
Turkish_German SAGT only spoken 2184 36934 n/a - spoken data only derive from `sent_id` prefix (everything before the trailing `-<number>`) 2 draft ↗
Ukrainian ParlaMint mixed 7142 109166 yes per README description derive `# document_id = NSDC_UA_28_Feb2014` for the NSDC-sourced sentences (`sent_id`… `text_en`: remove (single occurrence, placeholder value `undefined undefined`) 1 draft ↗
Vietnamese TueCL only spoken 100 1888 n/a - spoken data only 0
Western_Armenian ArmTDP mixed 6644 121432 yes via the `document_id` prefix, which already encodes genre `doc_title`: `document_id` already exists separately (e.g. `spoken-002R`);… 1 draft ↗
Western_Sierra_Puebla_Nahuatl MesoTree mixed 3024 19191 no `text[spa]` → `text_spa`; `text[eng]` → `text_eng` `text[orig]` → `text_original`; `text[morf]` → `text_morphemic`; `text[gloss]` → `text_glossing`; `text[orig_omitlan]`: corpus-specific (sentence-level) - verify against metadata.html; `text[orig_smt]`: corpus-specific… `user_id`: corpus-specific (speaker/paragraph-level) - verify against…; `finished`: corpus-specific (speaker/paragraph-level) - verify against…; `location`: corpus-specific (speaker/paragraph-level) - verify against…;… `timestamp` → `sound_alignment_begin` 13 draft ↗
Yiddish YiTB mixed 3113 27954 yes via the sentence-level `genre` field `text_en` → `text_eng` `rtl`: corpus-specific (speaker/paragraph-level) - verify against…; `source`: corpus-specific (speaker/paragraph-level) - verify against… 3 draft ↗
Zazaki ZSD only spoken 200 1371 n/a - spoken data only add `# document_id = Seyristane_dialogue` corpus-wide (single document) `text_en` → `text_eng` 0 draft ↗

Harmonization Pipeline:

  1. is spoken part clearly identifiable?
    1. yes > add # modality = spoken to relevant sentences
    2. no > open ISSUE with text XXXX

Workflow