Guidelines for Spoken Language UD Treebanks
Current spoken treebanks
This page is a slim recap generated from per-treebank data files in workgroups/spoken-data/treebanks/. Each treebank has its own page with full metadata review details and a manual-check list. Advice references the standardized naming conventions in Metadata harmonisation. Use the search box, filters, or click a column header to sort. Contributor names are listed on each treebank’s page; contact emails are kept separately on the contributors and contacts page.
Note: this comparison was carried out semi-automatically with the help of Claude (Anthropic); errors or misunderstandings are possible, so please double-check anything unclear before acting on it.
Data version: the analysis below is based on the latest available version of each treebank (the master/dev branch HEAD as of May 2026); please check the treebank’s own repository for any updates published afterwards.
| treebank | type | sentences | tokens | spoken identifiable? | how identified | document_id | translations | other text fields | speaker metadata | sound alignment | general notes | items to check | issue draft |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Abaza ATB | only spoken | 98 | 652 | n/a - spoken data only | `text_name`: this is a document identifier but is wrongly repeated on every… | `text_orth` → `text_morphemic`; `text_transcription` → `text_transliteration` | no speaker metadata exists at all. The 6 `text_name` filenames each seem to encode one… | 3 | draft ↗ | ||||
| Alemannic DIVITAL | mixed | 977 | 19334 | yes | via the `form` field: all 97 documents carry a `# form = ...` value | `document_id`: OK | `author` → `speaker_id` | 1 | draft ↗ | ||||
| Beja Autogramm | only spoken | 763 | 11948 | n/a - spoken data only | derive from recording basename | `text_en` → `text_eng` | `sent_timecode` → `sound_alignment_begin` | 2 | draft ↗ | ||||
| Bokota ChibErgIS | only spoken | 406 | 2713 | n/a - spoken data only | derive from recording basename; `sound_url` → move to document level | `text_en` → `text_eng` | `text_ortho` → `text_orthographic` | `sent_timecode` → `sound_alignment_begin` | 1 | draft ↗ | |||
| Bororo BDT | mixed | 21384 | 160356 | no | data-quality issue flagged (see page) | 1 | draft ↗ | ||||||
| Cantonese HK | only spoken | 1004 | 13918 | n/a - spoken data only | add `# document_id` at sentences 1, 411, 548, and 651, using the proposed ids above… | no speaker metadata exists; the legislative-council portion (sentences 651-1004) clearly… | 2 | draft ↗ | |||||
| Central_Romani Selice | only spoken | 0 | 0 | no | corpus appears empty (0 sentences) | 0 | — | ||||||
| Chinese HK | only spoken | 1004 | 9874 | n/a - spoken data only | add `# document_id` at sentences 1, 411, 548, and 651, using the proposed ids above… | no speaker metadata exists; the legislative-council portion (sentences 651-1004) clearly… | 5 | draft ↗ | |||||
| Chukchi HSE | only spoken | 1004 | 5389 | n/a - spoken data only | derive from `sent_id` prefix (everything before `:`) | `text[eng]` → `text_eng`; `text[rus]` → `text_rus` | `text[eng']` → `text_eng_literal`; `text[phon]` → `text_phonetic` | `timestamp` → `sound_alignment_begin` | 0 | draft ↗ | |||
| Classical_Nahuatl FloCo | mixed | 0 | 0 | no | corpus appears empty (0 sentences) | 0 | — | ||||||
| Czech PDTC | mixed | 213897 | 3432078 | yes | via the `document_id` prefix `pdtsc` | `global.Entity`: corpus-specific (coreference/entity annotation, project-wide) -… | 1 | draft ↗ | |||||
| Danish DDT | mixed | 5512 | 100733 | no | 1 | draft ↗ | |||||||
| Dargwa Mehweb | only spoken | 0 | 0 | no | corpus appears empty (0 sentences) | 0 | — | ||||||
| English CHILDES | only spoken | 48183 | 289817 | n/a - spoken data only | `corpus_name`: recompose: sort sentences by `original_sent_id` within each… | `child_name` → `speaker_id`; `child_age` → `speaker_age`; `child_gender` → `speaker_gender`; `chi l d`: **data bug**, not a real field: a single malformed line (`# chi l d… | 2 | draft ↗ | |||||
| English ESLSpok | only spoken | 2320 | 21312 | n/a - spoken data only | derive from `sent_id` prefix (everything before `_<number>` | no speaker metadata exists; each document is one L2 English speaker's interview session… | 1 | draft ↗ | |||||
| English GENTLE | mixed | 1334 | 17619 | yes | via `meta::genre = esports` | `document_id`: OK | `speaker` → `speaker_id` | 3 | draft ↗ | ||||
| English GUM | mixed | 14353 | 252284 | yes | via `meta::genre` | `document_id`: OK | `speaker` → `speaker_id` | 3 | draft ↗ | ||||
| French ParisStories | only spoken | 2776 | 42257 | n/a - spoken data only | derive from `sent_id` prefix (everything before the trailing `_<number>`); `sound_url` → move to document level | `speaker` → `speaker_id` | 1 | draft ↗ | |||||
| French Rhapsodie | only spoken | 3209 | 43691 | n/a - spoken data only | derive from `sent_id` prefix (everything before the trailing `-<number>`); `sound_url` → move to document level | `speaker`: corpus-specific turn-position label (`L1`, `L2`, ...), distinct from… | 1 | draft ↗ | |||||
| Frisian_Dutch Fame | only spoken | 400 | 3729 | n/a - spoken data only | `document_id`: OK | `text_switch`: OK | `speaker` (3rd segment, e.g. `sp0013f`) → `speaker_id`; `speaker` (2nd segment: `male`/`female`/`child`, 285/114/2 sentences) → `speaker_gender`; `speaker` (1st segment: `fr`/`nl`): split out; not a standard… | 2 | draft ↗ | ||||
| Gheg GPS | only spoken | 966 | 15990 | n/a - spoken data only | derive from `sent_id` prefix (everything before the trailing `_<number>`) | consider adding `speaker_residence` (`Prishtina`/`Zurich`), derived from the `P`/`Z`…; consider a corpus-specific `speaker_generation` field (`G1`/`G2`/`G3`) - the…; add `speaker_age` if per-speaker ages are available… | 1 | draft ↗ | |||||
| Greek GDT | mixed | 2521 | 61773 | yes | via the source-outlet component embedded in `document_id` | `document_id`: make tags: doc_id | 2 | draft ↗ | |||||
| Greek Lesbian | mixed | 625 | 6624 | yes | via the `oral_corpus` field, which marks sentences drawn from audio recordings as opposed to published… | `text_el` → `text_ell` | `text__el`: corpus-specific (sentence-level) - verify against metadata.html | `source` (`Gender:...`): split out as `speaker_gender`; `source` (`Location:...`): split out as `speaker_residence`; `source` (`Date:...`): split out as a corpus-specific date field (no standard field covers…; `source`… | 1 | draft ↗ | |||
| Hausa NorthernAutogramm | only spoken | 1305 | 15324 | n/a - spoken data only | `text_en` → `text_eng` | `text_ortho` → `text_orthographic` | `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) | 1 | draft ↗ | ||||
| Hausa SouthernAutogramm | only spoken | 1927 | 14398 | n/a - spoken data only | `text_en` → `text_eng` | `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) | 1 | draft ↗ | |||||
| Hausa WesternAutogramm | mixed | 775 | 13862 | no | `text_en` → `text_eng` | `text_ortho` → `text_orthographic` | `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) | 2 | draft ↗ | ||||
| Hebrew IAHLTknesset | mixed | 2883 | 50499 | no | `document_id`: OK | `speaker` → `speaker_id` | 1 | draft ↗ | |||||
| Highland_Puebla_Nahuatl ITML | mixed | 1260 | 10018 | yes | via `sent_id`: sentences from spoken material carry a `.eaf` | `text[spa]` → `text_spa` | `text[orig]` → `text_transcription`; `text[gloss]` → `text_glossing`; `text[glosa]`: typo | 2 | draft ↗ | ||||
| Ika ChibErgIS | only spoken | 628 | 5307 | n/a - spoken data only | `text_en` → `text_eng` | `text_phrase-gls-es` → `text_esp`; `text_phrase-gls-tl`: not sure what this is | `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) | 1 | draft ↗ | ||||
| Italian KIParlaForest | only spoken | 2221 | 18050 | n/a - spoken data only | 0 | draft ↗ | |||||||
| Japanese JDD | only spoken | 0 | 0 | no | corpus appears empty (0 sentences) | 0 | — | ||||||
| Khoekhoe KDT | mixed | 3589 | 27611 | yes | via the `document_id` prefix, which names the source type: `book` | 2 | draft ↗ | ||||||
| Khunsari AHA | mixed | 10 | 74 | n/a | `text_en` → `text_eng`; `text_fa` → `text_fas` | 0 | — | ||||||
| Komi_Zyrian IKDP | only spoken | 214 | 2304 | n/a - spoken data only | derive from `sent_id` prefix identifying the source recording (please confirm… | `text_en`: make tags: text_en; `text_ru`: text_rus; `text_end`: text_en | 1 | draft ↗ | |||||
| Latvian LVTB | mixed | 19580 | 330318 | no | 1 | draft ↗ | |||||||
| Ligurian GLT | mixed | 316 | 6568 | no | 2 | draft ↗ | |||||||
| Naija NSC | only spoken | 9241 | 140837 | n/a - spoken data only | derive from `sent_id` prefix identifying the source recording | `text_en` → `text_eng` | `text_ortho` → `text_orthographic` | 1 | draft ↗ | ||||
| Nayini AHA | mixed | 10 | 78 | n/a | `text_en` → `text_eng`; `text_fa` → `text_fas` | 0 | — | ||||||
| Nenets Tundra | only spoken | 170 | 1272 | n/a - spoken data only | `doc_title_`: use as `# document_id` (rename/repurpose the field) | `text_en` → `text_eng`; `text_ru` → `text_rus` | `text_p`: unclear | 2 | draft ↗ | ||||
| Nheengatu CompLin | mixed | 2839 | 26444 | no | `title` → `document_id` | `text_eng`, `text_por`, `text_rus`: OK | `speaker` → `speaker_id`; `speaker_gender`: OK | 2 | draft ↗ | ||||
| Northwest_Gbaya Autogramm | only spoken | 403 | 2692 | n/a - spoken data only | `text_fr` → `text_fra` | `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) | 1 | draft ↗ | |||||
| Norwegian NynorskLIA | only spoken | 5250 | 55410 | n/a - spoken data only | `dialect`: OK | 0 | draft ↗ | ||||||
| Persian Seraji | mixed | 5997 | 151627 | no | `text_en` → `text_eng` | 2 | draft ↗ | ||||||
| Pesh ChibErgIS | only spoken | 524 | 4275 | n/a - spoken data only | derive from `sent_id` prefix identifying the source recording | `text_en` → `text_eng` | `text_phrase-gls-es` → `text_spa`; `text_phrase-gls-it` → `text_phonetic`; `text_phrase-gls-pro` → `text_prosodic`; `text_phrase-gls-tl`: corpus-specific - verify against metadata.html; `text_phrase-gls-de`:… | `sent_timecode` → split (`sound_alignment_begin`, `sound_alignment_end`, and `duration`) | 0 | draft ↗ | |||
| Polish LFG | mixed | 17246 | 130967 | yes | via the sentence-level `genre` field | 3 | draft ↗ | ||||||
| Scottish_Gaelic ARCOSG | mixed | 4748 | 86139 | yes | via the letter prefix of `document_id` | `speaker` → `speaker_id` | 0 | draft ↗ | |||||
| Skolt_Sami Giellagas | mixed | 261 | 3049 | yes | per README description | derive from `sent_id` source-identifier prefix | `text_fi` → `text_fin`; `text_en` → `text_eng` | 3 | draft ↗ | ||||
| Slovenian SST | only spoken | 6121 | 98393 | n/a - spoken data only | 1 | draft ↗ | |||||||
| Soi AHA | mixed | 8 | 55 | yes | per README description | `text_en` → `text_eng`; `text_fa` → `text_fas` | declared type/genre may not match actual content - see modality note | 0 | draft ↗ | ||||
| South_Levantine_Arabic MADAR | mixed | 100 | 789 | no | declared type/genre may not match actual content - see modality note | 1 | draft ↗ | ||||||
| Spanish COSER | only spoken | 539 | 7987 | n/a - spoken data only | `turn_time` → split (`sound_alignment_begin` and `sound_alignment_end`; derive `duration`); `time` → split (`sound_alignment_begin` and `sound_alignment_end` (same as…); `turn_time` → split (`sound_alignment_begin` and… | 1 | draft ↗ | ||||||
| Swedish_Sign_Language SSLC | only spoken | 203 | 1610 | n/a - spoken data only | derive from `sent_id` prefix (everything before the first `:`) | 0 | draft ↗ | ||||||
| Telugu_English TECT | only spoken | 97 | 456 | no | declared type/genre may not match actual content - see modality note | 1 | draft ↗ | ||||||
| Turkish_English BUTR | only spoken | 58 | 441 | no | `text_en` → `text_eng` | declared type/genre may not match actual content - see modality note | 2 | draft ↗ | |||||
| Turkish_German SAGT | only spoken | 2184 | 36934 | n/a - spoken data only | derive from `sent_id` prefix (everything before the trailing `-<number>`) | 2 | draft ↗ | ||||||
| Ukrainian ParlaMint | mixed | 7142 | 109166 | yes | per README description | derive `# document_id = NSDC_UA_28_Feb2014` for the NSDC-sourced sentences (`sent_id`… | `text_en`: remove (single occurrence, placeholder value `undefined undefined`) | 1 | draft ↗ | ||||
| Vietnamese TueCL | only spoken | 100 | 1888 | n/a - spoken data only | 0 | — | |||||||
| Western_Armenian ArmTDP | mixed | 6644 | 121432 | yes | via the `document_id` prefix, which already encodes genre | `doc_title`: `document_id` already exists separately (e.g. `spoken-002R`);… | 1 | draft ↗ | |||||
| Western_Sierra_Puebla_Nahuatl MesoTree | mixed | 3024 | 19191 | no | `text[spa]` → `text_spa`; `text[eng]` → `text_eng` | `text[orig]` → `text_original`; `text[morf]` → `text_morphemic`; `text[gloss]` → `text_glossing`; `text[orig_omitlan]`: corpus-specific (sentence-level) - verify against metadata.html; `text[orig_smt]`: corpus-specific… | `user_id`: corpus-specific (speaker/paragraph-level) - verify against…; `finished`: corpus-specific (speaker/paragraph-level) - verify against…; `location`: corpus-specific (speaker/paragraph-level) - verify against…;… | `timestamp` → `sound_alignment_begin` | 13 | draft ↗ | |||
| Yiddish YiTB | mixed | 3113 | 27954 | yes | via the sentence-level `genre` field | `text_en` → `text_eng` | `rtl`: corpus-specific (speaker/paragraph-level) - verify against…; `source`: corpus-specific (speaker/paragraph-level) - verify against… | 3 | draft ↗ | ||||
| Zazaki ZSD | only spoken | 200 | 1371 | n/a - spoken data only | add `# document_id = Seyristane_dialogue` corpus-wide (single document) | `text_en` → `text_eng` | 0 | draft ↗ |
Harmonization Pipeline:
- is spoken part clearly identifiable?
yes> add# modality = spokento relevant sentencesno> open ISSUE with text XXXX
Workflow
- Per-treebank data (all reviewed fields, per-field advice, and a manual-check list) lives in
workgroups/spoken-data/treebanks/<Treebank>.md- one file per treebank. - This index is a generated recap; edit the per-treebank files, not this table, then regenerate.
- Per-treebank issue drafts summarizing needed changes are available in
workgroups/spoken-data/issue_drafts/and linked from the “issue draft” column above. - I hope I didn’t get anything wrong. -L