Abaza ATB
Overview
| type | only spoken |
| available since | 2.11 |
| link | https://github.com/UniversalDependencies/UD_Abaza-ATB |
| genre | spoken |
| contact | alexeykochevoy@gmail.com |
| sentences | 98 |
| tokens | 652 |
Issue draft: UD_Abaza-ATB
Modality identification
Is spoken part clearly identifiable? N/A
Metadata review
Every sentence carries exactly the same six comment fields, plus one MISC feature on (nearly) every token:
| Field | Level | Example | Standard name (metadata.md) | Advice |
|---|---|---|---|---|
text_orth |
sentence | # text_orth = сарА сы-хьиз фатима-пI |
this seems a morpheme-segmented orthographic form (hyphens mark morpheme boundaries, stress marked) | rename to text_morphemic |
text_transcription |
sentence | # text_transcription = sará sə-χ'iz fatima-ṗ |
Latin-script rendering of the Cyrillic orthographic form (not a phonetic/IPA transcription) | rename to text_transliteration |
text_rus |
sentence | # text_rus = Меня зовут Фатима. |
translation field; project convention uses 3-letter ISO codes | OK |
text_name |
sentence (but constant across all sentences of one recording — 6 distinct values over 98 sentences) | # text_name = Professija_AjsanovaFB_11072017_checked.eaf |
this seems a document identifier, wrongly repeated per-sentence instead of set once per document | convert to # newdoc id = ... at the first sentence of each of the 6 documents (drop the per-sentence repetition and the .eaf extension) |
Not present, worth considering
- Speaker metadata: none at all (
speaker_id,speaker_role, etc. absent). The 6text_namefilenames encode a speaker per recording (e.g.AjsanovaFB,SanashokovaCKh,DzhuzhuevKM,BidzhevaTA,AsanaevaFM— initials suggest one speaker/informant per file). Oncetext_nameis converted tonewdoc id, aspeaker_idcould plausibly be derived from the same filename component. - Genre: no
# genre = ...field, even though topics are recoverable from filenames. These read as personal narrative/interview elicitations — could add# genre = narrativeorinterviewper document. sound_url: the corpus homepage (lingconlab.ru/spoken_abaza) implies underlying audio recordings exist; no link is included in the.conllu. Worth asking whether individual recording URLs can be shared.