home edit page issue tracker

This page pertains to UD version 2.

Metadata harmonisation: align spoken-data fields with UniDive naming conventions

Back to Turkish_English BUTR · Back to index

Repo: https://github.com/UniversalDependencies/UD_Turkish_English-BUTR

Cross-posting from the UniDive WG1 T1.5 (spoken language guidelines) metadata harmonisation review. We compared UD_Turkish_English-BUTR’s current CoNLL-U metadata against the proposed naming conventions (see also the full treebank status table). This is a suggestion for maintainers to review - feel free to push back on anything that doesn’t fit the corpus. The comparison was carried out semi-automatically with the help of Claude (Anthropic); errors or misunderstandings are possible, so please double-check anything unclear.

1. Is the spoken portion identifiable?

type says only spoken, but the corpus is actually mixed. The README documents # medium as “Communication medium (Written or Spoken), where known”, and in the data it’s only present on 19 of 58 sentences (Spoken or Written).

Suggestion: Could you confirm modality for the remaining 39 sentences? In the meantime, medium maps directly onto our # modality field.

2. Sentence-level (naming conventions)

Field Suggestion
text_en rename to text_eng (ISO 639-3)
medium rename to # modality (values spoken/written, lowercase); only present on 19/58 sentences - please confirm the rest

Implementation notes


This issue was prepared as part of the UniDive WG1 T1.5 spoken language guidelines effort. Happy to help implement these changes ourselves if that’s easier than doing it on your end - just let us know.