home edit page issue tracker

This page pertains to UD version 2.

Metadata harmonisation: align spoken-data fields with UniDive naming conventions

Back to Norwegian NynorskLIA · Back to index

Repo: https://github.com/UniversalDependencies/UD_Norwegian-NynorskLIA

Cross-posting from the UniDive WG1 T1.5 (spoken language guidelines) metadata harmonisation review. We compared UD_Norwegian-NynorskLIA’s current CoNLL-U metadata against the proposed naming conventions (see also the full treebank status table). This is a suggestion for maintainers to review - feel free to push back on anything that doesn’t fit the corpus. The comparison was carried out semi-automatically with the help of Claude (Anthropic); errors or misunderstandings are possible, so please double-check anything unclear.

1. Speaker-level (naming conventions)

Field Suggestion
dialect OK as corpus-specific field, but speaker_id is currently embedded within it - please split speaker_id out into its own field

Implementation notes

Needs a small script


This issue was prepared as part of the UniDive WG1 T1.5 spoken language guidelines effort. Happy to help implement these changes ourselves if that’s easier than doing it on your end - just let us know.