home edit page issue tracker

This page pertains to UD version 2.

Metadata harmonisation: align spoken-data fields with UniDive naming conventions

Back to Ligurian GLT · Back to index

Repo: https://github.com/UniversalDependencies/UD_Ligurian-GLT

Cross-posting from the UniDive WG1 T1.5 (spoken language guidelines) metadata harmonisation review. We compared UD_Ligurian-GLT’s current CoNLL-U metadata against the proposed naming conventions (see also the full treebank status table). This is a suggestion for maintainers to review - feel free to push back on anything that doesn’t fit the corpus. The comparison was carried out semi-automatically with the help of Claude (Anthropic); errors or misunderstandings are possible, so please double-check anything unclear.

1. Is the spoken portion identifiable?

This treebank mixes spoken and written material but its .conllu files don’t explicitly mark which sentences are spoken. We looked for a pattern in the data (a weak signal):

Finding: Only 12 documents total, with mixed short prefixes - too few for an automatic pattern, but small enough to check by hand. The README mentions a radio broadcast among the sources, but we couldn’t work out which document_id(s) or sentences it corresponds to.

Evidence: document_id values: bdl-c00, cairo, esl-c01, wp-arba, wp-tintin, and others (12 total, no dominant separator/prefix pattern).

Suggestion: Since there are only 12 documents, could you confirm per-document whether each is spoken or written (e.g. do wp-* mean Wikipedia/written, is cairo a transcribed story), and specifically which document(s) correspond to the radio broadcast mentioned in the README?

2. Sentence-level (naming conventions)

Field Suggestion
parallel_id corpus-specific (sentence-level) - verify against metadata.html

Implementation notes

Needs manual input from maintainers