Is spoken part clearly identifiable? N/A - spoken data only
Metadata review
doc (and paragraphs) metadata
(none found) - no document_id exists at all. But sent_id already encodes it: e.g. ParisStories_2020_maisonAbondonnee_1 is document ParisStories_2020_maisonAbondonnee, sentence 1. 86 distinct documents across 2776 sentences. sound_url (currently repeated per sentence, present on 2749/2776 sentences - 27 sentences in one document lack it) should move to document level once document_id exists.
Field
Advice
—
derive # document_id from the sent_id prefix (everything before the trailing _<number>)
sound_url
move to document level, set once per document_id
transcription and annotation levels available
Field
Advice
macrosyntax
change to text_macrosyntax
tags
corpus-specific (only 1 occurrence, value TODO) - please confirm what this represents