Is spoken part clearly identifiable? N/A - spoken data only
Metadata review
languages and translation(s)
Field
Advice
text_en
change to text_eng
transcription and annotation levels available
Field
Advice
text_ortho
change to text_orthographic
morphemic_text
change to text_morphemic
speaker metadata
Field
Advice
speaker_id
OK
doc (and paragraphs) metadata
Field
Advice
—
no document_id exists at all (0 occurrences across 406 sentences) - but the 54 distinct sound_url values (e.g. SAB-TXT-AN-00000-01.WAV) already identify document boundaries; derive # document_id from the recording basename and set it once per document
sound_url
currently repeated on every sentence - move to document level once document_id exists
sent metadata
Field
Advice
sent_timecode
change to sound_alignment_begin, sound_alignment_end and duration