Is spoken part clearly identifiable? N/A - spoken data only
Metadata review
doc (and paragraphs) metadata
Field
Advice
—
no document_id exists at all (0 occurrences across 763 sentences) - but the 18 distinct sound_url values (one per recording, e.g. BEJ_MV_NARR_01_SHELTER.WAV) already identify document boundaries; derive # document_id from the recording basename and set it once per document, then move sound_url there too
languages and translation(s)
Field
Advice
text_en
change to text_eng
transcription and annotation levels available
Field
Advice
phonetic_text
change to text_phonetic
speaker metadata
Field
Advice
speaker_id
OK
sent metadata
Field
Advice
sent_timecode
change to sound_alignment_begin, sound_alignment_end and duration