No # newdoc id exists, but document boundaries are fully recoverable: the README documents that sent_id encodes <filename>:<sentence_number>, where <filename> matches the text’s name on the source corpus site (chuklang.ru). Splitting sent_id on : gives 65 distinct document prefixes (e.g. Abramovich, GUM, Katyusha) across the 1004 sentences.
Field
Advice
—
derive # newdoc id from the sent_id prefix (everything before :), set once at each document’s first sentence
languages and translation(s)
Field
Advice
text[eng]
change to text_eng
text[eng']
change to text_eng_literal
text[rus]
change to text_rus
transcription and annotation levels available
Field
Advice
text[phon]
change to text_phonetic
sent metadata
Field
Advice
timestamp
change to sound_alignment_begin, sound_alignment_end and duration (only 8 sentences carry this field)