South_Levantine_Arabic MADAR
Overview
| type | mixed |
| available since | 2.7 |
| link | https://github.com/UniversalDependencies/UD_South_Levantine_Arabic-MADAR |
| genre | spoken social |
| contributors | Zahra, Shorouq |
| sentences | 100 |
| tokens | 789 |
Issue draft: UD_South_Levantine_Arabic-MADAR
Modality identification
Is spoken part clearly identifiable? No - despite genre listing spoken, the README indicates the 100 sentences are written: they’re manually translated (not transcribed) short conversational tourism-related texts from the MADAR Parallel Corpus, itself derived from the written Basic Traveling Expression Corpus (BTEC). There’s no indication any sentence is a transcription of actual speech - flagged to maintainers to confirm or drop spoken from genre.