Data Science Wire

Annotation design for a dialect with no standard orthography: what we changed after the first pass failed

Reddit r/LanguageTechnology1w4 min read

Working on annotation for Tunisian Arabic customer service conversations. Derja, arabizi, French, frequently all three inside one message. This comes out of a job, so I am being vague about the source, but the question is a methods question and there is nothing to promote. The first pass at a labelling scheme failed in the way these usually do. Two people could not reliably produce the same labels, and neither could one person a fortnight apart. Posting what we changed, and two things I have not solved, in case anyone here has worked the same problem. Constraints, for anyone who has not worked

Read the full story at Reddit r/LanguageTechnology

More in Machine Learning