Date & time
10 a.m. – 1 p.m.
This event is free
School of Graduate Studies
ER Building
2155 Guy St.
Room 1222
Yes - See details
When studying for a doctoral degree (PhD), candidates submit a thesis that provides a critical review of the current state of knowledge of the thesis subject as well as the student’s own contributions to the subject. The distinguishing criterion of doctoral graduate research is a significant and original contribution to knowledge.
Once accepted, the candidate presents the thesis orally. This oral exam is open to the public.
Subjective natural language processing (NLP) tasks, such as sexism, irony, hate speech, and offensive language detection, often involve multiple valid interpretations, leading to substantial disagreement among annotators. Conventional approaches typically collapse these annotations into a single gold label, discarding useful information about the diversity of human perspectives. This thesis investigates how annotation disagreement can be modeled and evaluated in a way that better reflects attitudes towards subjective tasks.
The thesis first examines whether annotator labels and meta-information can improve classification. It introduces annotation-aware methods that condition predictions on demographic and other annotation-level features, including an attention-based architecture that produces different outputs for different meta-information bundles. A systematic analysis across all feasible demographic feature combinations shows that performance is not uniformly improved by adding more metadata. Instead, the usefulness of annotation information depends on dataset characteristics, feature balance, and soft-label variance.
The thesis then turns to evaluation. It shows that hard-label metrics such as F1 often hide the effects of disagreement-aware training, and that soft-label and calibration metrics provide a more faithful assessment of model behavior under ambiguity. Together, these results support a target-aware view of subjective classification tasks in which disagreement is treated as meaningful signal rather than noise, and evaluation is aligned with the intended annotation distribution.
© Concordia University