Skip to main content
Thesis defences

PhD Oral Exam - Narjes Tahaei, Computer Science

Annotation-Aware Language Models as Representations of Disagreement and Perspective


Date & time
Tuesday, August 25, 2026
10 a.m. – 1 p.m.
Cost

This event is free

Organization

School of Graduate Studies

Contact

Dolly Grewal

Where

ER Building
2155 Guy St.
Room 1222

Accessible location

Yes - See details

When studying for a doctoral degree (PhD), candidates submit a thesis that provides a critical review of the current state of knowledge of the thesis subject as well as the student’s own contributions to the subject. The distinguishing criterion of doctoral graduate research is a significant and original contribution to knowledge.

Once accepted, the candidate presents the thesis orally. This oral exam is open to the public.

Abstract

Subjective natural language processing (NLP) tasks, such as sexism, irony, hate speech, and offensive language detection, often involve multiple valid interpretations, leading to substantial disagreement among annotators. Conventional approaches typically collapse these annotations into a single gold label, discarding useful information about the diversity of human perspectives. This thesis investigates how annotation disagreement can be modeled and evaluated in a way that better reflects attitudes towards subjective tasks.

The thesis first examines whether annotator labels and meta-information can improve classification. It introduces annotation-aware methods that condition predictions on demographic and other annotation-level features, including an attention-based architecture that produces different outputs for different meta-information bundles. A systematic analysis across all feasible demographic feature combinations shows that performance is not uniformly improved by adding more metadata. Instead, the usefulness of annotation information depends on dataset characteristics, feature balance, and soft-label variance.

The thesis then turns to evaluation. It shows that hard-label metrics such as F1 often hide the effects of disagreement-aware training, and that soft-label and calibration metrics provide a more faithful assessment of model behavior under ambiguity. Together, these results support a target-aware view of subjective classification tasks in which disagreement is treated as meaningful signal rather than noise, and evaluation is aligned with the intended annotation distribution.

Back to top

© Concordia University