Date & time
9 a.m. – 12 p.m.
In-person
This event is free
Concordia University, School of Graduate Studies
Engineering, Computer Science and Visual Arts Integrated Complex
1515 Ste-Catherine St. W.
Room 3.309
Yes - See details
When studying for a doctoral degree (PhD), candidates submit a thesis that provides a critical review of the current state of knowledge of the thesis subject as well as the student’s own contributions to the subject. The distinguishing criterion of doctoral graduate research is a significant and original contribution to knowledge.
Once accepted, the candidate presents the thesis orally. This oral exam is open to the public.
Abstract A Unified Bayesian Framework for Unsupervised Learning: From Mixture Model Algorithms to Large Language Model Personalization Ornela Bregu, Ph.D. Concordia University, 2026 This thesis presents a unified Bayesian framework for unsupervised learning of highdimensional count data, with applications ranging from topic modeling and image recognition to large language model personalization. The central model is the Dirichlet Compound Negative Multinomial (DCNM) distribution, which captures overdispersion, burstiness, and positive feature correlations that simpler models such as the multinomial distribution cannot handle. Using a rising-factorial parametrization of the DCNM, we derive minorization-maximization equations for finite mixture models that guarantee monotone likelihood ascent, are analytically tractable without extensive Hessian computations, and are robust to initialization. Model selection is addressed via the minimum message length criterion, which simultaneously determines the number of mixture components and the relevant feature set. We extend the framework in two directions. First, we replace the Dirichlet prior with the more flexible generalized Dirichlet, yielding the generalized Dirichlet compound negative multinomial mixture model, which captures both positive and negative feature correlations within a conjugate, analytically tractable hierarchy. Second, we introduce feature saliency weights into the DCNM mixture (DCNM-FS) and develop an online learning algorithm that updates parameters and saliency weights incrementally from streaming data, enabling scalable topic discovery without retraining. Finally, the learned DCNM-FS mixture is embedded as a Bayesian prior in an instruction-tuned large language model personalisation framework. User preferences are modelled as a finite mixture of latent types. Then, an online Bayesian posterior over types is updated after each observed interaction and injected into the large language model context as a structured preference profile.
© Concordia University