Date & time
1 p.m. – 4 p.m.
This event is free
School of Graduate Studies
Engineering, Computer Science and Visual Arts Integrated Complex
1515 Ste-Catherine St. W.
Room 1.162
Yes - See details
When studying for a doctoral degree (PhD), candidates submit a thesis that provides a critical review of the current state of knowledge of the thesis subject as well as the student’s own contributions to the subject. The distinguishing criterion of doctoral graduate research is a significant and original contribution to knowledge.
Once accepted, the candidate presents the thesis orally. This oral exam is open to the public.
Label-Efficient 3D Point Cloud Understanding: From Representation Learning to Foundation Model Adaptation Point clouds are the native output of modern 3D sensors that power autonomous driving, robotics, and augmented reality, yet their irregular, permutation-invariant, and geometrically rich structure-together with the high cost of 3D annotation-makes fully supervised learning impractical. In this thesis, we propose novel methods for label-efficient 3D perception that span from learning representations directly on raw point clouds to adapting pre-trained 3D foundation models under limited supervision and distribution shift. We first address representation learning by proposing CrossMoCo, a cross-modal momentum-contrastive framework that jointly pre-trains point cloud and image encoders to learn transferable representations. We further propose AllMatch that targets the limitation of pseudo-labelling in semi-supervised learning, so that every unlabelled sample-not only high-confidence ones-contributes, recovering near fully supervised accuracy with a small fraction of the labels. We next address adaptation by proposing MCFT, an adapter-free fine-tuning paradigm that curbs overfitting at no additional inference cost. We further propose ReFine3D that extends this idea to multi-modal 3D vision-language models, adding multi-view consistency and text-diversity regularization for robustness to novel classes, data corruptions, and cross-dataset shift. Finally, we propose SAGE by introducing an end-to-end, encoder-free 3D multi-modal large language model that resolves semantic misalignment and resolution mismatch challenges by treating point clouds as a foreign language. Together, these contributions advance robust, scalable 3D understanding toward practical, annotation-light deployment.
© Concordia University