Skip to main content
Thesis defences

PhD Oral Exam - Xingnan Zhou, Civil Engineering

Multi-Scale Deep Learning Frameworks for Trajectory Prediction, Scene Understanding, and Perception in Autonomous Driving


Date & time
Tuesday, August 18, 2026
10 a.m. – 1 p.m.
Format

In-person

Cost

This event is free

Organization

School of Graduate Studies

Contact

Dolly Grewal

Where

Engineering, Computer Science and Visual Arts Integrated Complex
1515 Ste-Catherine St. W.
Room 003.309

Accessible location

Yes - See details

When studying for a doctoral degree (PhD), candidates submit a thesis that provides a critical review of the current state of knowledge of the thesis subject as well as the student’s own contributions to the subject. The distinguishing criterion of doctoral graduate research is a significant and original contribution to knowledge.

Once accepted, the candidate presents the thesis orally. This oral exam is open to the public.

Abstract

Reliable autonomous vehicle operation at urban intersections requires competence across the prediction–perception stack, yet current deep learning models do not fully exploit the structural priors that govern real-world driving. This thesis argues that incorporating structural priors at multiple scales improves the accuracy, interpretability, and robustness of autonomous driving models, substantiated through five interconnected studies. Chapter 3 introduces a Turn-Aware LSTM that encodes turning intent via cumulative heading change. On UAV footage from a signalised intersection in Châteauguay, Québec, it reduces Final Displacement Error by 15–20% for turning manoeuvres at a 3-second horizon. Chapter 4 establishes lane graph conditioning as an architecture-agnostic inductive bias for trajectory prediction. A lightweight module (under 50,000 parameters), evaluated on 89,258 Waymo scenarios, reduces ADE by 9.3% at 3 s and, at 8 s, minADE by 26.7%, minFDE by 32.6%, and miss rate by 42.7%. Chapter 5 recovers real-world locations for anonymised scenarios by matching intersection topologies to OpenStreetMap via WayGraph's 48-dimensional star pattern descriptor. Across 56,797 scenarios it achieves 90% top-1 accuracy under synthetic OSM-to-OSM matching (33–55% under realistic noise), and a geographic audit reveals fourfold under-sampling of signalised intersections (4.3% vs. 17.8% city-wide). Chapter 6 develops a spatial attention visualisation framework for Transformer-based prediction, projecting attention onto bird’s-eye-view scenes. Applied to MTR-Lite, it identifies a tunnel vision pattern: failed predictions are associated with lower attention entropy (5.72 vs. 5.94 bits) and elevated self-attention, with cyclist targets showing higher miss rates (88.1% vs. 54.0%) in illustrative case studies. Left-turn manoeuvres trigger a general agent-to-map attention shift (−19.9% vehicle attention, p < 0.001), suggesting possible VRU-related attention vulnerabilities, though the cyclist-specific interaction did not reach significance (p = 0.51, n = 17) and requires validation on larger samples. Chapter 7 shows that viewpoint diversity suppresses false positives via dual-camera LiDAR fusion with vehicle- and drone-mounted cameras. Across 2,600 CARLA frames and ten seeds, symmetric late fusion improves mAP by +0.92 percentage points (p = 0.001) and reduces false positives by 13% on a representative seed. These studies support the thesis argument across five scales—kinematic, road topology, geographic, attentional, and sensor geometry—and outline integrated future research directions. Keywords: autonomous driving, trajectory prediction, lane graph conditioning, GPS-free localisation, attention visualisation, sensor fusion, intersection safety, traffic safety, vulnerable road users

Back to top

© Concordia University