Skip to main content
Thesis defences

PhD Oral Exam - Criscent Birungi, Mathematics and Statistics

Essays on Stochastic Control and Reinforcement Learning for Lifecycle Retirement and Annuitization


Date & time
Wednesday, August 26, 2026
9 a.m. – 12 p.m.
Cost

This event is free

Organization

School of Graduate Studies

Contact

Dolly Grewal

Where

J.W. McConnell Building
1400 De Maisonneuve Blvd. W.
Room 921-4

Accessible location

Yes - See details

When studying for a doctoral degree (PhD), candidates submit a thesis that provides a critical review of the current state of knowledge of the thesis subject as well as the student’s own contributions to the subject. The distinguishing criterion of doctoral graduate research is a significant and original contribution to knowledge.

Once accepted, the candidate presents the thesis orally. This oral exam is open to the public.

Abstract

The decision to convert accumulated wealth into lifetime income has become increasingly complex due to rising longevity risk, heterogeneous health trajectories, and rising labor-force participation at older ages. This thesis applies stochastic control, optimal stopping, and continuous-time reinforce-ment learning to a sequence of life-cycle problems in which consumption, labor supply, portfolio choice, and the timing of irreversible annuitization are determined jointly under age-dependent or health-dependent mortality. It comprises four connected papers. 

The first paper studies optimal annuitization with labor income under both the deterministic Gom-pertz law and a stochastic affine mortality process, deriving semi-analytical policies governed by endogenous wealth thresholds for labor and retirement. Because the effective discount rate is the mortality-adjusted rate, both the discounting of future annuity income and the mortality credit rise with age through the same hazard; the results show that beyond a threshold age, rising mortality credits lower the annuitization threshold despite the age-increasing effective discount rate. 

The second paper introduces internal habit formation into the joint control and stopping problem. Its solution yields three behavioral regimes, defensive labor at low wealth relative to the habit, a "work-to-retire" phase, and an abrupt de-risking of the portfolio to the constant Merton fraction at retirement, once the implicit safe asset provided by human capital is given up. Subjective mortality beliefs shift the retirement threshold, offering a potential explanation for the observed reluctance to annuitize. 

The third paper replaces age-based mortality with a stochastic vitality process, in which death occurs when a latent health reserve is first depleted. Higher vitality volatility lowers optimal consumption, with the effect concentrated at high vitality and vanishing as health becomes fragile, and affects the timing of annuitization non-monotonically in volatility. Because vitality risk is modeled as orthogonal to and unhedgeable by market risk, it generates no hedging demand and the optimal risky-asset share is the constant Merton fraction; health uncertainty therefore affects intertemporal allocation rather than financial risk-taking under this independence assumption, which correlated health-market shocks would overturn. 

The fourth paper develops a continuous-time reinforcement-learning framework for the pre-retirement problem and benchmarks it against the homothetic case of the first paper's semi-analytical solu-tion. Formulated as model-based policy iteration with simulation-based policy evaluation, it recov-ers the classical portfolio, consumption, and labor policies exactly. Because policy improvement is analytic, the value-function time profile is the only object learned from data and is identifiable only at intermediate exploration and its estimator degrades as exploration vanishes, exposing an exploration-identifiability tradeoff. 

Together the four papers extend life-cycle annuitization modeling to richer mortality, habit, and la-bor structures, and complement their analytical solutions with a simulation-based method validated against the tractable case and intended for settings where no closed-form solution exists.

Back to top

© Concordia University