| October 1, 2022
Project overview
Designing a reward function for automated trading is difficult: a strategy must balance returns and risk, while financial observations are noisy and limited. Inverse reinforcement learning (IRL) offers a way to infer an investor’s objectives from observed decisions. The inferred reward can then guide a reinforcement learning agent when it selects and rebalances a portfolio.
This project, carried out by a team from the MMF specialisation, studies how this approach can be applied to portfolio allocation. It begins with a review of IRL and related learning methods [1-4], then focuses on the combination of inverse and direct reinforcement learning proposed for asset allocation [5,6]. Rather than relying on private portfolio manager records, the students aim to generate demonstration trajectories using classical allocation strategies, including mean-variance optimisation.
The objective is to implement and compare GIRL and T-REX for reward inference, and G-Learner for the subsequent portfolio decisions. Historical end-of-day prices provide a basis for backtesting across different assets. Comparisons with conventional portfolio optimisation, including covariance matrix filtering, help assess sensitivity to risk estimates and the choice of demonstrations. The study also considers adapting Upside-Down Reinforcement Learning [1] to financial data, and exploring trading agents in an ABIDES limit order book simulation. The expected deliverables are a literature review, reproducible implementations and benchmarks, and a scientific report documenting the methods and their limitations.
References
[1] Jürgen Schmidhuber. “Reinforcement Learning Upside Down: Don’t Predict Rewards – Just Map Them to Actions.” arXiv:1912.02875 (2020).
[2] Yiqing Xu, Wei Gao and David Hsu. “Receding Horizon Inverse Reinforcement Learning.” NeurIPS 35 (2022).
[3] Gerogiannis. “Inverse Reinforcement Learning.” (2022).
[4] Lili Chen et al. “Decision Transformer: Reinforcement Learning via Sequence Modeling.” arXiv:2106.01345 (2021).
[5] Igor Halperin, Jiayu Liu and Xiao Zhang. “Combining Reinforcement Learning and Inverse Reinforcement Learning for Asset Allocation Recommendations.” arXiv:2201.01874 (2022).
[6] Matthew Dixon and Igor Halperin. “G-Learner and GIRL: Goal Based Wealth Management with Reinforcement Learning.” arXiv:2002.10990 (2020).