Decision trees with complex conditions for trading

| October 1, 2024

Project overview

The Lusis Chair on AI and Finance focuses on two main research areas:

  1. Fraud detection in credit card payments
  2. Algorithmic trading

Both fields require scalable, interpretable, and high-speed classification and prediction methods, with a particular emphasis on large-scale fraud detection. While deep learning offers advantages in terms of model complexity and expressive power, anomaly detection on tabular data remains challenging.

Decision trees are widely used for tabular data classification, typically relying on simple node conditions, where a variable is compared to a threshold. This project explores enhanced decision trees with more complex conditions, such as:

  • Oblique conditions, which involve linear combinations of multiple features rather than single-variable comparisons.
  • Multi-variable conditions, where several features are evaluated simultaneously against multiple threshold values, allowing for richer and more flexible decision rules.

The objective is to design and implement a decision tree model capable of handling multi-variable conditions and to compare its performance with standard single-variable and oblique-condition trees. The study can use fraud detection or algorithmic trading datasets to assess the effectiveness of these enhanced decision trees in a financial setting.

The supplied reading covers conditional inference trees [1, 2], oblique splits [3], fuzzy rule based decision trees [4], and modular rule induction with PRISM, including its parallel variant [5–7]. Development uses Python/PyTorch on Linux, with access to CPU, GPU and large-memory computing resources; distributed computation may be needed.

Deliverables include:

  • Development and testing of a decision tree model incorporating multi-variable conditions.
  • Comparative analysis with conventional and oblique decision trees.
  • Application to fraud detection or trading datasets to evaluate effectiveness.
  • A final report summarizing the modeling process, experimental results, and performance evaluation.

References

[1] Hothorn, T., Hornik, K., & Zeileis, A. ctree: Conditional Inference Trees. partykit vignette. PDF.

[2] Schweinberger, M. An introduction to conditional inference trees in R: Basics of tree-based models, section “Problems”. Course notes.

[3] Decision trees vs Oblique decision trees. Data Science Stack Exchange. Discussion.

[4] Wang, X., Liu, X., Pedrycz, W., & Zhang, L. (2015). Fuzzy rule based decision trees. Pattern Recognition, 48(1), 50–59. https://doi.org/10.1016/j.patcog.2014.08.001

[5] Cendrowska, J. (1987). PRISM: An algorithm for inducing modular rules. International Journal of Man-Machine Studies, 27(4), 349–370. https://doi.org/10.1016/S0020-7373(87)80003-2

[6] Kennedy, W. B. (2024). PRISM-Rules in Python: A simple python rules-induction system. Towards Data Science. Article.

[7] Stahl, F., Adda, M., & Bramer, M. (2008). P-Prism: A computationally efficient approach to scaling up classification rule induction. Artificial Intelligence in Theory and Practice II, 276, 77–86. https://doi.org/10.1007/978-0-387-09695-7_8