Discovering financial factors with language model agents

October 1, 2026

Project overview

Can a language model agent discover useful financial factors more effectively than random or evolutionary search? Inspired by AQuA [1], this project studies whether remembering previous experiments improves the search process. A factor is a mathematical expression that transforms historical market observations into a score for ranking assets by their expected future returns.

The study uses twelve months of public Binance data for a fixed universe of forty liquid cryptocurrency perpetual contracts, sampled every five minutes, with a one-hour prediction horizon. Students build a restricted expression language based on the operators of 101 Formulaic Alphas [2], drawing on Qlib [3] for vectorized evaluation. An isolated evaluator checks formulas and measures predictive quality; the agent cannot change the data, labels or evaluation rules. Candidate factors are combined while limiting redundant signals, following the motivation for synergistic factor collections [4].

Four approaches share the same expression space and budget of distinct evaluated formulas: random search, genetic search, a grammar-constrained local LLM with experimental memory, and the same agent without informative memory. The agent proposes an economic hypothesis, translates it into a formula and receives evaluation feedback. This comparison isolates the contribution of the LLM and its memory.

Chronological data splits, non-anticipation checks, three independent runs per method, block bootstrap [6, 7] and multiple-testing controls [5, 8] support the comparison. Factor pools are frozen before the final test period is opened. The main criterion is the out-of-sample information coefficient: the Spearman correlation between asset scores and subsequent returns. Computational expenditure and exploration diversity are also recorded. Transaction costs and portfolio construction are outside this study’s scope, so predictive results do not establish trading profitability. Deliverables include a reproducible implementation, experiment logs and a comparative report; finding no advantage for the LLM is a valid outcome.

References

[1] Guo, J., et al. (2026). AQuA: Recursively Self-Improving Quantitative Trading Research Agents. arXiv:2608.12841.

[2] Kakushadze, Z. (2016). 101 Formulaic Alphas. Wilmott Magazine, 2016(84), 72–80. arXiv:1601.00991.

[3] Yang, X., et al. (2020). Qlib: An AI-oriented Quantitative Investment Platform. arXiv:2009.11189.

[4] Yu, S., et al. (2023). Generating Synergistic Formulaic Alpha Collections via Reinforcement Learning. Proceedings of KDD 2023, 5476–5486. DOI: 10.1145/3580305.3599831.

[5] Benjamini, Y., and Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological), 57(1), 289–300.

[6] Hansen, L. P., and Hodrick, R. J. (1980). Forward Exchange Rates as Optimal Predictors of Future Spot Rates: An Econometric Analysis. Journal of Political Economy, 88(5), 829–853.

[7] Politis, D. N., and Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association, 89(428), 1303–1313.

[8] Harvey, C. R., Liu, Y., and Zhu, H. (2016). … and the Cross-Section of Expected Returns. The Review of Financial Studies, 29(1), 5–68.