Large language models and news sentiment for algorithmic trading

| October 1, 2025

Project overview

Algorithmic trading involves sequential decisions in an uncertain environment. Alongside prices and technical indicators, financial news, company reports and market commentary may contain information relevant to these decisions. This project investigates how large language models (LLMs) can extract that information and whether it can improve trading signals for specific financial instruments.

The initial brief identifies TradExpert [1], a system combining specialist LLMs, as a starting point, and points students towards the financial AI workshop at ICLR 2025 [2]. The study focuses on news sentiment for Apple stock (AAPL), drawing on research into media sentiment [3, 4], financial language models [5, 6, 7] and event-based stock prediction [8]. Rather than representing each article with a single positive or negative score, the objective is to examine richer descriptions, including sentiment intensity, uncertainty, confidence, expected market impact and impact horizon.

The work includes comparing language models, checking the consistency of their structured outputs, and aggregating article-level information while accounting for publication time, article age and importance. The resulting features are studied alongside market variables, distinguishing overnight and intraday horizons. Predictive relationships must then be assessed separately from their practical use in a trading strategy.

Evaluation compares sentiment-based strategies with market-only models and a buy-and-hold baseline, using returns, Sharpe and Sortino ratios, and maximum drawdown. Generalization and backtest overfitting [9, 10] are central concerns, together with the timing of available information and transaction costs. Combining strategies can also be explored through portfolio allocation [11]. The expected deliverables are a Python/PyTorch model and experimental pipeline quantifying the contribution of LLMs, and a report written as a scientific article.

References

[1] Ding, Q., Shi, H., Guo, J., and Liu, B. (2025). TradExpert: Revolutionizing Trading with Mixture of Expert LLMs. arXiv:2411.00782, revised version.

[2] Advances in Financial AI Workshop at ICLR 2025: Accepted papers.

[3] Tetlock, P. C. (2007). Giving Content to Investor Sentiment: The Role of Media in the Stock Market. The Journal of Finance, 62(3), 1139-1168.

[4] Boudoukh, J., Feldman, R., Kogan, S., and Richardson, M. (2019). Information, Trading, and Volatility: Evidence from Firm-Specific News. The Review of Financial Studies, 32(3), 992-1033.

[5] Araci, D. (2019). FinBERT: Financial Sentiment Analysis with Pre-trained Language Models. arXiv:1908.10063.

[6] Lopez-Lira, A., and Tang, Y. (2023). Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models. SSRN.

[7] Yang, H., Liu, X.-Y., and Wang, C. D. (2023). FinGPT: Open-Source Financial Large Language Models. arXiv:2306.06031.

[8] Ding, X., Zhang, Y., Liu, T., and Duan, J. (2015). Deep Learning for Event-Driven Stock Prediction. Proceedings of the 24th International Joint Conference on Artificial Intelligence (IJCAI).

[9] Gu, S., Kelly, B., and Xiu, D. (2020). Empirical Asset Pricing via Machine Learning. The Review of Financial Studies, 33(5), 2223-2273.

[10] Bailey, D. H., Borwein, J. M., López de Prado, M., and Zhu, Q. J. (2017). The Probability of Backtest Overfitting. Journal of Computational Finance, 20(4), 39-69.

[11] Markowitz, H. (1952). Portfolio Selection. The Journal of Finance, 7(1), 77-91.

[12] Wu, S., et al. (2024). Large Language Models in Finance: A Survey.