| October 1, 2025
Project overview
Large language models (LLMs) can achieve strong predictive performance while remaining overconfident when they make mistakes. Accuracy and F1 scores alone do not establish whether a model’s confidence is trustworthy. This MMF project investigates uncertainty quantification for transformer models, initially focusing on classification of tabular data, where reliable confidence estimates remain an open research question.
The initial study brings together two starting points: a benchmark of LLM uncertainty across language tasks [1], and TabLLM, which converts table rows into natural-language descriptions for zero-shot and few-shot classification [2]. The objective is to assess how uncertainty methods can accompany such predictions, rather than evaluating classification performance alone. The distinction between uncertainty caused by noisy data and uncertainty caused by limited model knowledge provides a theoretical foundation [3].
The work includes reviewing conformal prediction, particularly adaptive prediction sets, and confidence indicators such as maximum softmax probability, entropy and perplexity. These approaches provide baselines for an investigation of Monte Carlo dropout [4]: repeated stochastic forward passes produce a distribution of predictions whose variability can be measured. Students are expected to implement an evaluation pipeline, compare uncertainty estimates across several datasets, and examine whether those estimates help identify unreliable predictions. The methodological reading also considers repeated-answer consistency [5] and efficient ensembles of attention models [6].
The chair’s broader trading brief motivates potential applications to LLM-based sentiment analysis and transformer-based time-series prediction. It proposes evaluating whether uncertainty estimates can improve strategy performance; the initial experiments on tabular classification establish a more general methodological basis for that application. Deliverables include reproducible Python/PyTorch experiments, documented notebooks, and a report comparing the approaches and their limitations.
References
[1] Ye, F., Yang, M., Pang, J., Wang, L., Wong, D. F., Yilmaz, E., Shi, S., and Tu, Z. (2024). Benchmarking LLMs via Uncertainty Quantification. NeurIPS 2024, Datasets and Benchmarks Track.
[2] Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X., and Sontag, D. (2023). TabLLM: Few-shot Classification of Tabular Data with Large Language Models. AISTATS, PMLR 206, 5549-5581.
[3] Hüllermeier, E., and Waegeman, W. (2021). Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods. Machine Learning, 110, 457-506.
[4] Gal, Y., and Ghahramani, Z. (2016). Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning. ICML, PMLR 48, 1050-1059.
[5] Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D. (2023). Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR 2023.
[6] Mühlematter, D. J., et al. (2025). LoRA-Ensemble: Efficient Uncertainty Modelling for Self-Attention Networks. arXiv:2405.14438, version 4.
[7] Zhang et al. (2023). Prompting Ensembles for Robustness and Uncertainty Estimation.