JEPA for Code Generation and Optimization

October 1, 2026

Project overview

This 2026–2027 research initiation project for the MMF / SDI tracks, supervised within the Lusis Chair and LISN, investigates whether a predictive world model can guide the production of correct, efficient programs while reducing expensive compiler evaluations. CompilerDream learns the effects of compilation passes [1]; Joint-Embedding Predictive Architectures (JEPA) instead predict representations rather than reconstruct observations [2]. The EB-JEPA library provides examples of action-conditioned models and planning [3]. Adapting these ideas to symbolic program transformations is the central research question.

Students will first study small programs in LLVM intermediate representation. An encoder represents the current program, and an action-conditioned predictor learns how a compilation pass changes that representation. Regularization will address representation collapse. Model-guided search will select transformation sequences, which are then applied and checked by the compiler. A hierarchical H-JEPA variant is an optional extension for transformations at several scales.

To connect optimization with generation, a pretrained code generator will propose short functions from specifications and tests. JEPA will guide candidate selection and optimization; the generator remains responsible for producing program text. Experiments will use public corpora, a single language and executable tasks, with separate training and evaluation programs.

Comparisons will include CompilerDream, standard LLVM configurations and search without a model. Generation will be evaluated with and without JEPA guidance under comparable candidate and compute budgets. Measures include compilation validity, tests passed, execution time, code size, search cost and compiler calls. Unseen programs and longer sequences will probe generalization and accumulated prediction errors. Performance gains remain hypotheses to evaluate; passing tests does not establish general semantic equivalence.

Expected deliverables are a reproducible prototype, a documented code repository and a research article reporting controlled experiments, ablations and limitations.

References

[1] Deng, C., Wu, J., Feng, N., Wang, J., & Long, M. (2025). CompilerDream: Learning a Compiler World Model for General Code Optimization. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, V.2, 486–497. DOI: 10.1145/3711896.3736887.

[2] Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y., & Ballas, N. (2023). Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. CVPR, 15619–15629. Paper.

[3] Terver, B., Balestriero, R., Dervishi, M., Fan, D., Garrido, Q., Nagarajan, T., Sinha, K., Zhang, W., Rabbat, M., LeCun, Y., & Bar, A. (2026). A Lightweight Library for Energy-Based Joint-Embedding Predictive Architectures. arXiv preprint, version 3. arXiv:2602.03604v3.