Back to Search Start Over

Generating Commit Messages from Git Diffs

Authors :
van Hal, S. R. P.
Post, M.
Wendel, K.
Publication Year :
2019

Abstract

Commit messages aid developers in their understanding of a continuously evolving codebase. However, developers not always document code changes properly. Automatically generating commit messages would relieve this burden on developers. Recently, a number of different works have demonstrated the feasibility of using methods from neural machine translation to generate commit messages. This work aims to reproduce a prominent research paper in this field, as well as attempt to improve upon their results by proposing a novel preprocessing technique. A reproduction of the reference neural machine translation model was able to achieve slightly better results on the same dataset. When applying more rigorous preprocessing, however, the performance dropped significantly. This demonstrates the inherent shortcoming of current commit message generation models, which perform well by memorizing certain constructs. Future research directions might include improving diff embeddings and focusing on specific groups of commits.

Details

Database :
arXiv
Publication Type :
Report
Accession number :
edsarx.1911.11690
Document Type :
Working Paper