CoUpJava: A Dataset of Code Upgrade Histories in Open-Source Java Repositories (MSR 2025 - Data and Tool Showcase Track)

Who

Kaihang Jiang, Bihui Jin, Pengyu Nie

Track

MSR 2025 Data and Tool Showcase Track

Time Zone

The program is currently displayed in (GMT-04:00) Eastern Time (US & Canada).

Use conference time zone: (GMT-04:00) Eastern Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Mon 28 Apr 2025 17:00 - 17:05 at 215 - Software evolution and analysis Chair(s): Mauricio Verano Merino

Abstract

Modern programming languages are constantly evolving, introducing new language features and APIs to enhance software development practices. Software developers frequently face the challenge of upgrading their codebase to adapt new programming language versions, which is a tedious and time-consuming process. Recently, large language models (LLMs) have demonstrated potential in automating various code generation and editing tasks, suggesting their applicability in automating code upgrade efforts as well. Despite their promise, there exists no benchmark for evaluating the code upgrade ability of LLMs, as distilling relevant code changes related to programming language evolution from real-world software repositories’ commit histories is a complex challenge. In this work, we introduce CoUpJava, the first large-scale dataset for code upgrade in Java. CoUpJava comprises 10,697 code upgrade samples, distilled from the commit histories of 1,379 open-source Java repositories and covering Java versions 7–23. The dataset is divided into two subsets: CoUpJava-Fine, which captures fine-grained method-level refactorings towards new language features, and CoUpJava-Coarse, which includes coarse-grained repository-level changes encompassing new language features, standard library APIs, and build system upgrades. Our proposed dataset provides high-quality samples by filtering irrelevant and noisy changes and verifying the compilability of upgraded code. Moreover, CoUpJava reveals diversity in code upgrade scenarios, ranging from small, fine-grained refactorings to large-scale repository modifications.

Kaihang Jiang

University of Waterloo

Bihui Jin

University of Waterloo

Canada

Pengyu Nie

University of Waterloo

Canada

Time Zone

The program is currently displayed in (GMT-04:00) Eastern Time (US & Canada).

Use conference time zone: (GMT-04:00) Eastern Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

Display full programSpecify a time band

Save

Session Program

Mon 28 Apr
Displayed time zone: Eastern Time (US & Canada) change

16:00 - 17:30	Software evolution and analysisData and Tool Showcase Track / Technical Papers / Industry Track at 215 Chair(s): Mauricio Verano Merino Vrije Universiteit Amsterdam

16:00 10m Talk		50 Years of Programming Language Evolution through the Software Heritage looking glass Technical Papers Adèle Desmazières Sorbonne Unversité, Roberto Di Cosmo Inria, France / University of Paris Diderot, France, Valentin Lorentz Inria Foundation Pre-print
16:10 10m Talk		It Works (only) on My Machine: A Study on Reproducibility Smells in Ansible Scripts Technical Papers Ghazal Sobhani Dalhousie University, Israat Haque Dalhousie University, Tushar Sharma Dalhousie University Pre-print
16:20 10m Talk		Are the Majority of Public Computational Notebooks Pathologically Non-Executable? Technical Papers Waris Gill Virginia Tech, Muhammad Ali Gulzar Virginia Tech, Tien Nguyen Virginia Tech Pre-print
16:30 10m Talk		Understanding Test Deletion in Java Applications Technical Papers Suraj Bhatta North Dakota State University, Frank Kendemah North Dakota State University, Ajay Jha North Dakota State University Pre-print
16:40 10m Talk		A Public Benchmark of REST APIs Technical Papers Alix Decrop University of Namur, Sara Eraso University of Valle, Xavier Devroey University of Namur, Gilles Perrouin Fonds de la Recherche Scientifique - FNRS & University of Namur Pre-print
16:50 5m Talk		What Do Contribution Guidelines Say About Software Testing? Technical Papers Bruna Pereira Falcucci UFMG, Felipe Gomide UFMG, Andre Hora UFMG Pre-print Media Attached
16:55 5m Talk		Measuring InnerSource Value Industry Track Chamindra de Silva Citibank, Daniel Izquierdo-Cortazar Bitergia
17:00 5m Talk		CoUpJava: A Dataset of Code Upgrade Histories in Open-Source Java Repositories Data and Tool Showcase Track Kaihang Jiang University of Waterloo, Bihui Jin University of Waterloo, Pengyu Nie University of Waterloo
17:05 5m Talk		EvoChain: A Framework for Tracking and Visualizing Smart Contract Evolution Data and Tool Showcase Track Ilham Qasse Reykjavik University, Mohammad Hamdaqa Polytechnique Montreal, Björn Þór Jónsson Reykjavik University
17:10 5m Talk		CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance Data and Tool Showcase Track Kunal Suresh Pai UC Davis, Prem Devanbu University of California at Davis, Toufique Ahmed IBM Research Pre-print
17:15 5m Talk		RefExpo: Unveiling Software Project Structures through Advanced Dependency Graph Extraction Data and Tool Showcase Track Vahid Haratian Bilkent Univeristy, Pouria Derakhshanfar JetBrains Research, Vladimir Kovalenko JetBrains Research, Eray Tüzün Bilkent University
17:20 5m Talk		HyperAST: Incrementally Mining Large Source Code Repositories Data and Tool Showcase Track Quentin Le Dilavrec TU Delft, Netherlands, Andy Zaidman TU Delft Pre-print