http://hdl.handle.net/1893/38384| Appears in Collections: | Computing Science and Mathematics Conference Papers and Proceedings |
| Peer Review Status: | Refereed |
| Author(s): | Dovhoshliubnyi, llia Soroush, Nima Sami, Ashkan Brownlee, Alexander |
| Contact Email: | alexander.brownlee@stir.ac.uk |
| Title: | What Do AI Agents Actually Change? An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests Anonymous Anonymous Institution |
| Citation: | Dovhoshliubnyi l, Soroush N, Sami A & Brownlee A (2026) What Do AI Agents Actually Change? An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests Anonymous Anonymous Institution. In: <i>Search-Based Software Engineering</i>. 18th Symposium on Search-Based Software Engineering 2026 (SSBSE 2026) Challenge Track, Montreal, Canada, 05.07.2026-06.07.2026. https://doi.org/10.1007/978-3-032-30699-9_12 |
| Issue Date: | 10-Jul-2026 |
| Date Deposited: | 5-May-2026 |
| Conference Name: | 18th Symposium on Search-Based Software Engineering 2026 (SSBSE 2026) Challenge Track |
| Conference Dates: | 2026-07-05 - 2026-07-06 |
| Conference Location: | Montreal, Canada |
| Abstract: | AI coding agents are black boxes: we cannot inspect how they generate code, but we can inspect what they change. This distinction matters for search-based software engineering (SBSE), where techniques such as genetic improvement depend on mutation operators that reflect how code is actually transformed. Of the 33,596 agent PRs in the AIDev dataset, less than 400 target performance (fewer than 1%), making each successful case a valuable window into otherwise opaque agent behaviour. We classify 1,254 performance-relevant diff hunks from 216 of these PRs, spanning five agent systems, against the 18-category syntactic mutation taxonomy of Even-Mendoza et al. (2025) using an LLM-as-a-judge pipeline. Three categories dominate: name modification (36.9%), object creation (26.3%), and type change (22.6%), a profile strikingly different from prior genetic improvement corpora where no change accounted for 84%. Each agent commits to a distinctive mutation vocabulary, and each performance strategy activates a largely disjoint category subset. Agent identity and target strategy are therefore informative priors that narrow the effective SBSE operator space from 18 categories to a handful per context. Replication package: https://anonymous.4open.science/r/ssbse-challenge-2026-710C/ |
| Status: | AM - Accepted Manuscript |
| Rights: | This version of the article has been accepted for publication, after peer review (when applicable) and is subject to Springer Nature’s AM terms of use, but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/978-3-032-30699-9_12. |
| File | Description | Size | Format | |
|---|---|---|---|---|
| What Do AI Agents Actually Change An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests.pdf | Fulltext - Accepted Version | 421.06 kB | Adobe PDF | Under Embargo until 2027-07-11 Request a copy |
Note: If any of the files in this item are currently embargoed, you can request a copy directly from the author by clicking the padlock icon above. However, this facility is dependent on the depositor still being contactable at their original email address.
This item is protected by original copyright |
Items in the Repository are protected by copyright, with all rights reserved, unless otherwise indicated.
The metadata of the records in the Repository are available under the CC0 public domain dedication: No Rights Reserved https://creativecommons.org/publicdomain/zero/1.0/
If you believe that any material held in STORRE infringes copyright, please contact library@stir.ac.uk providing details and we will remove the Work from public display in STORRE and investigate your claim.
