Examining LLMs Ability to Summarize Code Through Mutation-Analysis (AIware 2026 - Main Track)

Who

Lara Khatib, Michael Pu, Bogdan Vasilescu, Mei Nagappan

Track

AIware 2026 Main Track

This program is tentative and subject to change.

Time Zone

The program is currently displayed in (GMT-04:00) Eastern Time (US & Canada).

Use conference time zone: (GMT-04:00) Eastern Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

By setting a time band, the program will dim events that are outside this time window. This is useful for (virtual) conferences with a continuous program (with repeated sessions).
The time band will also limit the events that are included in the personal iCalendar subscription service.

Display full programSpecify a time band

Save

When

Mon 6 Jul 2026 09:40 - 09:45 at MB 1.210 - Coding Agents, Software Testing, and Code Understanding

Abstract

As developers increasingly rely on LLM-generated code summaries for documentation, testing, and review, it is important to study whether these summaries accurately reflect what the program actually does. LLMs often produce confident descriptions of what the code looks like it should do (intent), while missing subtle edge cases or logic changes that define what it actually does (behavior). We present a mutation-based evaluation methodology that directly tests whether a summary truly matches the code’s logic. Our approach generates a summary, injects a targeted mutation into the code, and checks if the LLM updates its summary to reflect the new behavior. We validate it through three experiments totalling 624 mutated samples across 62 programs. First, on 12 controlled synthetic programs with 324 mutations varying in type (statement, value, decision) and location (beginning, middle, end). We find that summary accuracy decreases sharply with complexity from 76.5% for single functions to 17.3% for multi-threaded systems, while mutation type and location exhibit weaker effects. Second, testing 150 mutated samples on 50 human-written programs from the Less Basic Python Problems (LBPP) dataset confirms the same failure patterns persist as models often describe algorithmic intent rather than actual mutated behavior with a summary accuracy rate of 49.3%. Furthermore, while a comparison between GPT-4 and GPT-5.2 shows a substantial performance leap (from 49.3% to 85.3%) and an improved ability to identify mutations as “bugs”, both models continue to struggle with distinguishing implementation details from standard algorithmic patterns. This work establishes mutation analysis as a systematic approach for assessing whether LLM-generated summaries reflect program behavior rather than superficial textual patterns.

Lara Khatib

University of Waterloo

Canada

Michael Pu

University of Waterloo

Bogdan Vasilescu

Carnegie Mellon University

United States

Mei Nagappan

University of Waterloo

Canada

This program is tentative and subject to change.

Time Zone

The program is currently displayed in (GMT-04:00) Eastern Time (US & Canada).

Use conference time zone: (GMT-04:00) Eastern Time (US & Canada)Select other time zone

The GMT offsets shown reflect the offsets at the moment of the conference.

Time Band

Display full programSpecify a time band

Save

Session Program

Mon 6 Jul
Displayed time zone: Eastern Time (US & Canada) change

08:50 - 10:30	Coding Agents, Software Testing, and Code UnderstandingArXiv Track / Main Track at MB 1.210

08:50 5m Talk		Collaborator or Assistant? How AI Coding Agents Partition Work Across Pull Request Lifecycles Main Track Young Jo Chung , Safwat Hassan University of Toronto
08:55 5m Talk		When Code Authors Are Agents: A Large-Scale Study of Human–Agent Collaboration in Pull Requests Main Track Anthonia Oluchukwu Njoku École Polytechnique de Montréal, Université de Montréal, CA, Zohreh Sharafi Polytechnique Montréal, Foutse Khomh Polytechnique Montréal
09:00 5m Talk		Understanding Conversational Patterns in Multi-Agent Programming: A Case Study On Fibonacci Game Development Main Track Srijita Basu Chalmers University of Technology and University of Gothenburg, Viktor Kjellberg Göteborg University, SE, Simin Sun , Bengt Haraldsson Chalmers University of Technology and University of Gothenburg, Scania CV AB, Md Abu Ahammed Babu Volvo Cars AB, Wilhelm Meding Ericsson, Farnaz Fotrousi Chalmers University of Technology and University of Gothenburg, Miroslaw Staron Chalmers University of Technology and University of Gothenburg
09:05 5m Talk		Recovering from Misbehaviors in Coding Agents Main Track Rahul Nanda Facebook, US, Chandra Maddila Meta Platforms, Inc., Smriti Jha Facebook, US, Euna Mehnaz Khan , Satish Chandra Meta Platforms, Inc., Matteo Paltenghi University of Stuttgart
09:10 5m Talk		Configuring Agentic AI Coding Tools: An Exploratory Study Main Track Matthias Galster University of Canterbury, Seyedmoein Mohsenimofidi Heidelberg University, Jai Lal Lulla Singapore Management University, Muhammad Auwal Abubakar Otto-Friedrich Universität Bamberg, DE, Christoph Treude Singapore Management University, Sebastian Baltes Heidelberg University Pre-print
09:15 5m Talk		Execution Control Matters: Deterministic and Agentic Tool Orchestration for LLM-Based Code Translation Main Track Naing Oo Lwin Bucknell University, US, Rajesh Kumar Bucknell University, US
09:20 5m Talk		Developer Experience with AI Coding Agents: HTTP Behavioral Signatures in Documentation Portals ArXiv Track Oleksii Borysenko Cisco DevNet
09:25 5m Talk		VISOR: A Vision-Language Model-based Test Oracle for Testing Robots Main Track Prasun Saurabh Simula Research Laboratory, NO, Pablo Valle Mondragon University, Aitor Arrieta Mondragon University, Shaukat Ali Simula Research Laboratory and Oslo Metropolitan University, Paolo Arcaini National Institute of Informatics
09:30 5m Talk		Fixpad++: Automated Bug Fix Verification Using LLM Agents Main Track Mustafa Özkan İr Bilkent University, Bilkent University, TR, Mehmet Dedeler Bilkent University, Bilkent University, TR, Anil Koyuncu Bilkent University, Eray Tüzün Bilkent University
09:35 5m Talk		Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation Main Track Éric Jacopin Cosmic AI, FR
09:40 5m Talk		Examining LLMs Ability to Summarize Code Through Mutation-Analysis Main Track Lara Khatib University of Waterloo, Michael Pu University of Waterloo, Bogdan Vasilescu Carnegie Mellon University, Mei Nagappan University of Waterloo
09:45 5m Talk		Testing AIware Systems: A Software Engineering Survey Main Track Karla Gonzalez Royal Military College of Canada, Mariam El Mezouar Royal Military College
09:50 5m Talk		TestMap: Evidence Infrastructure for Foundation-Model-Assisted Test Generation ArXiv Track Hunter Leary Virginia Tech, Luke Hanuska Virginia Tech, Chris Brown Virginia Tech
09:55 5m Talk		Metamorphic Testing for Clinical ML Models: A Framework Proposal and Pilot Study ArXiv Track Jie JW Wu Michigan Technological University, USA, Feiyu E Michigan Technological University, USA, Bo Chen Michigan Technological University, USA
10:00 5m Talk		An Empirical Study of Reasoning Steps in Thinking Code LLMs Main Track Haoran Xue York University, CA, Gias Uddin York University, Canada, Song Wang York University
10:05 5m Talk		Can LLMs really reason about Code? Studying how well LLMs understand the relation between Input, Code, and Output Main Track Norman Becker CISPA Helmholtz Center for Information Security, DE, Tural Mammadov CISPA Helmholtz Center for Information Security, Andreas Zeller CISPA Helmholtz Center for Information Security
10:10 5m Talk		How Robustly do LLMs Understand Execution Semantics? Main Track Claudio Spiess University of California, Davis, Premkumar Devanbu UC Davis, Earl T. Barr University College London
10:15 5m Talk		Program-as-Weights: A Programming Paradigm for Fuzzy Functions ArXiv Track Wentao Zhang University of Waterloo, Liliana Hotsko University of Waterloo, Woojeong Kim Cornell University, Pengyu Nie University of Waterloo, Stuart Shieber Harvard University, Yuntian Deng University of Waterloo
10:20 10m Live Q&A		Joint Q&A and Discussion Main Track