AIware 2026
Mon 6 - Tue 7 July 2026 Montreal, Canada
co-located with FSE 2026

AI coding assistants and autonomous agents are becoming integral to software development workflows, reshaping how code is produced, reviewed, and maintained. While recent research has focused mainly on the capabilities and impacts of productivity of these systems, much less attention has been paid to accountability: who is responsible when agents generate, modify, or recommend code? In practice, accountability is defined through the Terms of Service (ToS) and related policy documents that govern the use of AI-powered development tools.

In this vision paper, we present a comparative analysis of the Terms of Service for widely used AI coding assistants and agent-enabled development tools. We examine how these documents allocate ownership, responsibility, liability, and disclosure obligations between tool providers and software developers, and we identify common patterns and divergences between providers. Our analysis reveals a consistent tendency to shift responsibility for correctness, safety, and legal compliance onto users, as well as substantial variation in how providers address issues such as indemnification, data reuse, and acceptable use.

Based on these findings, we argue that existing policy frameworks are poorly aligned with increasingly agent-mediated and autonomous software development workflows. We outline a research roadmap for accountable agents in software engineering, identifying challenges and opportunities for modeling responsibility, designing governance artifacts, developing tooling that supports accountability, and conducting empirical studies of developers’ perceptions and practices.

Christoph Treude is an Associate Professor of Computer Science at Singapore Management University. His work spans empirical and automated software engineering (SE), artificial intelligence (AI) and software engineering (AI&SE), human-AI collaboration, and AI for science. He has authored over 200 scientific publications in collaboration with more than 300 co-authors. His research has been recognized with five best paper awards, including three ACM SIGSOFT Distinguished Paper Awards, and has received funding from Google, Facebook, DST, and through an ARC Discovery Early Career Researcher Award (2018–2020). Before joining SMU, he held academic positions at the University of Melbourne and the University of Adelaide, and postdoctoral appointments at McGill University, the University of São Paulo, and the Federal University of Rio Grande do Norte. He currently serves on the editorial boards of the IEEE Transactions on Software Engineering and Empirical Software Engineering (Springer), and as Open Science Editor for the Journal of Systems and Software (Elsevier). He is co-chairing the program committee for FSE 2026 and regularly serves on program committees across leading software engineering venues.

Tue 7 Jul

Displayed time zone: Eastern Time (US & Canada) change

14:00 - 15:30
Human Factors, Responsible AIware, and Benchmarks & DatasetsBenchmark & Dataset Track / Main Track at MB 1.210
Chair(s): Diego Elias Costa Concordia University, Canada
14:00
5m
Talk
Is Artificial Intelligence an Elixir to the Software Engineering Community? An Empirical Study among ManagersACM SIGSOFT Distinguished Paper Award
Main Track
Xin Zhao Seattle University, Brian Vu Seattle University, US, Sitesh Pattanaik Donald Bren School of Information and Computer Sciences, University of California, Irvine, US
DOI
14:05
5m
Talk
Towards AI as a Collaborative Partner: A Taxonomy of AI Agent Behavior in Software Engineering
Main Track
Tao Dong Google, Sherry Shi Google, Harini Sampath , Andrew Macvean Google, Inc.
DOI Pre-print
14:10
5m
Talk
Auditing Who Appears to Belong: A Large-Scale Empirical Study of Bias in Deployed Text-to-Image Systems for Software Engineering
Main Track
Mohamad Kassab Boston University
DOI
14:15
5m
Talk
Operationalizing Ethics for AI Agents: How Developers Encode Values into Repository Context Files
Main Track
Christoph Treude Singapore Management University, Sebastian Baltes Heidelberg University, Marc Cheong the University of Melbourne
DOI Pre-print
14:20
5m
Talk
Accountable Agents in Software Engineering: An Analysis of Terms of Service and a Research Roadmap
Main Track
Christoph Treude Singapore Management University
DOI Pre-print
14:25
5m
Talk
SOSecure: The Wisdom of the Crowd for Safer AI-Generated Code
Main Track
Manisha Mukherjee Carnegie Mellon University, Vincent J. Hellendoorn Google DeepMind
DOI
14:30
5m
Talk
SecVulEval: Context-Aware Benchmarking of LLMs for Vulnerability DetectionAIware Best Benchmark/Dataset Paper Award
Benchmark & Dataset Track
Md Basim Uddin Ahmed York University, CA, Nima Shiri Harzevili York University, Jiho Shin York University, Hung Viet Pham York University, Song Wang York University
DOI
14:35
5m
Talk
SecMutBench: Evaluating LLM-Generated Security Tests via Mutation-Based Vulnerability Detection
Benchmark & Dataset Track
Mariam ALMutairi Virginia Polytechnic Institute and State University, US
DOI
14:40
5m
Talk
CrossCommitVuln-Bench: A Dataset of Multi-commit Python Vulnerabilities Invisible to Per-Commit Static Analysis
Benchmark & Dataset Track
Arunabh Majumdar Independent Researcher, IN
DOI
14:45
5m
Talk
REBench: A Procedural, Fair-by-Construction Benchmark for LLMs on Stripped-Binary Types and Names
Benchmark & Dataset Track
Jun Yeon Won Ohio State University, Columbus, US, Xin Jin Meta, Shiqing Ma University of Massachusetts at Amherst, Zhiqiang Lin The Ohio State University
DOI
14:50
5m
Talk
RustBuildEq: A Benchmark for Binary Equivalence under Build Variability
Benchmark & Dataset Track
Elliott Wen The University of Auckland, Chenye Ni , Valerio Terragni University of Auckland, Jens Dietrich Victoria University of Wellington
DOI
14:55
5m
Talk
TOGBench: A Developer-Written Multi-variant Dataset and Benchmark Suite for Test Oracle Generation
Benchmark & Dataset Track
Tasfia Tasnim University of Texas at Dallas, US, Matthew B Dwyer University of Virginia, Soneya Binta Hossain University of Texas at Dallas
DOI
15:00
5m
Talk
HEJ-Robust: A Robustness Benchmark for LLM-Based Automated Program Repair
Benchmark & Dataset Track
Fazle Rabbi Concordia University, Jinqiu Yang Concordia University
DOI
15:05
5m
Paper
JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks
Benchmark & Dataset Track
Yiran Wang Linköping University, José Antonio Hernández López Department of Computer Science and Systems, University of Murcia, Ulf Nilsson Linköping University, Daniel Varro Linköping University / McGill University
DOI Pre-print
15:10
5m
Talk
AgentTelemetry: A Fault Detection Benchmark and Toolkit for LLM Agent Observability
Benchmark & Dataset Track
DOI
15:15
15m
Live Q&A
Joint Q&A
Benchmark & Dataset Track