AIware 2026
Mon 6 - Tue 7 July 2026 Montreal, Canada
co-located with FSE 2026
Events (11 results)

VeriTrans: Fine-Tuned LLM-Assisted NL→PL Translation via a Deterministic Neuro-symbolic Pipeline

Main Track When: Mon 6 Jul 2026 14:00 - 14:05 People: Xuan Liu, Dheeraj Kodakandla, Kushagra Srivastva, Mahfuza Farooque

… , all executed via fixed API configuration (temperature= 0; fine-tuning runs use … runtime, and all prompts/responses and timing metadata are logged to enable …

AgentTelemetry: A Fault Detection Benchmark and Toolkit for LLM Agent Observability

Benchmark & Dataset Track When: Tue 7 Jul 2026 15:10 - 15:15 People: Krishna Chaitanya Balusu

all nine span kinds are necessary: removing any one makes at least one fault … not statistically significant at $n{=}24$). All code, data, and benchmark configurations …

Co-located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation

Main Track When: Mon 6 Jul 2026 09:35 - 09:40 People: Éric Jacopin

… %) across all models except Claude 3.5 Haiku, which strips all doctests; (2) separated …

CrossCommitVuln-Bench: A Dataset of Multi-commit Python Vulnerabilities Invisible to Per-Commit Static Analysis

Benchmark & Dataset Track When: Tue 7 Jul 2026 14:40 - 14:45 People: Arunabh Majumdar

… % across all 30 vulnerabilities — 93% of chains are invisible to per-commit SAST …

Metamorphic Testing for Clinical ML Models: A Framework Proposal and Pilot Study

ArXiv Track When: Mon 6 Jul 2026 09:55 - 10:00 People: Jie JW Wu, Feiyu E, Bo Chen

… on the UCI Heart Disease dataset, where all three clinical models tested (AUROC …

JunoBench: A Benchmark Dataset of Crashes in Python Machine Learning Jupyter Notebooks

Benchmark & Dataset Track When: Tue 7 Jul 2026 15:05 - 15:10 People: Yiran Wang, José Antonio Hernández López, Ulf Nilsson, Daniel Varro

… a unified execution environment that reliably reproduces all crashes. In addition …

How Robustly Do LLMs Understand Execution Semantics?

Main Track When: Mon 6 Jul 2026 10:10 - 10:15 People: Claudio Spiess, Premkumar Devanbu, Earl T. Barr

… point to limitations in the way all models understand code, and reinforce …

Wink: Recovering from Misbehaviors in Coding Agents

Main Track When: Mon 6 Jul 2026 09:05 - 09:10 People: Rahul Nanda, Chandra Maddila, Smriti Jha, Euna Mehnaz Khan, Satish Chandra, Matteo Paltenghi

… Problems, and Tool Call Failures, which we find occur in about 30% of all agent …

Detecting Unsoundness in Neural Network Verifiers via Concrete–Abstract Consistency

Main Track When: Mon 6 Jul 2026 14:30 - 14:35 People: Kaijie Liu, Yulei Sui

… by considering all possible specified behaviors, the soundness …

Can LLMs Really Reason about Code? Studying How Well LLMs Understand the Relation between Input, Code, and Output

Main Track When: Mon 6 Jul 2026 10:05 - 10:10 People: Norman Becker, Tural Mammadov, Andreas Zeller

… -weight models
achieve the strongest performance across all datasets …

From Features to Data and Domain Knowledge: Reflections on Two Decades of AI for Software

Keynotes When: Mon 6 Jul 2026 11:40 - 12:00 People: Lin Tan

… correctness to maintenance and security. LLMs are powerful across all