Auditing Who Appears to Belong: A Large-Scale Empirical Study of Bias in Deployed Text-to-Image Systems for Software Engineering
Generative image systems are increasingly embedded in software engineering artifacts such as slides, documentation, and recruiting collateral, shaping implicit signals about who is seen to “belong.” We present a mixed-methods empirical audit of 880 images generated by four widely used text-to-image models (GPT-4o/DALL·E 3, Llama-4/Emu, Qwen3-235B-A22B, Stable Diffusion) using 22 demographically neutral prompts, intentionally uniform to isolate default model priors, varying role, seniority, team context, geography, and language. Independent human annotations, triangulated with automated raters and validated through per-category agreement analysis, capture both demographic representation (gender, race/ethnicity, age) and portrayal cues (setting, attire, props, emotion). We analyze intersectional distributions and benchmark them against occupational reference statistics using risk ratios, JensenShannon divergence, and the Theil index. Across models, outputs consistently converge on a narrow archetype: young men dominate (95.8% male, 88% under 40), women and older professionals are rare, and several racial and ethnic groups are systematically underrepresented relative to workforce baselines. Prompt variation modestly shifts racialized appearance but leaves gender imbalance largely intact, while model differences are primarily of degree rather than direction. We translate these findings into actionable implications for AI-aware software engineering practice, including representational regression tests in CI/CD pipelines and diversity-aware generation defaults, arguing that evaluation of AI systems in software engineering must account for the societal signals conveyed by generated imagery alongside functional performance.