Human Evaluation of Synthetic Videostroboscopic Laryngeal Images Generated Using StyleGAN3.

TitleHuman Evaluation of Synthetic Videostroboscopic Laryngeal Images Generated Using StyleGAN3.
Publication TypeJournal Article
Year of Publication2026
AuthorsMorrison DA, Mohanty AS, Sulica L, Khosravi P, Rameau A
JournalLaryngoscope
Date Published2026 Aug 26
ISSN1531-4995
Abstract

OBJECTIVE: To evaluate the perceptual realism of synthetic videostroboscopic laryngeal images and determine the relationship between training duration and clinician classification accuracy.

METHODS: Synthetic images were generated using StyleGAN3 from two age-stratified datasets: Dataset A (≥ 65 years) and Dataset B (< 65 years). A total of 114 clinicians evaluated a randomized, balanced-controlled 36-image survey drawn from a 144-image pool containing real frames and five StyleGAN3 training intervals (5120-25,000 kimg). Accuracy was analyzed across image classes, experience levels, and viewing devices.

RESULTS: Clinicians demonstrated significantly higher mean accuracy identifying true clinical frames than synthetic generations (70.9% vs. 57.8%, p < 0.001). Overall survey accuracy was 59.9% ± 14.2%, with realism peaking at 20,120 kimg (44.8% accuracy). Accuracy dropped significantly between 5120 kimg (79.6%) and 10,120 kimg (59.1%, p < 0.001), with diminishing returns thereafter. No significant accuracy difference existed between Datasets A and B (p = 0.27). The effect of specialty reached borderline significance (p = 0.059), though limited by severe subgroup imbalances (e.g., n = 5 fellows, n = 2 residents). Computer users (63.9%) were significantly more accurate than phone users (56.2%, p = 0.014). A confounder analysis confirmed specialty and device choice were statistically independent (p = 0.164).

CONCLUSION: StyleGAN3 produces synthetic laryngeal images demonstrating high perceptual realism. The significant real-versus-synthetic accuracy gap highlights that while true frames remain distinct, synthetic generations achieve profound ambiguity. Perceptual realism plateaus early (10,120-20,120 kimg), suggesting moderate training durations can produce human-perceived realism while reducing computational and environmental demands.

LEVEL OF EVIDENCE: N/A.

DOI10.1002/lary.70870
Alternate JournalLaryngoscope
PubMed ID42649563
Grant ListK76 AG079040 / AG / NIA NIH HHS / United States
OT2 OD032720 / / NIH Common Fund /