Bias in Student Evaluations of Surgical Attendings: Role of Gender, Race, Age and Experience
Publication type
journal article
Publication date
June 2026
Author(s)
Theis, Claudia
Jacob, Anusha S.
Kapadia, Muneera Rehana
Pascarella, Luigi
Publisher
Elsevier BV
Language
English
Discipline(s)
Geographical area
Abstract
Introduction
Student evaluations influence faculty promotion but may reflect implicit bias. We assessed whether surgeon gender, race, age, and experience were associated with differences in student evaluations at a large academic medical center.
Materials and methods
This retrospective cohort study included 149 surgical attendings evaluated by medical students at the University of North Carolina between 2016 and 2020. Quantitative evaluation items were rated on 5-point Likert scales and analyzed using surgeon-level summary comparisons (Wilcoxon rank-sum and Kruskal–Wallis tests), evaluation-level multivariable logistic regression, and generalized estimating equations (GEEs) to account for clustering of multiple evaluations per attending. Qualitative free-text comments were analyzed using a validated natural language processing framework and summarized on a 5-point sentiment scale.
Results
A total of 149 surgical attendings were evaluated, including 39 women (26.2%) and 110 men (73.8%). The racial and ethnic distribution was 102 White (68.5%), 28 Asian (18.8%), 13 Black (8.7%), and 6 Latino (4.0%). In total, 2475 quantitative evaluations were analyzed, of which 1542 (62.3%) included narrative comments. The median time in practice was 14 y (Q1-Q3: 9-22), and the median time at the University of North Carolina was 11 y (Q1-Q3: 7-19.5). Median composite quantitative evaluation scores were lower for women than for men (4.42 [4.32-4.61] versus 4.61 [4.40-4.82] on a 5-point scale; P = 0.002). In GEE models accounting for clustering of evaluations within attendings and adjusting for age, race/ethnicity, and years in practice, women remained independently associated with lower evaluation scores (β = −0.206; 95% confidence interval: −0.294 to −0.118; P < 0.001). Composite evaluation scores did not differ significantly across racial or ethnic groups (Kruskal–Wallis P = 0.40), and race was not a significant predictor in adjusted models (GEE P = 0.13). Attendings aged ≥50 y had lower unadjusted evaluation scores than younger colleagues (P = 0.03); however, age was not independently associated with evaluation scores after clustering adjustment (GEE P = 0.14). Qualitative sentiment scores were uniformly high (median 4.0 [Q1-Q3: 4.0-4.5]) and did not differ by gender, race, or age.
Conclusions
Women surgeons received lower quantitative scores than their male colleagues, despite similarly positive qualitative feedback. Age differences were observed in unadjusted analyses but were attenuated after accounting for clustering, whereas no significant differences were observed by race or years in practice. Although absolute score differences were modest, these findings suggest that commonly used student evaluation metrics may reflect gender-related bias and should be interpreted cautiously in academic surgical assessment.
Part of
Journal of Surgical Research
ISSN
0022-4804
Volume
322