TY - JOUR T1 - Effect of rater training on the reliability of technical skill assessments: a randomized controlled trial JF - Canadian Journal of Surgery JO - CAN J SURG SP - 405 LP - 411 DO - 10.1503/cjs.015917 VL - 61 IS - 6 AU - Reagan L. Robertson AU - Ashley Vergis AU - Lawrence M. Gillman AU - Jason Park Y1 - 2018/12/01 UR - http://canjsurg.ca/content/61/6/405.abstract N2 - Background: Rater training improves the reliability of observational assessment tools but has not been well studied for technical skills. This study assessed whether rater training could improve the reliability of technical skill assessment.Methods: Academic and community surgeons in Royal College of Physicians and Surgeons of Canada surgical subspecialties were randomly allocated to either rater training (7-minute video incorporating frame-of-reference training elements) or no training. Participants then assessed trainees performing a suturing and knot-tying task using 3 assessment tools: a visual analogue scale, a task-specific checklist and a modified version of the Objective Structured Assessment of Technical Skill global rating scale (GRS). We measured interrater reliability (IRR) using intraclass correlation type 2.Results: There were 24 surgeons in the training group and 23 in the no-training group. Mean assessment tool scores were not significantly different between the 2 groups. The training group had higher IRR than the no-training group on the visual analogue scale (0.71 v. 0.46), task-specific checklist (0.46 v. 0.33) and GRS (0.71 v. 0.61). However, confidence intervals were wide and overlapping for all 3 tools.Conclusion: For education purposes, the reliability of the visual analogue scale and GRS would be considered “good” for the training group but “moderate” for the no-training group. However, a significant difference in IRR was not shown, and reliability remained below the desired level of 0.8 for high-stakes testing. Training did not significantly improve assessment tool reliability. Although rater training may represent a way to improve reliability, further study is needed to determine effective training methods. ER -