The Accuracy Evaluation Shared Task As A Retrospective Reproduction Study

Craig Thomson, Ehud Reiter

GenChal - Thursday 07/21 12:00 EST
Abstract: We investigate the data collected for the Accuracy Evaluation Shared Task as a retrospective reproduction study. The shared task was based upon errors found by human annotation of computer generated summaries of basketball games. Annotation was performed in three separate stages, with texts taken from the same three systems and checked for errors by the same three annotators. We show that the mean count of errors was consistent at the highest level for each experiment, with increased variance when looking at per-system and/or per-errortype breakdowns.