Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI labs models. Youll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based