ReadGlim
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild — ReadGlim