TL;DR
Human evaluation of generated text often lacks transparency and consistency, leading to unreliable assessments. A large-scale analysis was conducted on human evaluation protocols for long-form text generation, reviewing 284 papers from CL conference publications.
✦ Why It Matters
Engineers and researchers should prioritize transparent evaluation protocols to ensure reliable assessments of text generation models.
Key Takeaways
Full Summary
Human evaluation is essential for determining the quality of generated text, yet many current practices lack transparency and thorough documentation. This study performed a comprehensive analysis of human evaluation protocols specifically for long-form text generation tasks, focusing on 284 papers from the Computational Linguistics (CL) conference publications between 2023 and 2025.
The methodology included both manual reviews and LLM-assisted analysis to assess the rigor of evaluation protocols. Results revealed that many studies did not provide sufficient details on their evaluation methods, leading to questions about the reliability of their findings.
For instance, only a small percentage of papers clearly documented their evaluation criteria and processes. These insights underscore the need for standardized and transparent evaluation practices in the field, which could enhance the reproducibility of results and improve the overall quality of generated text.
Related