試験
通常
試験時間: 00:00問題時間: 00:00
問題一覧
問題 3 / 65
What are the main problems with using only ROUGE scores to evaluate summarization models?
①Because ROUGE is a metric that directly measures the fluency of generated text, it can evaluate meaning.
②ROUGE is a metric for evaluating a model's inference speed.
③ROUGE only evaluates n-gram overlap and cannot adequately assess factuality or consistency.
④ROUGE is exclusively for evaluating translation and cannot be used for summarization.
⑤ROUGE always ends up giving long summaries high scores.
