試験

問題 3 / 65

What are the main problems with using only ROUGE scores to evaluate summarization models?

Because ROUGE is a metric that directly measures the fluency of generated text, it can evaluate meaning.
ROUGE is a metric for evaluating a model's inference speed.
ROUGE only evaluates n-gram overlap and cannot adequately assess factuality or consistency.
ROUGE is exclusively for evaluating translation and cannot be used for summarization.
ROUGE always ends up giving long summaries high scores.

当サイトでは、ユーザー体験の向上を目的としてCookieを使用しています。サイトの利用を継続することで、Cookieの使用に同意したものとみなされます。