試験
通常
試験時間: 00:00問題時間: 00:00
問題一覧
問題 37 / 65
Which evaluation metric is most appropriate when assessing the improvement effects of a conversational AI?
①Only look at the model's floating-point operation count.
②Training cost reduction rate
③Using only automatic language metrics such as BLEU and ROUGE.
④Evaluate based on actual users' task success rates and satisfaction.
⑤Minimize the average number of tokens in response messages.
