試験

問題 16 / 65

Which evaluation method is most reliable when comparing multiple prompt candidates to choose the best one?

Automatic comparison of generated token lengths and character counts
Comparison of average model inference times
Human evaluation of output quality
Comparison of model parameter counts and sizes
Choose randomly and use.

当サイトでは、ユーザー体験の向上を目的としてCookieを使用しています。サイトの利用を継続することで、Cookieの使用に同意したものとみなされます。