AWS Certification Exam AI Practitioner (Practice Questions) set1

問題 18 / 65

What is the most effective measure to keep inference latency low for high-frequency requests in a real-time text generation service?

Batch-process at night and return a response.
Output a large amount of logs and analyze them later.
Provide hot inference capacity by pre-provisioning resources such as provisioned concurrency and dedicated inference instances.
Add a multi-stage model chain to generate responses.
Save the results to S3 and have the client retrieve them later.

当サイトでは、ユーザー体験の向上を目的としてCookieを使用しています。サイトの利用を継続することで、Cookieの使用に同意したものとみなされます。