試験

問題 31 / 65

Which design modification is effective if you want a model that uses positional encodings, like a Vision Transformer, to operate without further training when the input size changes?

Keep using fixed absolute positional embeddings.
Always resize the input image to a fixed resolution before tokenizing.
Increase the number of Transformer layers to enhance expressive power.
Replace with a convolutional stem that takes optical properties into account.
Introduce relative positional encoding.

当サイトでは、ユーザー体験の向上を目的としてCookieを使用しています。サイトの利用を継続することで、Cookieの使用に同意したものとみなされます。