This is the default working method for anything that has to look like a specific thing. You settle composition, lighting, product and wardrobe as a still (cheap, fast, reviewable) and only then spend a video render on it. Text-to-video is for exploration; image-to-video is for delivery.
It also puts the gates in the right place. A label that fails at the still stage costs one image. The same failure discovered after motion costs the clip.
Is image-to-video better than text-to-video?
For controlled work, almost always. You approve the frame before you pay for the motion, and the model has far less room to reinvent the subject.