Most weak results come from prompts that describe a category instead of a picture. "A nice product photo" gives the model nothing to aim at. "A ceramic mug on a pale oak table, morning light from the left, shallow depth of field" gives it four decisions it no longer has to guess.

Say what's in the frame

Name the subject, the surface it's on, and what's behind it. Vague backgrounds are where results go generic, because the model fills the gap with an average of everything it has seen.

Describe the light

Lighting does more for realism than any other single word you can add. Soft window light, hard overhead sun, golden hour, overcast, studio softbox, each produces a visibly different image. If you say nothing about light, you get whatever the model defaults to.

Set the framing

Close-up, mid shot, wide shot. Eye level, low angle, overhead. These are cheap words that change the composition completely, and they're the first thing to adjust when a result is technically fine but doesn't look like what you pictured.

Cut the adjectives that don't decide anything

"Stunning", "beautiful", "high quality", "amazing" don't constrain the image, because they don't describe anything visual. Swap them for the concrete detail you actually meant. "Professional" means nothing; "plain gray backdrop, even lighting, no shadow on the wall" means something.

Change one thing at a time

When a result is close, resist rewriting the whole prompt. Adjust the single element that's wrong and run it again. Rewriting everything makes it impossible to tell which change did what, and it burns credits learning nothing.

Length isn't the point

A long prompt full of filler is worse than a short one full of decisions. Every word should rule something out.