Text prompts describe content well and geometry badly. Guidance models solve that by accepting a structural input alongside the prompt.
The main guidance types
- Edge detection. Follows the outline of a reference image. Useful for keeping a composition while changing the subject entirely.
- Depth maps. Preserves the spatial relationship of a scene. Useful for interiors and landscapes where perspective matters.
- Pose skeletons. Fixes the position of a human figure. Standard practice for character work.
- Segmentation. Assigns regions to specific elements, giving the tightest layout control at the cost of more setup.
Setting guidance strength
Too low and the structure is ignored; too high and the output looks traced, with visible artefacts along edges. Start in the middle and adjust in small steps while watching where the image departs from the reference.
Combining guidance
Depth plus pose handles most character-in-environment work. Stacking three or more guidance inputs usually creates conflicts and produces a muddled result.
Comments (0)
Log in to join the discussion
Log InNo comments yet