← Node reference

Vision Captioner

Utility

Uses AI vision to describe what's in an image. Choose from multiple vision-language models and customize the description instructions. Outputs a text caption you can use as a prompt or combine with other text nodes.

Runs
In your browser
Inputs
1
Outputs
1

Inputs

PortTypeID
Imageimageimage-in

Outputs

PortTypeID
Captiontexttext-out
What this connects into (24)

Computed from the port compatibility rules, so it stays correct as nodes change.

Node ID vision-caption. Open it in Studio with Tab, then type the name.