
Vision–language alignment architecture
A model architecture diagram with image encoder, text encoder, projection layers, and a contrastive alignment objective.
- Field
- Computer science
- Style
- Flat schematic
- Model
- GPT Image 2.5
- Aspect ratio
- 3:2
Prompt
Create a publication figure in the visual style of Nature Machine Intelligence: a model architecture diagram: image encoder (patches, transformer blocks), text encoder, projection heads into a shared embedding space, contrastive similarity matrix, and a downstream decoder, clean boxes and arrows. flat vector schematic, white background, precise English labels.




