ControlNet
ControlNet adds spatial control (edges, depth, pose, segmentation) to pretrained text-to-image diffusion models without retraining the base model.
A trainable copy of the diffusion network is attached to the frozen model through zero-initialised convolutions, so a conditioning map steers layout while the base model's quality is preserved. Works with Stable Diffusion and trains robustly on small (<50k) and large (>1M) datasets.
- Date
- Friday, 10 February 2023
- Lab
- Stanford University
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Training-set range tested | <50k to >1M samples abstract says training holds up well across both | authors |
arXiv v1 date 2023-02-10; authors Lvmin Zhang, Anyi Rao, Maneesh Agrawala. Code released on GitHub per the paper page.
Sources
This record was checked against its sources on 6 October 2026. How we check