ON COLOR ALIGNMENT IN VAE LATENT SPACES
AND ITS APPLICATIONS

1 Computer Vision Center (CVC), Barcelona, Spain  |  2 Universitat Autònoma de Barcelona, Barcelona, Spain
3 City University of Hong Kong (Dongguan), China  |  4 City University of Hong Kong, China
5 Universitat de València, Spain
† Corresponding author
Input Image x
Input Image x
Latent Encoder Enc.
Latent Encoder Enc.
+ δ Shift
Decoded Output D(z + δ)
Neutral Baseline
Ch 2: +0.00  |  Ch 3: +0.00  |  Ch 1: 0.00

Color Space

+ Ch 3 +0.00
- Ch 2
-0.00
+ Ch 2
+0.00
- Ch 3 -0.00
Ch 1: Brightness 0.0 (Neutral)
-Δ (Darker) 0 (Baseline) +Δ (Brighter)

A color space hides in plain sight in the latent space of an autoencoder. Train a VAE with three latent channels on natural images, encode an image (left), add a constant offset δ to one channel, and decode: the image changes color while preserving its content. Sweeping two channels (right; Ch 2 across columns, Ch 3 across rows) spans an opponent-color plane and all intermediate mixtures, while the remaining channel (bottom row) controls brightness.

Abstract

Variational autoencoders (VAEs) are a key part of modern text-to-image models, which generate images within their latent space. VAEs are known to disentangle the main factors of variation in the data, and color is known to be one of the most structured of these in natural images: decorrelating it yields one luminance axis and two opponent-color axes. Color should therefore be expected to emerge as a distinct factor in the VAE latent space. Yet how these latent spaces represent color remains largely unexplored. In this work, we show that the VAEs of text-to-image models share a color subspace aligned with brightness and opponent-colors. Through a linear approximation of the encoder and targeted latent steering, we find this subspace consistently across a broad range of VAEs, from SD1.5 to FLUX.2 and Z-Image. Building on this characterization, we propose three applications: ColorTuning, which achieves state-of-the-art in precise numerical color generation on the fine-grained CSS3/X11 system of GenColorBench, saturation control, to adjust the global chromatic intensity, and color transfer, to change the palette to match a reference.

Color in Autoencoder Latent Space

Traversing principal latent directions in pretrained VAEs reveals an opponent-color plane and brightness axis that emerge spontaneously from natural image statistics.

FLUX.1 Latent Channel Traversals

FLUX.1-dev (16 latent channels): Sweeping principal directions spans an orthogonal opponent-color plane (ch2 across columns, ch3 across rows) and independent brightness control while preserving image structure.

FLUX.2 Latent Channel Traversals

FLUX.2-dev (32 latent channels): Higher dimensional latent spaces preserve the identical canonical chromatic organization, demonstrating model-agnostic universality across architectures.

SD3 Latent Channel Traversals

Stable Diffusion 3 (16 latent channels): Leading latent components isolate opponent-color transitions across natural scenes without affecting spatial geometry.

SDXL Latent Channel Traversals

SDXL (4 latent channels): In compact 4-channel latents, color information aligns directly with dominant channels, confirming that efficient image compression inherently isolates chromaticity.

Applications

Building on the characterization of the color basis, we propose three generation-time applications in text-to-image diffusion models.

ColorTuning: Numerical Color Precision

Exploits the discovered orthogonal directions to steer generation towards exact numerical colors (HEX, RGB, CIELAB). Replaces the numerical code with an ISCC-NBS Level 2 proxy color name to initiate diffusion in the target chromatic basin. At a calibrated gate step s, predicts clean latent ẑ0, segments the object via SAM, and measures the CIELAB color. A lightweight ResMLP predicts the displacement along the color basis (u1, u2, u3), injected inside the object mask with a linearly decaying schedule.

Saturation Control

Continuously modulates the chroma of the generated scene or specific objects without altering semantic structure or hue. Preserves lightness Li and hue angle while scaling chroma to (Li, (1 − α)ai, (1 − α)bi) for reduction factor α ∈ [0, 1]. Evaluates dense spatial displacements across the grid, synthesizing images within narrower color gamuts directly at generation time without post-processing.

Color Transfer

Steers the scene's color distribution to match either a discrete color palette or an exemplar reference image. Extracts principal colors via CIELAB k-means clustering with CIEDE2000 distance constraints, and uses the most chromatic color as a prompt proxy. At the gate step, semantic regions are identified and steered toward target palette colors via the latent color basis, faithfully reproducing the reference palette while generating the prompt content.

Interactive Demonstration

Explore the three generation-time modes.

Select Object:
“A banana on a kitchen countertop”
Select Target Color Swatch:
Generated object with steered color
Target Color: #DC143C
Selected Color Palette
Selected Color Palette or Reference
Generated Output
Generated image with transferred color
Select Scene:
Saturation Setting: 100% Saturation
100% 90% 80% 70% 60% 50%
Saturation Control Steered Generation

BibTeX Citation

@article{santamaria2025coloralignment,
  title={On Color Alignment in VAE Latent Spaces and Its Applications},
  author={Santamaria, Julian D. and Wang, Kai and Malo, Jes{\'u}s and Vazquez-Corral, Javier and Gomez-Villa, Alexandra},
  journal={arXiv preprint arXiv:XXXX.XXXXX},
  year={2025}
}