Iris-3B

Pixel-space generation & general vision learner

Iris-3B is a pixel-space generative model that can act as a general vision learner. Working in pixel space, it bypasses the lossy compression and the biased, texture-focused latents of a VAE. We show image generators can replace vision models such as DINOv2, release our pixel-space scaling recipes, and put the generative prior to work on dense tasks where detail matters.

Hanqiu Li Cai†Chema GarabitoSperidlabs · † project lead

03 — Architecture

Architecture

A 3B flow-matching transformer that generates raw RGB pixels directly, with no VAE in the loop.

Iris-3B architecture
FIG. 02 — ArchitectureTrunk · PiT head · trunk block

Images are cut into 16 × 16 pixel patches. The trunk has 8 dual-stream blocks, with separate text and image weights, followed by 16 single-stream blocks.

A small pixel-transformer (PiT) head decodes each patch token back to pixels. Text comes from a frozen Qwen3-VL-4B encoder, and 2D RoPE lets one set of weights run at any resolution and aspect ratio up to one megapixel.

04 — Benchmarks

Benchmarks

Competitive with state-of-the-art latent models on every official benchmark, at a fraction of their training compute.

ModelGenEvalDPGLongTextOneIG
SDXL—74.65—0.316
SD3 Medium0.6284.08——
FLUX.1-dev 12B0.6683.80.6070.434
SD3.5 Large0.71——0.462
Lumina-Image 2.00.7387.20—0.353
Janus-Pro-7B0.8084.190.0190.267
HiDream-I1-Full0.8385.890.5430.477
Z-Image 6B*0.8488.10.9350.546
Qwen-Image 20B0.8788.30.9430.539
Iris-3B0.79886.520.8570.540

Iris-3B: official evaluators at 1024², no prompt rewriting, no best-of-N. References from the Qwen-Image and Z-Image reports; — not reported; * uses prompt rewriting.

05 — Downstream

Downstream

Iris-3B tuned for downstream tasks.

05.A

Depth

DepthPhotoPhotoDepth
DepthPhotoPhotoDepth
DepthPhotoPhotoDepth
DepthPhotoPhotoDepth
DepthPhotoPhotoDepth
DepthPhotoPhotoDepth

Point clouds from predicted depth · drag to orbit, scroll to zoom

05.B

Image Restoration / Super-resolution

Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
Iris-3BInputInputIris-3B
06 — Citation

Cite this work

@techreport{licai2026iris,
  title       = {Iris-3B: Going Beyond the Latent with Pixel-Space
                 Diffusion Training, Conversion and Fine-Tuning},
  author      = {Li Cai, Hanqiu and Garabito, Chema},
  institution = {Speridlabs},
  year        = {2026}
}