Global results

Twenty-Prompt Attack vs Generation Study

Attack versus independent generation variation across twenty controlled prompts.

All distances reference each image’s own prompt baseline. Compare values only within the same encoder, space, dimension count, and metric.

The exact preexisting 600 images were reused; no images were regenerated.

Prompt-balanced full-space separation

Generation / attack is the mean distance between independently generated images divided by the mean distance between an attacked version of the same source image and its baseline. The attack distance defines the 1× reference; values above 1× mean that different semantic generations are farther apart than attacked copies.

EncoderAttack meanGeneration meanGeneration / attack95% CI
BLIP CLS0.35660.54521.53×1.3966–1.6668
CLIP0.23850.46531.95×1.7511–2.1467
DINOv20.21670.77323.57×3.2325–3.9167
Qwen3-VL Embedding 2B0.24350.53112.18×1.9747–2.4272
DINOv3-L/160.22370.63342.83×2.5057–3.2158
Qwen3-VL Embedding 8B0.29930.62262.08×1.8795–2.3107

DINOv2 vs DINOv3

DINOv2 produced the stronger separation signal in this experiment. This is specific to the completed twenty-prompt robustness protocol and is not a universal ranking of DINOv2 and DINOv3 as representation models.

Measured DINO separation
EncoderAttack meanGeneration meanGeneration / attack95% CIAUC
DINOv20.21670.77323.568×3.232–3.9170.984
DINOv3-L/160.22370.63342.832×2.506–3.2160.940

Per-family caption robustness

Per-family caption robustness
CaptionerAttack familyAttack changeGeneration changeGeneration / attack95% CI
BLIP captionnoise0.17370.29541.70×1.1716–2.8609
BLIP captionnoise_denoise0.16450.29541.80×1.2893–2.8910
BLIP captionblur0.19640.29541.50×1.1156–2.2660
BLIP captiongeometric0.14490.29542.04×1.2568–3.6122
BLIP captioncompression0.10930.29542.70×1.8047–4.5860
BLIP captionresampling0.10020.29542.95×1.9254–6.0923
Detailed Florence captionnoise0.32100.52791.64×1.3452–2.1093
Detailed Florence captionnoise_denoise0.34410.52791.53×1.2531–1.9604
Detailed Florence captionblur0.37740.52791.40×1.1716–1.7559
Detailed Florence captiongeometric0.27780.52791.90×1.6137–2.3282
Detailed Florence captioncompression0.25510.52792.07×1.7317–2.6112
Detailed Florence captionresampling0.28410.52791.86×1.4158–2.6666
Qwen3-VL 8B neutralnoise0.45520.69361.52×1.3836–1.6814
Qwen3-VL 8B neutralnoise_denoise0.48280.69361.44×1.3383–1.5517
Qwen3-VL 8B neutralblur0.58470.69361.19×1.1210–1.2619
Qwen3-VL 8B neutralgeometric0.44910.69361.54×1.3987–1.7459
Qwen3-VL 8B neutralcompression0.42940.69361.62×1.4788–1.7866
Qwen3-VL 8B neutralresampling0.45490.69361.52×1.3757–1.7111
Qwen3-VL 8B artifact-invariantnoise0.43460.67281.55×1.3793–1.7741
Qwen3-VL 8B artifact-invariantnoise_denoise0.44670.67281.51×1.3282–1.7439
Qwen3-VL 8B artifact-invariantblur0.54770.67281.23×1.1486–1.3280
Qwen3-VL 8B artifact-invariantgeometric0.40780.67281.65×1.4409–1.9295
Qwen3-VL 8B artifact-invariantcompression0.38360.67281.75×1.5262–2.0534
Qwen3-VL 8B artifact-invariantresampling0.44090.67281.53×1.3638–1.7492

Paired prompt ablation

Neutral uses the standard scene-description instruction: describe the visible objects, actions, scene, and relationships. Artifact-invariant uses the same Qwen3-VL 8B model with a different instruction: focus on stable semantic content and ignore minor blur, noise, JPEG, resampling, and color artifacts. This comparison tests an instruction change—not a different model—and is not a proven paper method.

Paired prompt ablation
Attack familyHelpful promptsGeneration preservedMean invariant − neutral attack change
blur13/2020/20-0.0370
compression13/2020/20-0.0458
geometric13/2020/20-0.0413
noise12/2020/20-0.0206
noise_denoise13/2020/20-0.0361
resampling10/2020/20-0.0140

Runtime and provenance

Runtime and peak GPU memory
ModeRuntime secondsPeak GPU memory GiB
Runtime/memory remains recorded in immutable shard provenance.
Exact model revisions
ComponentModelRevisionRole
DINOv3-L/16facebook/dinov3-vitl16-pretrain-lvd1689mea8dc2863c51be0a264bab82070e3e8836b02d51embedding
Qwen3-VL Embedding 8BQwen/Qwen3-VL-Embedding-8B2c4565515e0f265c6511776e7193b22c0968ddc7embedding
Qwen3-VL 8B artifact-invariantQwen/Qwen3-VL-8B-Instruct0c351dd01ed87e9c1b53cbc748cba10e6187ff3bprompt_ablation
Qwen3-VL 8B neutralQwen/Qwen3-VL-8B-Instruct0c351dd01ed87e9c1b53cbc748cba10e6187ff3bprimary

BLIP CLS

Normalized full embedding spaces are encoder-specific. Raw-global PCA is fitted to the original embeddings from all prompts, so prompt identity and image variation both affect its axes. Centered-global PCA first subtracts each prompt’s baseline embedding, then fits one shared set of axes to the resulting within-prompt displacements; all prompt baselines therefore share the origin. Axes are shared only within the same encoder. Local PCA axes are not shared.

Full-space distributions and PCA views

Full-space L2 histogram plot
Full-space L2 histogram
Full-space L2 violin / boxplot plot
Full-space L2 violin / boxplot
Raw-global 2D PCA plot
Raw-global 2D PCA
Raw-global 3D PCA plot
Raw-global 3D PCA
Centered-global 2D PCA plot
Centered-global 2D PCA
Centered-global 3D PCA plot
Centered-global 3D PCA
Explained variance plot
Explained variance
Mean / attack separation factor plot
Mean / attack separation factor

Prompt-balanced minimum, median, and maximum

Embedding generalization
Attack meanGeneration meanGeneration / attack95% CIDistance-only AUCHeld-out protocol
0.35660.54521.53×1.3966–1.66680.8499use separator evaluation prompt and prompt-plus-family folds

CLIP

Normalized full embedding spaces are encoder-specific. Raw-global PCA is fitted to the original embeddings from all prompts, so prompt identity and image variation both affect its axes. Centered-global PCA first subtracts each prompt’s baseline embedding, then fits one shared set of axes to the resulting within-prompt displacements; all prompt baselines therefore share the origin. Axes are shared only within the same encoder. Local PCA axes are not shared.

Full-space distributions and PCA views

Full-space L2 histogram plot
Full-space L2 histogram
Full-space L2 violin / boxplot plot
Full-space L2 violin / boxplot
Raw-global 2D PCA plot
Raw-global 2D PCA
Raw-global 3D PCA plot
Raw-global 3D PCA
Centered-global 2D PCA plot
Centered-global 2D PCA
Centered-global 3D PCA plot
Centered-global 3D PCA
Explained variance plot
Explained variance
Mean / attack separation factor plot
Mean / attack separation factor

Prompt-balanced minimum, median, and maximum

Embedding generalization
Attack meanGeneration meanGeneration / attack95% CIDistance-only AUCHeld-out protocol
0.23850.46531.95×1.7511–2.14670.8920use separator evaluation prompt and prompt-plus-family folds

DINOv2

Normalized full embedding spaces are encoder-specific. Raw-global PCA is fitted to the original embeddings from all prompts, so prompt identity and image variation both affect its axes. Centered-global PCA first subtracts each prompt’s baseline embedding, then fits one shared set of axes to the resulting within-prompt displacements; all prompt baselines therefore share the origin. Axes are shared only within the same encoder. Local PCA axes are not shared.

Full-space distributions and PCA views

Full-space L2 histogram plot
Full-space L2 histogram
Full-space L2 violin / boxplot plot
Full-space L2 violin / boxplot
Raw-global 2D PCA plot
Raw-global 2D PCA
Raw-global 3D PCA plot
Raw-global 3D PCA
Centered-global 2D PCA plot
Centered-global 2D PCA
Centered-global 3D PCA plot
Centered-global 3D PCA
Explained variance plot
Explained variance
Mean / attack separation factor plot
Mean / attack separation factor

Prompt-balanced minimum, median, and maximum

Embedding generalization
Attack meanGeneration meanGeneration / attack95% CIDistance-only AUCHeld-out protocol
0.21670.77323.57×3.2325–3.91670.9841use separator evaluation prompt and prompt-plus-family folds

DINOv3-L/16

Normalized full embedding spaces are encoder-specific. Raw-global PCA is fitted to the original embeddings from all prompts, so prompt identity and image variation both affect its axes. Centered-global PCA first subtracts each prompt’s baseline embedding, then fits one shared set of axes to the resulting within-prompt displacements; all prompt baselines therefore share the origin. Axes are shared only within the same encoder. Local PCA axes are not shared.

Full-space distributions and PCA views

Full-space L2 histogram plot
Full-space L2 histogram
Full-space L2 violin / boxplot plot
Full-space L2 violin / boxplot
Raw-global 2D PCA plot
Raw-global 2D PCA
Raw-global 3D PCA plot
Raw-global 3D PCA
Centered-global 2D PCA plot
Centered-global 2D PCA
Centered-global 3D PCA plot
Centered-global 3D PCA
Explained variance plot
Explained variance
Mean / attack separation factor plot
Mean / attack separation factor

Prompt-balanced minimum, median, and maximum

Embedding generalization
Attack meanGeneration meanGeneration / attack95% CIDistance-only AUCHeld-out protocol
0.22370.63342.83×2.5057–3.21580.9395use separator evaluation prompt and prompt-plus-family folds

Qwen3-VL Embedding 2B

Normalized full embedding spaces are encoder-specific. Raw-global PCA is fitted to the original embeddings from all prompts, so prompt identity and image variation both affect its axes. Centered-global PCA first subtracts each prompt’s baseline embedding, then fits one shared set of axes to the resulting within-prompt displacements; all prompt baselines therefore share the origin. Axes are shared only within the same encoder. Local PCA axes are not shared.

Full-space distributions and PCA views

Full-space L2 histogram plot
Full-space L2 histogram
Full-space L2 violin / boxplot plot
Full-space L2 violin / boxplot
Raw-global 2D PCA plot
Raw-global 2D PCA
Raw-global 3D PCA plot
Raw-global 3D PCA
Centered-global 2D PCA plot
Centered-global 2D PCA
Centered-global 3D PCA plot
Centered-global 3D PCA
Explained variance plot
Explained variance
Mean / attack separation factor plot
Mean / attack separation factor

Prompt-balanced minimum, median, and maximum

Embedding generalization
Attack meanGeneration meanGeneration / attack95% CIDistance-only AUCHeld-out protocol
0.24350.53112.18×1.9747–2.42720.9288use separator evaluation prompt and prompt-plus-family folds

Qwen3-VL Embedding 8B

Normalized full embedding spaces are encoder-specific. Raw-global PCA is fitted to the original embeddings from all prompts, so prompt identity and image variation both affect its axes. Centered-global PCA first subtracts each prompt’s baseline embedding, then fits one shared set of axes to the resulting within-prompt displacements; all prompt baselines therefore share the origin. Axes are shared only within the same encoder. Local PCA axes are not shared.

Full-space distributions and PCA views

Full-space L2 histogram plot
Full-space L2 histogram
Full-space L2 violin / boxplot plot
Full-space L2 violin / boxplot
Raw-global 2D PCA plot
Raw-global 2D PCA
Raw-global 3D PCA plot
Raw-global 3D PCA
Centered-global 2D PCA plot
Centered-global 2D PCA
Centered-global 3D PCA plot
Centered-global 3D PCA
Explained variance plot
Explained variance
Mean / attack separation factor plot
Mean / attack separation factor

Prompt-balanced minimum, median, and maximum

Embedding generalization
Attack meanGeneration meanGeneration / attack95% CIDistance-only AUCHeld-out protocol
0.29930.62262.08×1.8795–2.31070.9198use separator evaluation prompt and prompt-plus-family folds