Digitala Vetenskapliga Arkivet

Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Representation-Conditioned Diffusion Models for Guided Training Data Generation
Linköping University, Department of Science and Technology, Media and Information Technology. Linköping University, Faculty of Science & Engineering.ORCID iD: 0000-0003-3476-1986
Linköping University, Department of Science and Technology, Media and Information Technology. Linköping University, Faculty of Science & Engineering.ORCID iD: 0000-0002-7765-1747
Linköping University, Department of Science and Technology, Media and Information Technology. Linköping University, Faculty of Science & Engineering. Linköping University, Center for Medical Image Science and Visualization (CMIV).ORCID iD: 0000-0002-9217-9997
2026 (English)In: Synthetic Data for Computer Vision Workshop (SynData4CV), 2026Conference paper, Poster (with or without abstract) (Refereed)
Abstract [en]

Data availability remains a critical bottleneck in many deep learning applications. Large-scale datasets are often expensive to collect, curate and annotate, which can limit the scalability and applicability of supervised learning methods. In this work, we evaluate the classification performance of models trained on synthetic image datasets produced by generative deep learning. In particular, we use latent diffusion models conditioned on learned representations from DINOv2, DINOv3, and CLIP. Our results demonstrates that this representation-conditioned formulation significantly outperforms class-conditioned generation by a large margin (+10.76 p.p. top-1 accuracy on ImageNet100), by improving sample quality and mode coverage. Furthermore, by scaling the size of the synthetic dataset, we are able to outperform a classifier trained on the real data (+2.0 p.p top-1 accuracy).We also demonstrate how generated images can be used for augmentation purposes, outperforming classical augmentation methods, and how the conditioning space can be used for sample filtering to further improve training value. Collectively, these findings highlight that representation-conditioned diffusion models provide a promising approach for augmenting, complementing, or potentially replacing real-world datasets in large-scale visual learning tasks.

Place, publisher, year, edition, pages
2026.
Series
arXiv.org ; arXiv:2605.27495
Keywords [en]
DINOv2, DINOv3, CLIP, Diffusion Models, Synthetic data, Generative Data Augmentation
National Category
Computer Vision and Learning Systems Computer Sciences
Identifiers
URN: urn:nbn:se:liu:diva-226269OAI: oai:DiVA.org:liu-226269DiVA, id: diva2:2088235
Conference
The IEEE/CVF Conference on Computer Vision and Pattern Recognition 2026
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)Available from: 2026-07-26 Created: 2026-07-26 Last updated: 2026-07-26

Open Access in DiVA

No full text in DiVA

Search in DiVA

By author/editor
Karthikeyan, Nithesh ChandherUnger, JonasEilertsen, Gabriel
By organisation
Media and Information TechnologyFaculty of Science & EngineeringCenter for Medical Image Science and Visualization (CMIV)
Computer Vision and Learning SystemsComputer Sciences

Search outside of DiVA

GoogleGoogle Scholar

urn-nbn

Altmetric score

urn-nbn
Total: 58 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf