DeepFashion MultiModal

Fashion images with rich visual and language annotations for controlled multimodal research.

What the resource contains

  • 44,096 human fashion images
  • Parsing, pose-related and attribute annotations
  • Text descriptions describing clothing appearance

Potential uses

  • Research text-guided fashion editing
  • Train parsing and attribute models
  • Evaluate image-language alignment

Access & formats

Images / annotation files. External dependencies. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.

License & reuse

The publisher limits the dataset to noncommercial research and restricts redistribution and commercial derivatives.

Important limitations

Images of people and derived annotations must remain within publisher conditions. This is not a commercial catalog feed.

Source & attribution

yumingj — original GitHub project ↗

Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.

Related resources

Fashion & product vision

Fashion-MNIST

zalandoresearch

A compact fashion-image classification benchmark suitable for learning and quick model comparisons.

Open license stated

IDX / image arrays · Repository / linked release

Fashion & product vision

FEIDEGGER

zalandoresearch

Fashion photographs paired with multiple German descriptions for multimodal retrieval research.

Review required

Images / text annotations · Repository / linked release

Fashion & product vision

DeepFashion2

switchablenorms

Clothing images with detailed annotations for detection, segmentation and consumer-to-shop retrieval.

Review required

Images / JSON annotations · External dependencies