MAVE

Multi-source product attribute extraction annotations built from Amazon product profiles.

What the resource contains

  • Product IDs and categories
  • Attribute keys, values and character-span evidence
  • Labels for roughly 2.2 million product profiles across 1,257 categories

Potential uses

  • Extract structured attributes
  • Evaluate span detection
  • Build product enrichment pipelines

Access & formats

JSONL. External dependencies. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.

License & reuse

The data license was not established from the reviewed README. Upstream Amazon content has separate access and use conditions.

Important limitations

The repository supplies labels, not the full product paragraphs. Full reconstruction requires the upstream 2018 metadata. Repository is archived.

Source & attribution

google-research-datasets — original GitHub project ↗

Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.

Related resources

Catalog & attributes

WDC Products

wbsg-uni-mannheim

A product entity-matching benchmark for identifying offers that describe the same product.

Review required

Structured records / benchmark splits · External dependencies

Catalog & attributes

WDC PAVE

wbsg-uni-mannheim

Product attribute extraction and normalization resources for turning heterogeneous offer text into consistent fields.

Review required

Processed dataset files · Repository files

Catalog & attributes

OA-Mine

xinyangz

Weakly supervised resources for discovering both attribute types and values in ecommerce product titles.

Review required

Task-specific data files · Repository files