MAVE
Multi-source product attribute extraction annotations built from Amazon product profiles.
What the resource contains
- Product IDs and categories
- Attribute keys, values and character-span evidence
- Labels for roughly 2.2 million product profiles across 1,257 categories
Potential uses
- Extract structured attributes
- Evaluate span detection
- Build product enrichment pipelines
Access & formats
JSONL. External dependencies. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.
License & reuse
The data license was not established from the reviewed README. Upstream Amazon content has separate access and use conditions.
Important limitations
The repository supplies labels, not the full product paragraphs. Full reconstruction requires the upstream 2018 metadata. Repository is archived.
Source & attribution
google-research-datasets — original GitHub project ↗
Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.
Related resources
WDC Products
wbsg-uni-mannheim
A product entity-matching benchmark for identifying offers that describe the same product.
Structured records / benchmark splits · External dependencies
WDC PAVE
wbsg-uni-mannheim
Product attribute extraction and normalization resources for turning heterogeneous offer text into consistent fields.
Processed dataset files · Repository files
OA-Mine
xinyangz
Weakly supervised resources for discovering both attribute types and values in ecommerce product titles.
Task-specific data files · Repository files