WDC PAVE

Product attribute extraction and normalization resources for turning heterogeneous offer text into consistent fields.

What the resource contains

  • Processed product datasets
  • Attribute annotations
  • Extraction and normalization evaluation resources

Potential uses

  • Normalize product characteristics
  • Compare attribute extraction models
  • Prepare consistent catalog schemas

Access & formats

Processed dataset files. Repository files. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.

License & reuse

No clear dataset license was identified in the reviewed source. Check repository terms and upstream rights.

Important limitations

Coverage and preprocessing vary by benchmark subset. Inspect each subset before combining labels or comparing results.

Source & attribution

wbsg-uni-mannheim — original GitHub project ↗

Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.

Related resources

Catalog & attributes

MAVE

google-research-datasets

Multi-source product attribute extraction annotations built from Amazon product profiles.

Review required

JSONL · External dependencies

Catalog & attributes

WDC Products

wbsg-uni-mannheim

A product entity-matching benchmark for identifying offers that describe the same product.

Review required

Structured records / benchmark splits · External dependencies

Catalog & attributes

OA-Mine

xinyangz

Weakly supervised resources for discovering both attribute types and values in ecommerce product titles.

Review required

Task-specific data files · Repository files