EcomInstruct Evaluation

Released evaluation tasks from EcomGPT covering product language understanding and generation.

What the resource contains

  • 12 evaluation datasets with 500 sampled instances each
  • English and Chinese tasks including attributes, classification and Q&A
  • Per-task metadata and test JSON files

Potential uses

  • Evaluate ecommerce instruction following
  • Compare attribute extraction
  • Test product classification and title generation

Access & formats

JSON. Repository files. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.

License & reuse

The reviewed README does not establish a single license covering all contributed source datasets. Check each upstream dataset.

Important limitations

The repository describes a larger 2.5-million-instance training corpus; this entry covers the explicitly released evaluation subsets, not an assumed full release.

Source & attribution

Alibaba-NLP — original GitHub project ↗

Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.

Related resources

Shopping AI benchmarks

Shopping MMLU

KL4805

A multitask benchmark for language models across shopping knowledge, reasoning, behavior and languages.

Review required

CSV / JSON · Repository / linked release

Shopping AI benchmarks

WebShop

princeton-nlp

A simulated shopping environment with products and human instructions for training and evaluating web agents.

Review required

JSON / JSONL / environment assets · External dependencies