Shopping MMLU
A multitask benchmark for language models across shopping knowledge, reasoning, behavior and languages.
What the resource contains
- Multiple-choice questions with answer options
- Retrieval, ranking and entity-recognition tasks
- Generation inputs and target outputs
Potential uses
- Evaluate shopping assistants
- Compare multilingual shopping reasoning
- Measure task-specific model performance
Access & formats
CSV / JSON. Repository / linked release. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.
License & reuse
Apache-2.0 is identified for the repository. Check the released benchmark data and competition conditions separately.
Important limitations
A benchmark score does not establish reliable real-world purchasing behavior. Preserve the official task and evaluation definitions.
Source & attribution
KL4805 — original GitHub project ↗
Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.
Related resources
EcomInstruct Evaluation
Alibaba-NLP
Released evaluation tasks from EcomGPT covering product language understanding and generation.
JSON · Repository files
WebShop
princeton-nlp
A simulated shopping environment with products and human instructions for training and evaluating web agents.
JSON / JSONL / environment assets · External dependencies
Amazon Product Text Evaluation
amazon-science
Product-description evaluation data for assessing language-model outputs against product features.
Text / structured metadata · Repository files