Shopping MMLU

A multitask benchmark for language models across shopping knowledge, reasoning, behavior and languages.

What the resource contains

  • Multiple-choice questions with answer options
  • Retrieval, ranking and entity-recognition tasks
  • Generation inputs and target outputs

Potential uses

  • Evaluate shopping assistants
  • Compare multilingual shopping reasoning
  • Measure task-specific model performance

Access & formats

CSV / JSON. Repository / linked release. Follow the project’s instructions on GitHub; some resources require external files, registration or approval.

License & reuse

Apache-2.0 is identified for the repository. Check the released benchmark data and competition conditions separately.

Important limitations

A benchmark score does not establish reliable real-world purchasing behavior. Preserve the official task and evaluation definitions.

Source & attribution

KL4805 — original GitHub project ↗

Documentation reviewed 2026-09-24. Review scope and licensing notes. Suggest a correction.

Related resources

Shopping AI benchmarks

WebShop

princeton-nlp

A simulated shopping environment with products and human instructions for training and evaluating web agents.

Review required

JSON / JSONL / environment assets · External dependencies