datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ecom-niverse
Ecom-niverse
What is Ecom-niverse
We construct a comprehensive e-commerce tokens dataset by refining a broad web dataset to isolate content with retail or shopping context. This curated corpus is intended for continual pre-training of LLMs and other Encoder-only models so they better understand product descriptions, prices, and other commerce-related text
Need for E-commerce pre-training Dataset
Generic web-crawled corpora often lack the focused coverage of… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/Ecom-niverse.economics_reports_v2
Vidore Benchmark 2 - World Economics report Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_v2.womens-clothing-ecommerce-reviews
Dataset Card for "womens-clothing-ecommerce-reviews"
Processed version of this dataset.
ai-ecosystem-daily
TensorFeed AI Ecosystem Daily
Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL.
Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.ECommerce-Women-Clothing-Reviewsopen-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.IndustryCorpus2_finance_economics
IndustryCorpus2: Finance & Economics
This repository contains the IndustryCorpus2: Finance & Economics domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_finance_economics.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.kaggle-womens-ecom-clothing-reviews
Women's Clothing E-Commerce Reviews (wide-row repack)
Repack of Kaggle dataset nicapotato/womens-ecommerce-clothing-reviews (CC0-1.0) into a single wide parquet with ZStandard level 9 compression.
Schema: one row per review — row_id (int64), clothing_id (int32), age (int32), rating (int8), recommended_ind (int8), positive_feedback_count (int32), title (string, nullable), review_text (string, nullable), division_name (string, nullable), department_name (string, nullable)… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/kaggle-womens-ecom-clothing-reviews.zh-vie_ecom
1688 zh-vi ecom
Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên
ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới".
Xem dữ liệu ở đâu
data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây:
bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự
dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi
(bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.Ecommerce_textafrica-world-bank-economics-time-series
Africa World Bank Economics Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal economic observations, proposed state labels, anomaly… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-economics-time-series.imfweoThis dataset is derived from the International Monetary Fund (IMF) World Economic Outlook (WEO) Database.
The use of IMF Data is governed by the IMF’s “The Use of IMF Data” special terms: you are allowed to download, extract, copy, create derivative works, publish and distribute the data provided you attribute the source, preserve data accuracy, declare any material transformations, and pass these terms on. See: IMF Copyright and Usage – The Use of IMF Data.
This data has been… See the full description on the dataset page: https://huggingface.co/datasets/econdataverse/imfweo.gspc-ai-economy-index
GSPC — ai adoption components facts (Eurostat)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Live axis name: ai-adoption-components — MEASURED as two Eurostat series (deterministic-facts, n=2). Not an index. No composite formula. No MEASURED-INDEX-v0.1 sticker (C-2026-0826-05: do not restore).
This Hub repo id keeps the legacy slug gspc-ai-economy-index for inbound links only. Cite the live axis name. Do not stamp an… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-ai-economy-index.neon4cast-scoresSnapshot of the Ecological Forecasting Initiative NEON Forecasting Challenge
Includes probabilistic forecasts, observations, and skill scores across all submitted forecasts over 5 challenge themes.
Ecom-niverseus-economic-events
US Economic Events
What the US government published, when the embargo lifted, and what the
number was at that moment — not what it has since been revised to.
2 994 official releases · 34 251 observations · 13 323 events ·
12 release families · 2010-01-07 to 2026-09-10
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
The number you remember is not the number that was published
Total nonfarm payrolls for May 2026, as… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/us-economic-events.ecommerce-behavior-data-from-multi-category-store_oct-nov_2019
eCommerce Behavior Data from Multi-Category Store
About the Dataset
This dataset contains behavioral data for 285 million user events from a large multi-category eCommerce store. The data spans 7 months (October 2019 - April 2020) and records various user interactions with products.
Dataset Overview
Time Frame: October 2019 - April 2020
Total Events: 285 million
Event Granularity: Each row represents an event associated with a product and a user.
Data Source:… See the full description on the dataset page: https://huggingface.co/datasets/kevykibbz/ecommerce-behavior-data-from-multi-category-store_oct-nov_2019.marqo-general-ecommerce-evalsemantic-ecology
NuBerea Semantic Ecology
Feature-indexed analysis layer for the study of canon formation, part of the
NuBerea corpus estate of biblical and patristic texts. It describes the
"semantic ecology" of ancient biblical and related literature — how semantic
domains are used across texts, how candidate texts were received over time,
and which measurable features accompany canonical inclusion — packaged as a
set of ready-to-load configurations.
Attribution
NuBerea project.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/semantic-ecology.olist-ecommerce-for-delivery-and-review-prediction
E-Commerce Analytics for Delivery and Review Prediction
This dataset was created for a datathon project. It's a cleaned and feature-engineered version of the public Olist Brazilian E-commerce dataset, specifically prepared to predict shipping delays and customer review scores.
Project Goals
Our project focuses on two key business problems:
Model 1 (Regression): Can we predict how delayed a shipment will be? This helps manage customer expectations proactively.
Model 2… See the full description on the dataset page: https://huggingface.co/datasets/miminmoons/olist-ecommerce-for-delivery-and-review-prediction.IndustryInstruction_Finance-Economics
IndustryInstruction: Finance & Economics
This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.zenodo-ecommerce-text
E-commerce Text Classification (wide-row repack)
Repack of Zenodo record 10.5281/zenodo.3355823 (Saurabh Gautam, CC-BY-4.0) into a single wide parquet with ZStandard level 9 compression.
Schema: one row per product, three columns — row_id (int64), category (string: Household, Books, Electronics, Clothing & Accessories), description (string).
50,424 rows retained. 1 upstream row dropped for null description.
Source DOI: 10.5281/zenodo.3355823. See CITATION.cff.
kovidore-v2-economic-beirKoViDoRe v2 : Economic trends
This dataset, Economic trends, is a corpus of periodic reports on major economic indicators in Korea, intended for complex-document understanding tasks. It is one of the 4 corpora comprising the KoViDoRe v2 Benchmark.
Links
Github: https://github.com/whybe-choi/kovidore-benchmark
Collection: https://huggingface.co/collections/whybe-choi/kovidore-benchmark-beir-v2
Data Generation Pipeline: https://github.com/whybe-choi/kovidore-data-generator… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/kovidore-v2-economic-beir.ecommerce-product-classification-by-categories
🛍️ E-Commerce Product Classification Dataset
Dataset Description
This dataset contains 31,851 realistic English product descriptions across 33 e-commerce categories, designed for training product classification models.
Key Features
✅ 31,851 samples with natural language variations
✅ 33 product categories covering major e-commerce segments
✅ 4 seller personas: Professional, Individual, Reseller, Minimal
✅ Realistic noise: Typos, abbreviations, casual language… See the full description on the dataset page: https://huggingface.co/datasets/Lezh1n/ecommerce-product-classification-by-categories.e_coli_rnasecommerce-user-behavior-dataafrica-algeria-algeria-economic-social-environmental-health-education-dev-5f717023
Algeria - Economic, Social, Environmental, Health, Education, Development and Energy | Africa (Algeria official open data)
75,901 rows - 1 Africa country - 1960-2025 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Algeria as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-algeria-algeria-economic-social-environmental-health-education-dev-5f717023.EcomRetrieval
EcomRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
EcomRetrieval
Task category
t2t
Domains
None
Reference
https://arxiv.org/abs/2203.03367
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["EcomRetrieval"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)
To learn more about how to run models on mteb task check out… See the full description on the dataset page: https://huggingface.co/datasets/mteb/EcomRetrieval.HUGO-Bench-Paper-Reproducibility
HUGO-Bench Paper Reproducibility
Supplementary data and reproducibility materials for the paper:
Vision Transformers for Zero-Shot Clustering of Animal Images: A Comparative Benchmarking Study - https://arxiv.org/abs/2602.03894
Hugo Markoff, Stefan Hein Bengtson, Michael Ørsted
Aalborg University, Denmark
Dataset Description
This repository contains complete experimental results, pre-computed embeddings, and execution logs from our comprehensive benchmarking study… See the full description on the dataset page: https://huggingface.co/datasets/AI-EcoNet/HUGO-Bench-Paper-Reproducibility.
