datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EconomicIndex
The Anthropic Economic Index
Overview
The Anthropic Economic Index provides insights into how AI is being incorporated into real-world tasks across the modern economy.
Data Releases
This repository contains multiple data releases, each with its own documentation:
Labor market impacts: Job exposure and task penetration data
2026-06-26 Release: Updated analysis with Artifacts and monthly aggregates
2026-03-24 Release: Updated analysis with Opus… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/EconomicIndex.Ecom-niverse
Ecom-niverse
What is Ecom-niverse
We construct a comprehensive e-commerce tokens dataset by refining a broad web dataset to isolate content with retail or shopping context. This curated corpus is intended for continual pre-training of LLMs and other Encoder-only models so they better understand product descriptions, prices, and other commerce-related text
Need for E-commerce pre-training Dataset
Generic web-crawled corpora often lack the focused coverage of… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/Ecom-niverse.geo-ecoeconomics_reports_v2
Vidore Benchmark 2 - World Economics report Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_v2.womens-clothing-ecommerce-reviews
Dataset Card for "womens-clothing-ecommerce-reviews"
Processed version of this dataset.
synthux-economy-r6-w034
SynthUX Computer-Use Dataset — economy sim r6
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r6-w034.synthux-economy-r6-w042
SynthUX Computer-Use Dataset — economy sim r6
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r6-w042.ai-ecosystem-daily
TensorFeed AI Ecosystem Daily
Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL.
Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.opencode-ecosystem-core
OpenCode Ecosystem Core
Ecossistema Python completo para orquestração de tarefas, memória metacognitiva, especificações SDD/TDD, integrações MCP e fluxos de pesquisa científica.
Visão Geral
O OpenCode Ecosystem Core é um ecossistema Python de código aberto que organiza o ciclo perceber → especificar → delegar → executar → verificar → refletir. Ele combina:
Interface de linha de comando (CLI) para diagnóstico, pesquisa e apresentações
Registro de… See the full description on the dataset page: https://huggingface.co/datasets/marceloclaro/opencode-ecosystem-core.ECommerce-Women-Clothing-Reviewsopen-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.IndustryCorpus2_finance_economics
IndustryCorpus2: Finance & Economics
This repository contains the IndustryCorpus2: Finance & Economics domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_finance_economics.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.zh-vie_ecom
1688 zh-vi ecom
Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên
ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới".
Xem dữ liệu ở đâu
data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây:
bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự
dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi
(bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.kaggle-womens-ecom-clothing-reviews
Women's Clothing E-Commerce Reviews (wide-row repack)
Repack of Kaggle dataset nicapotato/womens-ecommerce-clothing-reviews (CC0-1.0) into a single wide parquet with ZStandard level 9 compression.
Schema: one row per review — row_id (int64), clothing_id (int32), age (int32), rating (int8), recommended_ind (int8), positive_feedback_count (int32), title (string, nullable), review_text (string, nullable), division_name (string, nullable), department_name (string, nullable)… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/kaggle-womens-ecom-clothing-reviews.Ecommerce_textsynthux-economy-r5-w083
SynthUX Computer-Use Dataset — economy sim r5
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r5-w083.synthux-economy-r6-w031
SynthUX Computer-Use Dataset — economy sim r6
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r6-w031.africa-world-bank-economics-time-series
Africa World Bank Economics Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal economic observations, proposed state labels, anomaly… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-economics-time-series.imfweoThis dataset is derived from the International Monetary Fund (IMF) World Economic Outlook (WEO) Database.
The use of IMF Data is governed by the IMF’s “The Use of IMF Data” special terms: you are allowed to download, extract, copy, create derivative works, publish and distribute the data provided you attribute the source, preserve data accuracy, declare any material transformations, and pass these terms on. See: IMF Copyright and Usage – The Use of IMF Data.
This data has been… See the full description on the dataset page: https://huggingface.co/datasets/econdataverse/imfweo.synthux-economy-r6-w049
SynthUX Computer-Use Dataset — economy sim r6
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r6-w049.gspc-ai-economy-index
GSPC — ai adoption components facts (Eurostat)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Live axis name: ai-adoption-components — MEASURED as two Eurostat series (deterministic-facts, n=2). Not an index. No composite formula. No MEASURED-INDEX-v0.1 sticker (C-2026-0826-05: do not restore).
This Hub repo id keeps the legacy slug gspc-ai-economy-index for inbound links only. Cite the live axis name. Do not stamp an… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-ai-economy-index.neon4cast-scoresSnapshot of the Ecological Forecasting Initiative NEON Forecasting Challenge
Includes probabilistic forecasts, observations, and skill scores across all submitted forecasts over 5 challenge themes.
rawbench_ecoliEcom-niversesynthux-economy-r6-w048
SynthUX Computer-Use Dataset — economy sim r6
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r6-w048.synthux-economy-r6-w041
SynthUX Computer-Use Dataset — economy sim r6
Grounded visual computer-use trajectories generated by SynthUX. Each record is
one worker's device session inside a simulated company: a goal expands into a
node tree, terminal nodes drive real desktop-simulator apps (Terminal, Notes,
VS Code, Browser, Slack/Teams, Mail, Sheets, Slides, …) through low-level
mouse/keyboard input, and the observed trajectory is recorded as a screen video
plus per-node frames.
App-native… See the full description on the dataset page: https://huggingface.co/datasets/jacob-valdez/synthux-economy-r6-w041.Olist_Ecommerce_Datasetlibero-wo-ecotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 3917,
"total_frames": 567494,
"total_tasks": 73,
"total_videos": 0,
"total_chunks": 4,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3917"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/uclanecl/libero-wo-ecot.
