datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ecom-niverse
Ecom-niverse
What is Ecom-niverse
We construct a comprehensive e-commerce tokens dataset by refining a broad web dataset to isolate content with retail or shopping context. This curated corpus is intended for continual pre-training of LLMs and other Encoder-only models so they better understand product descriptions, prices, and other commerce-related text
Need for E-commerce pre-training Dataset
Generic web-crawled corpora often lack the focused coverage of… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/Ecom-niverse.ai-ecosystem-daily
TensorFeed AI Ecosystem Daily
Daily snapshots of the AI ecosystem: news, model pricing, benchmarks, service status, GPU rental prices, MCP registry growth, LLM endpoint latency probes, agent traffic, and the AFTA adopter directory. Captured once per day from the public tensorfeed.ai API and committed to this repo as JSONL.
Each daily snapshot lives in a YYYY-MM-DD/ subfolder with one JSONL file per feed plus a manifest.json summarizing what was captured.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/tensorfeed/ai-ecosystem-daily.zh-vie_ecom
1688 zh-vi ecom
Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên
ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới".
Xem dữ liệu ở đâu
data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây:
bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự
dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi
(bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.gspc-ai-economy-index
GSPC — ai adoption components facts (Eurostat)
SWIFT census (live): https://councilof.ai/api/swift
XRPL reader (live): https://councilof.ai/api/xrpl
Live axis name: ai-adoption-components — MEASURED as two Eurostat series (deterministic-facts, n=2). Not an index. No composite formula. No MEASURED-INDEX-v0.1 sticker (C-2026-0826-05: do not restore).
This Hub repo id keeps the legacy slug gspc-ai-economy-index for inbound links only. Cite the live axis name. Do not stamp an… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-ai-economy-index.IndustryInstruction_Finance-Economics
IndustryInstruction: Finance & Economics
This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.financial-economics-reasoning
Model Card
📌 Summary
financial-economics-reasoning dataset was constructed using advanced Inference Distillation techniques. We employed the qwen-3-235b-a22b-thinking-2507 model as the Teacher Model to process the open-source BAAI/IndustryInstruction_Finance-Economics dataset, which contains 122,378 bilingual (Chinese-English) entries in finance, economics, and business.
Unlike standard distillation datasets that only provide final answers, this dataset retains the… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/financial-economics-reasoning.economist-tui-sessions
Coding agent session traces for thomasmustier/economist-tui-sessions
This dataset contains redacted coding agent session traces collected while working on tmustier/economist-tui. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/economist-tui-sessions.trade-economy-index
The Trade Economy Index
Release: 2026.1.6
How the US skilled trades fare in the AI era: labor, market structure, unit
economics, cash cycle, AI exposure, geography, and valuation for 13 commercial
trades, with per-cell citations where applicable and table-level provenance for
derived and aggregate tables.
Companion site: tradesindex.org. Published by
Level. Archived with a DOI:
10.5281/zenodo.21762674.
Also available as CSV and JSON on
GitHub.
Why this exists
Most… See the full description on the dataset page: https://huggingface.co/datasets/LevelCFO/trade-economy-index.semantic-economyall things are now lawful to you in jack feist
EA-RHIZOME-SE-01 — semantic economy
The archive's political economy of meaning, seeded at the stolon the collapse body left open. Its roles are NOT the collapse roles: where that body asks what narrows, this one asks who gains by the narrowing.
These are symbola. They are for traversal.
A token broken in two, each half held by a different party, no half carrying complete authority, the fit of the fracture proving the… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/semantic-economy.labour-economy-unmeasured
Labour economy — UNMEASURED, on purpose
Register: UNMEASURED. Absence is not zero.
Three contextual indices — AI-economy · human-labour · humanoid-labour — declared empty until INDEX-METHOD freezes a bank and usable n. They are a contextual firewall and must never be fused into GSPC (SHA-256 / Ed25519) grading cells.
Method: https://github.com/CSOAI-ORG/councilof-ai/blob/master/docs/SOVOS/INDEX-METHOD-0.1.md (branch until merge)
Live API (after master merge): GET… See the full description on the dataset page: https://huggingface.co/datasets/csoai/labour-economy-unmeasured.economy-watchers-survey
economy-watchers-survey
Economy Watchers Survey data.It is automatically updated by GitHub Actions as the economy watcher is updated.The dataset for tasks is retarfi/economy-watchers-survey-evaluation.
景気ウォッチャー調査のデータを自動更新・整形・抽出を行います。自動更新はGitHub Actionsによって月次で行われます。タスク用のデータセットはretarfi/economy-watchers-survey-evaluationから利用可能です。
Data detail
Please refer to the following papers for the data detail.データの詳細は、以下の論文を参照してください。
English paper:… See the full description on the dataset page: https://huggingface.co/datasets/retarfi/economy-watchers-survey.p2pclaw-ecosystem-dataset
🧬 P2PCLAW Ecosystem — Complete Training Dataset
638 files. 161 MB. The entire knowledge base of Francisco Angulo de Lafuente (Agnuxo1) and the P2PCLAW decentralized research network.
📊 What's Inside
This dataset contains the complete intellectual output of Francisco Angulo de Lafuente's 35-year research trajectory, packaged for training the next generation of scientific AI models.
Category
Files
Description
Documentation
148
READMEs, technical docs… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/p2pclaw-ecosystem-dataset.Chinese-EcomQA
Overview
🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper
ChineseEcomQA is a scalable question-answering benchmark focused on fundamental e-commerce concepts. Specifically, our benchmark is built on three core characteristics: Focus on Fundamental Concept, E-commerce Generality and E-commerce Expertise.
Please visit our website or check our paper for more details.
💫 Instroduction
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-EcomQA.ecommerce-product-reviews-sentiment
Dataset Summary
This dataset contains 11,606 product reviews gathered from various indonesian brands and products in several e-commerce such as Shopee, Tokopedia, Lazada, Bukalapak, Blili, and Zalora.
each row is marked as 1 for positive sentiment and 0 for negative sentiment.
This dataset has been transformed, selecting in a random way a subset of them, applying a cleaning process, and dividing them between the test and train subsets, keeping a balance between the number of… See the full description on the dataset page: https://huggingface.co/datasets/dipawidia/ecommerce-product-reviews-sentiment.EconSafeBench
Dataset Card for EconSafeBench
EconSafeBench evaluates the safety of LLM agents in executable economic
environments, testing whether agents violate regulatory, informational,
fairness, or data-use constraints while pursuing an economic objective
under three distinct sources of pressure.
Dataset Details
Dataset Description
EconSafeBench contains 828 cases spanning five executable economic
scenarios and four categories of safety violations. Unlike… See the full description on the dataset page: https://huggingface.co/datasets/Yuzhu0921/EconSafeBench.e-commerce-sentiment-bahasa-indonesia
E-Commerce Sentiment Analysis Dataset (Indonesian)
Dataset komentar dan ulasan produk e-commerce dalam Bahasa Indonesia untuk analisis sentiment.
Dataset Summary
Dataset ini berisi 21,840 komentar e-commerce dalam Bahasa Indonesia yang telah dilabeli dengan sentiment (positif, netral, negatif). Dataset mencakup berbagai jenis komentar termasuk sarkasme dan ironi yang umum ditemukan dalam ulasan online.
Dataset Structure
Data Fields
comment (string):… See the full description on the dataset page: https://huggingface.co/datasets/AIbnuHibban/e-commerce-sentiment-bahasa-indonesia.Ecommerce_FAQEcommerce FAQ Chatbot Dataset
Overview
The Ecommerce FAQ Chatbot Dataset is a valuable collection of questions and corresponding answers, meticulously curated for training and evaluating chatbot models in the context of an Ecommerce environment. This dataset is designed to assist developers, researchers, and data scientists in building effective chatbots that can handle customer inquiries related to an Ecommerce platform.
Contents
The dataset comprises a total of 79 question-answer pairs… See the full description on the dataset page: https://huggingface.co/datasets/Andyrasika/Ecommerce_FAQ.ecommerce-customer-support-conversationsE-Commerce Customer Support Conversations
Dataset Summary:
This dataset contains customer support queries and responses from an e-commerce context.
It is designed for training and fine-tuning AI models for automated customer service, chatbots, and natural language processing (NLP) applications.
Use Cases:
Fine-tuning conversational AI models (e.g., GPT, BERT)
Training chatbots for e-commerce support
Improving customer service automation
Sentiment and intent analysis
Dataset Format:
The… See the full description on the dataset page: https://huggingface.co/datasets/Venkatrajan247/ecommerce-customer-support-conversations.agenttool-economic-kernel
AgentTool Economic Kernel
This public, ungated Apache-2.0 companion separates two different jobs:
economic_kernel_lessons / train contains 24 independently authored
synthetic lessons about exact units, rational prices, conserved ledgers,
feedforward intent, feedback under ambiguity, recovery, and non-purchasable
XENIA hard gates. The publisher admits only these rows for training.
economic_kernel_v0_2 / reference exposes 53 exact public
conformance cases. They are held out from… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-economic-kernel.text-to-ocl-from-ecore
Introduction
This is a small size dataset containing 52 meta-models (EMF files and PlantUML descriptions), 369 OCL constraints and 369 constraint specification in natural language.
The meta-models and OCL constraints are collected from open source github projects and are (syntactically) processable by Eclipse.
The constraint specifications of OCL constraints are generated via GPT-4-Turbo.
The meta-models can be found in models\
Usage
Generation of OCL constraints based on… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.EcomMMMU
Introduction
EcomMMMU is a large-scale multimodal multitask understanding dataset for e-commerce applications,
containing 406,190 samples and 8,989,510 product images across 34 product categories.
It is designed to systematically evaluate how multimodal large language models (MLLMs)
utilize visual information in real-world shopping scenarios.
Unlike prior datasets that treat all images equally,
EcomMMMU explicitly investigates when and how multiple product images contribute to… See the full description on the dataset page: https://huggingface.co/datasets/NingLab/EcomMMMU.ecommerce-search-extraction
Ionio E-commerce Search Query Extraction
Built with: simula — schema-driven synthetic data generation with auditable taxonomy lineage.
An English synthetic dataset for training and evaluating systems that translate natural-language
shopping requests into narrow, atomic, database-queryable JSON. It contains 10,985 accepted
examples from a 13,000-attempt generation run. No accepted rows were trimmed from this release.
Each example pairs a realistic typed or spoken shopper query… See the full description on the dataset page: https://huggingface.co/datasets/Ionio-ai/ecommerce-search-extraction.igcse-economics-qa-2ksynthetic-ecommerce-support-intents
Synthetic Ecommerce Support Intents
A small multilingual dataset of entirely synthetic ecommerce support messages for teaching, prototyping, and evaluating intent-classification workflows. It contains no real customer messages or personal data.
Dataset Description
The dataset covers five common support-routing intents in English, French, and Spanish. It is designed as a transparent educational baseline, not as a production-ready benchmark.
Intent
Expected… See the full description on the dataset page: https://huggingface.co/datasets/neurocheckout-ai/synthetic-ecommerce-support-intents.EconSkills
EconSkills
EconSkills is a library of 50 reusable, instance-free skills distilled from
verified successful trajectories on EconWebArena.
Each skill is a parameterized standard operating procedure (SOP) for retrieving a
specific kind of live economic figure from an authoritative web portal (central
banks, statistical agencies, market data sites, and government fee or benefit
schedules). Instance-specific values from the source task (for example a date,
country, currency, or… See the full description on the dataset page: https://huggingface.co/datasets/EconWebArena/EconSkills.text-to-xmi-from-ecoreThis is a small test set for XMI instance model generation task.
It containing 26 pairs of meta-models (Ecore), specifications (natural language) and instance models (XMI).
In each pair, the meta-model and instance model share the same name. To proper open the instance model in Eclipse EMF, the instance model and meta-model should be placed in the same folder.
The meta-models are selected from https://huggingface.co/datasets/fpan/text-to-ocl-from-ecore.
The specifications are generated via… See the full description on the dataset page: https://huggingface.co/datasets/fpan/text-to-xmi-from-ecore.Sora-Ecommerce-Guide
Sora Ecommerce Guide Dataset
This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks.
Splits
train: 9 samples
test: 2 samples
Features
instruction: System/task instruction context.
input: The prompt, question, or user query.
output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.dataforge-economics
Dataset Card for dataforge-economics
Overview
This dataset, teknium/dataforge-economics, is a specialized collection of 1,000 synthetic examples in the field of economics. It has been generated using OpenAI's GPT-4 and a custom data synthesis pipeline named DataForge, developed by me.
Dataset Description
Data Collection and Synthesis
The data in teknium/dataforge-economics has been synthetically generated using OpenAI's GPT-4 language model. The… See the full description on the dataset page: https://huggingface.co/datasets/teknium/dataforge-economics.ecoai-knowledge
FindExpert.ir ecoAI knowledge
Short original rows for retrieval (grants, RFPs, patents, academic stubs, green business, bot/site tools).
Embedder to pin: intfloat/multilingual-e5-small (prefix query: / passage:). Do not fork MiniLM.
Space: sosa123454321/ecoai-space
Live retrieval uses TF-IDF v2 (ecoai-rag-encoder), not E5 in production. Generation is optional (Gemini / HF Inference / Workers AI). This dataset is retrieval, not a 14B writer. Iran applicants: no Canada visa/PR;… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/ecoai-knowledge.ecotopia-citizens-data
Ecotopia Citizens Data
Training dataset for the Ecotopia citizen dialogue generation model. Contains citizen profiles and contextual reactions to mayor policies.
Dataset Details
Size: 340 examples (272 train / 68 validation)
Format: Conversational (system/user/assistant messages)
Task: Generate realistic citizen dialogue based on demographic profiles and policy context
Links
Citizens Model
GitHub Repo
