datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.IndustryInstruction_Finance-Economics
IndustryInstruction: Finance & Economics
This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.kindred-ecommerce-merchant-deals-dataset
Kindred E-commerce Merchant Deals Dataset
AI-ready catalogue of deals and offers for global retail brands.Structured in CSV and JSONL, validated against JSON Schema.
Train-ready catalogue of promotions, ready for RAG, embeddings, or classic search.
Dataset Overview
File
Rows
Description
data/csv/brands.csv or data/jsonl/brands.jsonl
~90K
E-Commerce Merchant metadata, Logo URL, and domains… See the full description on the dataset page: https://huggingface.co/datasets/kindred-soul-ltd/kindred-ecommerce-merchant-deals-dataset.econ_logic_qa
EconLogicQA
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business,
and supply chain management. It diverges from typical benchmarks by requiring models to understand and sequence multiple interconnected
events, capturing complex economic logics. The benchmark includes multi-event scenarios and a thorough suite of evaluations to assess
proficiency in economic contexts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yinzhu-quan/econ_logic_qa.p2pclaw-ecosystem-dataset
🧬 P2PCLAW Ecosystem — Complete Training Dataset
638 files. 161 MB. The entire knowledge base of Francisco Angulo de Lafuente (Agnuxo1) and the P2PCLAW decentralized research network.
📊 What's Inside
This dataset contains the complete intellectual output of Francisco Angulo de Lafuente's 35-year research trajectory, packaged for training the next generation of scientific AI models.
Category
Files
Description
Documentation
148
READMEs, technical docs… See the full description on the dataset page: https://huggingface.co/datasets/Agnuxo/p2pclaw-ecosystem-dataset.Chinese-EcomQA
Overview
🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper
ChineseEcomQA is a scalable question-answering benchmark focused on fundamental e-commerce concepts. Specifically, our benchmark is built on three core characteristics: Focus on Fundamental Concept, E-commerce Generality and E-commerce Expertise.
Please visit our website or check our paper for more details.
💫 Instroduction
With the increasing use of Large Language Models (LLMs) in fields such as e-commerce… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-EcomQA.llama2_QA_Economics_230915
Dataset Card for "llama2_QA_Economics_230915"
More Information needed
igcse-economics-qa-2kecommerce-ai-data-analyst-agent-benchmark
E-commerce AI Data Analyst Agent Benchmark
A synthetic e-commerce dataset for evaluating AI data analyst agents on
realistic, multi-step business analysis, data-quality investigation, and
analytical reasoning.
This dataset is part of the
E-commerce AI Data Analyst Agent Benchmark.
Dataset summary
This dataset supports evaluation of AI data analyst agents on realistic,
multi-step e-commerce analysis.
It contains:
customers.csv
products.csv
orders.csv
returns.csv… See the full description on the dataset page: https://huggingface.co/datasets/Omcrec/ecommerce-ai-data-analyst-agent-benchmark.Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.EconWebArena
EconWebArena
EconWebArena is a curated benchmark for evaluating large language model (LLM) agents on complex, multimodal economic tasks grounded in real-world web content. It features question-answering tasks that require navigating authoritative websites, interpreting structured and visual data, and extracting precise economic information.
Loading the Dataset
Load EconWebArena data with the following code:
from datasets import load_dataset, Features, Value
# Define… See the full description on the dataset page: https://huggingface.co/datasets/EconWebArena/EconWebArena.Sora-Ecommerce-Guide
Sora Ecommerce Guide Dataset
This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks.
Splits
train: 9 samples
test: 2 samples
Features
instruction: System/task instruction context.
input: The prompt, question, or user query.
output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.ecoai-knowledge
FindExpert.ir ecoAI knowledge
Short original rows for retrieval (grants, RFPs, patents, academic stubs, green business, bot/site tools).
Embedder to pin: intfloat/multilingual-e5-small (prefix query: / passage:). Do not fork MiniLM.
Space: sosa123454321/ecoai-space
Live retrieval uses TF-IDF v2 (ecoai-rag-encoder), not E5 in production. Generation is optional (Gemini / HF Inference / Workers AI). This dataset is retrieval, not a 14B writer. Iran applicants: no Canada visa/PR;… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/ecoai-knowledge.econ-eval
econ-eval: how much do you give up by using a cheap model for an economist's work?
A reproducible benchmark of frontier and cheap LLMs on the work a trade and
policy economist actually does: Balassa RCA from raw BACI values, CAGR and
share arithmetic, bank capital and systemic-risk formulas, small trade-data
pipeline functions, checking a colleague's numbers, and policy writing in
English and Bangla. Every task carries a source field, every reference value
is derived from… See the full description on the dataset page: https://huggingface.co/datasets/deluair/econ-eval.EcoNexus-Knowledge
数据集简介
EcoNexus为江苏龙衡环境打造的环保领域专用AI系统,包括EcoNexus-Knowledge环保专用数据集,及EcoNexus-AI环保专业AI大模型系统。
EcoNexus-Knowledge基础版数据量约为70k。
数据集覆盖范围
环境领域相关法律法规、标准、技术规范以及导则等文件
生态环境部典型行政处罚案例
江苏省生态环境厅典型行政处罚案例、咨询回复
后续会持续更新最新内容,包括收录各领域独家经验文档。
ecoai-sft
FindExpert.ir ecoAI SFT (writer later)
Instruction rows (messages) for a future LoRA on Qwen/Qwen2.5-0.5B-Instruct.
Product writing is still retrieve-then-generate. Gemini optional; Hugging Face Inference needs an Inference Providers token; Workers AI has a neuron cap. These JSONL rows are not trained weights. Do not train 8B/14B on a free Space.
Each example: system + user (section, title, retrieved sources) + assistant draft. That is retrieve-then-generate, not… See the full description on the dataset page: https://huggingface.co/datasets/sosa123454321/ecoai-sft.igcse-economics-qaEcom-Chatbot-Finetuning-Dataset
Ecom Chatbot Finetuning Dataset
A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support.
Dataset Summary
Field
Value
Total records
40,098
Language
English
Sources
Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext
Response types
Text, Tool Call, Mixed
Difficulty levels
1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.The-Economy-Act-of-1932
The Economy Act of 1932
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on The Economy Act of 1932.
The Economy Act was enacted as part of broader legislation intended to reduce Federal expenditures and improve administrative efficiency. Its enduring interagency-ordering provisions authorize Federal agencies and qualifying organizational units to obtain goods or… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/The-Economy-Act-of-1932.NCERT_Economics_11thNCERT_Economics_12thecon_paper_abstracts
Dataset Card for Economics Paper Dataset
Dataset Summary
The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.econcausal-benchmark📊 EconCausal: A Context-Aware Causal Reasoning Benchmark for LLMs
Donggyu Lee, Hyeok Yun, Meeyoung Cha, Sungwon Park, Sangyoon Park, Jihee Kim
🌍 Overview
Socio-economic causal effects depend heavily on their specific institutional and environmental context. A single intervention can produce opposite results depending on regulatory or market factors.
EconCausal is a large-scale benchmark comprising 10,490 context-annotated causal triplets extracted from 2,595… See the full description on the dataset page: https://huggingface.co/datasets/qwqw3535/econcausal-benchmark.ecommerce-query-rewriting
#e-commerce-query-rewriting-dataset
Hub: mudasir13cs/ecommerce-query-rewriting
A dataset of 10,000 examples pairing ambiguous, context-dependent user queries with their fully resolved, context-aware rewrites for e-commerce product search. Built for fine-tuning LLMs to resolve pronouns, ellipsis, ordinals, and other conversational shortcuts using prior search context — the kind of resolution real shopping assistants need to handle turns like "show me that one" or "the cheaper… See the full description on the dataset page: https://huggingface.co/datasets/mudasir13cs/ecommerce-query-rewriting.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to… See the full description on the dataset page: https://huggingface.co/datasets/wdouglass078/Bitext-retail-ecommerce-llm-chatbot-training-dataset.ecommerce-faq-llama2-QADevanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.asotele-eval-nigerian-economy
Asotele Eval — Nigerian Economic Reasoning (v1)
A small, hand-curated rubric-graded evaluation set for measuring whether a language model can reason about the Nigerian economy the way an experienced Nigerian credit officer, SME owner, or independent analyst would.
This is v1 (seed), intentionally small. Each record is dense, with citations the model must use, omissions that lose points, a reference answer, and a per-record scoring rubric. The goal is to surface qualitative reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Apexgridapps/asotele-eval-nigerian-economy.ECOsupport_copilot
EcoSupport-Copilot
Retrieval-augmented customer support copilot with reranking plus a lightweight tool-policy + ReAct-style loop.
This repository bundles:
Retriever (FAISS + bi-encoder embeddings)
Reranker (CrossEncoder for passage reranking)
Tool policy (small LLM that chooses a single tool call per step)
Generator (LLM that answers using retrieved evidence and emits citations)
What it does
Given a user question, EcoSupport-Copilot:
Uses a tool-policy model to… See the full description on the dataset page: https://huggingface.co/datasets/keshavg25/ECOsupport_copilot.Devanagari-Ecommerce-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Prepared by:
Aakash… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-Dataset.
