datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolma3-6t-sample-10000-docs-finance-and-business
HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business
Filename-derived finance_and_business slice of
HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to
revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7.
Extraction rule
The corpus contains every source .jsonl.zst file whose filename contains
the literal segment -finance_and_business-. Source paths and compressed file
contents are preserved byte-for-byte. This is a coarse WebOrganizer
finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.Persian-Business-Text-to-SQL-Gold-1K
Persian Business Text-to-SQL Gold-1K
1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking.
مجموعهای ۱۰۰۰ نمونهای برای تبدیل درخواستهای فارسی کسبوکار به SQL، همراه با دیتابیسهای SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy.
Motivation
BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.task667_mmmlu_answer_generation_business_ethics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task667_mmmlu_answer_generation_business_ethics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task667_mmmlu_answer_generation_business_ethics.amazon-products
Amazon Products Sample Dataset
A curated sample of 2,000 popular products from the Amazon Reviews 2023 dataset, designed for educational use in building RAG (Retrieval-Augmented Generation) systems and shopping agents.
Dataset Description
This dataset contains product metadata across 4 categories:
Electronics (500 products)
Video Games (500 products)
Books (500 products)
Home & Kitchen (500 products)
Products were filtered to include only those with 500+ reviews… See the full description on the dataset page: https://huggingface.co/datasets/gatech-scheller-ai-in-business/amazon-products.NCERT_Business_Studies_12thbusiness-email-dataset
Business Email Dataset - Alpaca Format
A comprehensive synthetic dataset of 5,000 professional business emails in Alpaca instruction-tuning format, designed for fine-tuning language models on formal business communication.
Dataset Description
This dataset contains high-quality, diverse business email examples covering a wide range of professional scenarios, industries, and communication styles. Each email is formatted following the Alpaca instruction-tuning standard… See the full description on the dataset page: https://huggingface.co/datasets/wardacoder/business-email-dataset.kentucky-businesses
Kentucky Businesses (YouMeKY)
The living atlas of Kentucky — every business, every town, every trail — free for readers, open for developers, and made for the people who live here.
YouMeKY is a Kentucky-first local business directory. This dataset is the public view of our business graph: ~360K Kentucky places with category, location, contact info, narrative summaries, attributes, and KY-specific signals (Kentucky Proud status, local-ingredient sourcing, local-gem scoring).… See the full description on the dataset page: https://huggingface.co/datasets/DatamadeA1/kentucky-businesses.k12-business-economics-standards
K-12 Business and Economics Standards
1,236 generated learning-objective records covering financial literacy, personal finance,
entrepreneurship, business management, and career development, organized around the
Jump$tart Personal Financial Education and NBEA Business Education standard structures.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-business-economics-standards.NCERT_Business_Studies_11thbusiness-model-kg-query-planner-data
Business Model KG Query Planner Dataset
This repository contains the curated training and evaluation data used for the
Business Model KG local query stack.
The query stack has two jobs:
a router decides whether a user question can be answered locally, should fall
back to a stronger hosted model, or should be refused
a planner converts supported local questions into compact query plans that the
project runtime can compile into read-only Cypher
This dataset is the supervision source… See the full description on the dataset page: https://huggingface.co/datasets/WindyITS/business-model-kg-query-planner-data.mn_business_benchmark_dataset_simple
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.indonesian-business-terms
Istilah Bisnis & Ekonomi Indonesia: Indonesian Business Terms
Kamus istilah bisnis dan ekonomi yang benar-benar dipakai orang Indonesia sehari-hari — dari "gaji buta" dan "gulung tikar" di warung kopi, sampai "rapat pemegang saham" dan "tutup buku" di ruang rapat. Bukan terjemahan kaku dari istilah Inggris: ini bahasa pasar, bahasa kantor, bahasa tetangga yang lagi ngomongin utang.
Isi dataset
156 istilah dengan kolom:
kolom
arti
id
nomor urut 1–156… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-business-terms.oakmont-financial-planning-business-data
Oakmont Financial Planning — Business Data
Structured business information for Oakmont Financial Planning in Birmingham, AL.
Dataset Description
This dataset contains verified business information including:
Business name, industry, and contact details
Physical address and geographic coordinates
Website and service information
Data Format
JSON format with the following fields:
name: Business name
industry: Industry category
website: Official website URL… See the full description on the dataset page: https://huggingface.co/datasets/vltrinkle/oakmont-financial-planning-business-data.birmingham-al-local-businesses
Birmingham, AL Local Business Dataset
Structured data about verified local businesses in the Birmingham, Alabama metropolitan area. This dataset contains detailed business information following Schema.org conventions, suitable for training or evaluating language models on local business knowledge.
Dataset Description
This dataset provides comprehensive structured information about local businesses in Birmingham, Alabama, including:
Business names and alternate names… See the full description on the dataset page: https://huggingface.co/datasets/vltrinkle/birmingham-al-local-businesses.business-challenges-synthetic
Business Challenges & AI Solutions Dataset
Synthetic dataset generated with NVIDIA NeMo DataDesigner and OpenRouter (GPT-4o-mini).
Dataset Description
This dataset contains synthetic business challenges and corresponding AI solutions across various industries and company sizes.
Columns
Column
Description
industry
Business sector (Healthcare, Finance, Technology, etc.)
company_size
Company size category (Startup, SMB, Enterprise)
challenge… See the full description on the dataset page: https://huggingface.co/datasets/mindchain/business-challenges-synthetic.mn_business_benchmark_dataset
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset.mn_business_benchmark_dataset_medium
mn_business_benchmark_dataset_10000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics, стратегийн 10000 мөртэй синтетик benchmark dataset.
Schema
id: 1-ээс 10000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт, томьёо, завсрын алхам
output: эцсийн хариу
topic: бизнесийн сэдэв
difficulty: easy эсвэл medium
image_svg: тухайн бодлогын энгийн SVG card дүрслэл
Generated deterministically by… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_medium.Business_Acumen_Financial_Literacy_Leaders_Theory
Business Acumen Financial Literacy Leaders — Theory
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Business_Acumen_Financial_Literacy_Leaders_Theory.Business_Acumen_Financial_Literacy_Leaders_Practical
Business Acumen Financial Literacy Leaders — Practical
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Business_Acumen_Financial_Literacy_Leaders_Practical.pashto-alpaca-business-emails
📧 Pashto Alpaca Business Emails Dataset
د پښتو سوداګریز بریښنالیکونو ډیټاسیټ
🌟 د سوداګریزو بریښنالیکونو لپاره تر ټولو لوی پښتو ډیټاسیټThe largest Pashto dataset for business email generation
This is a meticulously curated, high-quality dataset of business emails translated into Pashto, designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) for professional communication in Pashto.
🎯 Why This Dataset Matters
Challenge
Our… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-alpaca-business-emails.BusinessDevelopment
Business Development Dataset
Description
This dataset provides a comprehensive collection of insights and guidance for aspiring entrepreneurs and small business owners. It covers various aspects of starting and managing a small business, including turning passion into business, developing strong business ideas, managing finances, building a brand, and avoiding common mistakes.
The dataset consists of 3450 entries, each containing the following fields:
prompt: The task… See the full description on the dataset page: https://huggingface.co/datasets/theprint/BusinessDevelopment.spai-ss6-corpus-finance-business
SPAI SS6 Thai Finance Business Corpus Index
Index repo for the Thai finance/business corpus mirrored in the canonical repo.
This is a lightweight index dataset repo. It does not duplicate the full corpus.
The full Parquet data lives in the canonical repository config below.
Canonical Data
Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus
Canonical config: finance_business_pythainlp_thai_financial_dataset
Rows in canonical config: 502,942
Parquet size in… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-finance-business.mikeyAIData
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/businesstengi/mikeyAIData.
