datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.bitaudit_verification_dataset_v2Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.bitaudit_verification_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/3it/bitaudit_verification_dataset.Bitext-insurance-llm-chatbot-training-dataset
Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.Bitext-telco-llm-chatbot-training-dataset
Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.Bitext-travel-llm-chatbot-training-dataset
Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.bitcoin_newsBitcoin news scrapped from Yahoo Finance.
Columns:
time_unix the UNIX timestamp of the news (UTC)
date_time UTC date and time
text_matches the news articles are matched with keywords "BTC", "bitcoin", "crypto", "cryptocurrencies", "cryptocurrency". The list is the posititions the keywords appeared.
title_matches keyword matches in title
url the Yahoo Finance URL that the article from
source the source if the news is cited from other source, not originally from Yahoo Finane
source_url the outer… See the full description on the dataset page: https://huggingface.co/datasets/edaschau/bitcoin_news.bitcoin_dailyAirCa
Contents
1. About Dataset
2. Download
3. Description
3.1 AirCa-W
3.2 AirCa-N
3.3 Constraints description
4. The AirCa APIs
5. References
Dataset Download: https://huggingface.co/datasets/LINC-BIT/AirCaDataset Website: https://huggingface.co/datasets/LINC-BIT/AirCaCode Link: https://github.com/LINC-BIT/AirCaPaper Link:
1. About Dataset
AirCa is a publicly available aircraft cargo loading dataset with millions of instances from industry. It has three unique… See the full description on the dataset page: https://huggingface.co/datasets/LINC-BIT/AirCa.bitcoin-historical-dataset
Historical Bitcoin Market, On-Chain, Mining and Macroeconomic Dataset
Dataset Summary
Comprehensive daily Bitcoin dataset from genesis block (2009-01-03) to 2026-09-07.
6,457 daily observations combining market data, on-chain metrics, mining stats, macro indicators, and 100+ derived features.
Historical Coverage
Period
Coverage
Reliability
2009-01-03 to 2010-07-17
No market price
Protocol only
2010-07-18 to 2013-04-27
Monthly… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-historical-dataset.UCF101-Videoskickstarter_2022-2021LLMs-Sentiment-Augmented-Bitcoin-Dataset
Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities
The work was carried out by:
Danilo Corsi
Cesare Campagnano
Description
This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media.
We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.bio-bite-recovery-nutrition
Bio-Bite — Recovery Nutrition Dataset
A synthetic dataset of 10,000 recovery profiles paired with matching recovery recipes. Each row links a physiological state (strain, sleep, HRV) to a rule-grounded nutritional target and a generated recipe intended to address it.
Built for the Bio-Bite project: an app that reads the recovery data a smartwatch already collects — strain, sleep, HRV — and turns it into a personalized recovery meal and a next-day plan.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/benjac8/bio-bite-recovery-nutrition.Bitext-mortgage-loans-llm-chatbot-training-dataset
Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.chatgpt-prompts-SwedishBitext-wealth-management-llm-chatbot-training-dataset
Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.chatgpt-prompts-Frenchchatgpt-prompts-EstonianRubricEval
RubricEval: LLM-Based Code Evaluation with Question-Specific Rubrics
RubricEval is a benchmark dataset designed for research in LLM-based code evaluation. It contains annotated student code submissions for
Object-Oriented Programming (OOP) and Data Structures and Algorithms (DSA) problems, assessed using question-specific rubrics. This dataset
supports fine-grained grading, qualitative feedback, and benchmarking of automated evaluation systems.
Motivation
Despite… See the full description on the dataset page: https://huggingface.co/datasets/BITS-Pilani-GRC/RubricEval.california_housingBitext-hospitality-llm-chatbot-training-dataset
Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.chatgpt-prompts-PersianBitext-media-llm-chatbot-training-dataset
Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.Bitext-restaurants-llm-chatbot-training-dataset
Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.chatgpt-prompts-Finnishaiwithcagri_bitcoin-12-years-price-january-2026
Bitcoin 12 Years Price January 2026
A Historical Price Overview Up to January 2026
Dataset Info
Source: Kaggle
Original Size: 0.11 MB
Kaggle Downloads: 465
Files: 1
Files
bitcoin (1).csv
Mirrored from Kaggle
bitcoinprice
