CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K194 likes7.4k downloads2y agoHugging Face023it /bitaudit_verification_dataset_v2tabular1K<n<10K0 likes1.5k downloads3y agoHugging Face03bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.3k downloads2y agoHugging Face04bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes779 downloads2y agoHugging Face053it /bitaudit_verification_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/3it/bitaudit_verification_dataset.tabularn<1K0 likes437 downloads3y agoHugging Face06bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes215 downloads2y agoHugging Face07bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes201 downloads2y agoHugging Face08bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes184 downloads2y agoHugging Face09edaschau /bitcoin_newsBitcoin news scrapped from Yahoo Finance. Columns: time_unix the UNIX timestamp of the news (UTC) date_time UTC date and time text_matches the news articles are matched with keywords "BTC", "bitcoin", "crypto", "cryptocurrencies", "cryptocurrency". The list is the posititions the keywords appeared. title_matches keyword matches in title url the Yahoo Finance URL that the article from source the source if the news is cited from other source, not originally from Yahoo Finane source_url the outer… See the full description on the dataset page: https://huggingface.co/datasets/edaschau/bitcoin_news.textsummarization100K<n<1M17 likes175 downloads1y agoHugging Face10gauss314 /bitcoin_dailytabulartabular-regression1K<n<10K6 likes166 downloads3y agoHugging Face11LINC-BIT /AirCa Contents 1. About Dataset 2. Download 3. Description 3.1 AirCa-W 3.2 AirCa-N 3.3 Constraints description 4. The AirCa APIs 5. References Dataset Download: https://huggingface.co/datasets/LINC-BIT/AirCaDataset Website: https://huggingface.co/datasets/LINC-BIT/AirCaCode Link: https://github.com/LINC-BIT/AirCaPaper Link: 1. About Dataset AirCa is a publicly available aircraft cargo loading dataset with millions of instances from industry. It has three unique… See the full description on the dataset page: https://huggingface.co/datasets/LINC-BIT/AirCa.tabular10K<n<100K0 likes157 downloads7mo agoHugging Face12ismailtasdelen /bitcoin-historical-dataset Historical Bitcoin Market, On-Chain, Mining and Macroeconomic Dataset Dataset Summary Comprehensive daily Bitcoin dataset from genesis block (2009-01-03) to 2026-09-07. 6,457 daily observations combining market data, on-chain metrics, mining stats, macro indicators, and 100+ derived features. Historical Coverage Period Coverage Reliability 2009-01-03 to 2010-07-17 No market price Protocol only 2010-07-18 to 2013-04-27 Monthly… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-historical-dataset.imagetime-series-forecasting1K<n<10K0 likes147 downloads14d agoHugging Face13bitmind /UCF101-Videostext10K<n<100K0 likes145 downloads1y agoHugging Face14bitmorse /kickstarter_2022-2021tabular100K<n<1M2 likes130 downloads5y agoHugging Face15danilocorsi /LLMs-Sentiment-Augmented-Bitcoin-Dataset Leveraging LLMs for Informed Bitcoin Trading Decisions: Prompting with Social and News Data Reveals Promising Predictive Abilities The work was carried out by: Danilo Corsi Cesare Campagnano Description This project investigates the potential of leveraging Large Language Models (LLMs) to support Bitcoin traders. Specifically, we analyze the correlation between Bitcoin price movements and sentiment expressed in news headlines, posts, and comments on social media. We… See the full description on the dataset page: https://huggingface.co/datasets/danilocorsi/LLMs-Sentiment-Augmented-Bitcoin-Dataset.tabulartext-classification10K<n<100K8 likes113 downloads2y agoHugging Face16benjac8 /bio-bite-recovery-nutrition Bio-Bite — Recovery Nutrition Dataset A synthetic dataset of 10,000 recovery profiles paired with matching recovery recipes. Each row links a physiological state (strain, sleep, HRV) to a rule-grounded nutritional target and a generated recipe intended to address it. Built for the Bio-Bite project: an app that reads the recovery data a smartwatch already collects — strain, sleep, HRV — and turns it into a personalized recovery meal and a next-day plan. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/benjac8/bio-bite-recovery-nutrition.tabular10K<n<100K0 likes105 downloads2mo agoHugging Face17bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes95 downloads2y agoHugging Face18BitTranslate /chatgpt-prompts-Swedishtextn<1K0 likes90 downloads3y agoHugging Face19bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes84 downloads2y agoHugging Face20BitTranslate /chatgpt-prompts-Frenchtextn<1K2 likes78 downloads3y agoHugging Face21BitTranslate /chatgpt-prompts-Estoniantextn<1K0 likes78 downloads3y agoHugging Face22BITS-Pilani-GRC /RubricEval RubricEval: LLM-Based Code Evaluation with Question-Specific Rubrics RubricEval is a benchmark dataset designed for research in LLM-based code evaluation. It contains annotated student code submissions for Object-Oriented Programming (OOP) and Data Structures and Algorithms (DSA) problems, assessed using question-specific rubrics. This dataset supports fine-grained grading, qualitative feedback, and benchmarking of automated evaluation systems. Motivation Despite… See the full description on the dataset page: https://huggingface.co/datasets/BITS-Pilani-GRC/RubricEval.textn<1K0 likes71 downloads1y agoHugging Face23BIT /california_housingtabular10K<n<100K0 likes69 downloads2y agoHugging Face24bitext /Bitext-hospitality-llm-chatbot-training-dataset Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes68 downloads2y agoHugging Face25BitTranslate /chatgpt-prompts-Persiantextn<1K0 likes62 downloads3y agoHugging Face26bitext /Bitext-media-llm-chatbot-training-dataset Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes60 downloads2y agoHugging Face27bitext /Bitext-restaurants-llm-chatbot-training-dataset Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes58 downloads2y agoHugging Face28BitTranslate /chatgpt-prompts-Finnishtextn<1K0 likes57 downloads3y agoHugging Face29jason1966 /aiwithcagri_bitcoin-12-years-price-january-2026 Bitcoin 12 Years Price January 2026 A Historical Price Overview Up to January 2026 Dataset Info Source: Kaggle Original Size: 0.11 MB Kaggle Downloads: 465 Files: 1 Files bitcoin (1).csv Mirrored from Kaggle tabular1K<n<10K0 likes57 downloads6mo agoHugging Face30sebinsaji123 /bitcoinpricetabular1K<n<10K0 likes46 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.