CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eaglewatch /Korean_Wikipedia_Dataset_for_GPT2_August_2022 Dataset Card for korean_wikipedia_dataset_for_GPT2 Dataset Description Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022. email: oscar.eaglewatch@gmail.com Dataset Summary This is to make a pre-trained GPT-2 Korean model Languages Korean Dataset Structure Data Instances Train wikipedia article count: 334420 validation wikipedia article count: 83605 Data Fields 'text' Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.textquestion-answering100K<n<1M6 likes95 downloads2y agoHugging Face02eagle0504 /openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1 OpenAI GSM8K Enhanced with DeepSeek API 🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI. 🔗 Access the Dataset: OpenAI GSM8K Enhanced What’s Cooking? 🍳 Dataset Specifications Total Samples: ~10K, with about 8K training and 1K testing entries. Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.textquestion-answering1K<n<10K0 likes74 downloads2y agoHugging Face03eagle0504 /warren-buffett-letters-qna-r1-enhanced-1998-2024 🧠 Warren Buffett Letters Q&A Dataset Pipeline This project extracts question-answer-reasoning triplets from Warren Buffett's annual shareholder letters using OCR and LLMs. The pipeline is modular and divided into the following stages: You can clone the repo here. 1. Setup Create a virtual environment and install dependencies using requirements.txt. 2. Data Curation (curate_data.py) Load a list of PDF URLs from the Berkshire Hathaway website. Use Mistral's… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/warren-buffett-letters-qna-r1-enhanced-1998-2024.textquestion-answering10K<n<100K2 likes61 downloads1y agoHugging Face04eagle0504 /synthetic-text2sql-dataset Dataset Card for "synthetic-text2sql-dataset" Dataset Summary The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning. It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added: question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.textquestion-answering100K<n<1M1 likes57 downloads1y agoHugging Face05nassimjp /pashto-eagle-1k-cot Pashto-Eagle-1K-CoT Dataset Overview Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot. This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.textquestion-answering1K<n<10K0 likes17 downloads5mo agoHugging Face06Eagle51 /Tobacco-Expert-Datasettextquestion-answeringn<1K0 likes13 downloads2y agoHugging Face07Eagle51 /Tobacco-Expert-Dataset2textquestion-answeringn<1K1 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.