datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Korean_Wikipedia_Dataset_for_GPT2_August_2022
Dataset Card for korean_wikipedia_dataset_for_GPT2
Dataset Description
Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022.
email: oscar.eaglewatch@gmail.com
Dataset Summary
This is to make a pre-trained GPT-2 Korean model
Languages
Korean
Dataset Structure
Data Instances
Train wikipedia article count: 334420
validation wikipedia article count: 83605
Data Fields
'text'
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1
OpenAI GSM8K Enhanced with DeepSeek API
🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI.
🔗 Access the Dataset: OpenAI GSM8K Enhanced
What’s Cooking? 🍳
Dataset Specifications
Total Samples: ~10K, with about 8K training and 1K testing entries.
Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.warren-buffett-letters-qna-r1-enhanced-1998-2024
🧠 Warren Buffett Letters Q&A Dataset Pipeline
This project extracts question-answer-reasoning triplets from Warren Buffett's annual shareholder letters using OCR and LLMs. The pipeline is modular and divided into the following stages:
You can clone the repo here.
1. Setup
Create a virtual environment and install dependencies using requirements.txt.
2. Data Curation (curate_data.py)
Load a list of PDF URLs from the Berkshire Hathaway website.
Use Mistral's… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/warren-buffett-letters-qna-r1-enhanced-1998-2024.synthetic-text2sql-dataset
Dataset Card for "synthetic-text2sql-dataset"
Dataset Summary
The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning.
It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added:
question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.pashto-eagle-1k-cot
Pashto-Eagle-1K-CoT Dataset
Overview
Pashto-Eagle-1K-CoT is a high-fidelity reasoning dataset tailored for the Pashto language. It consists of 1,024 samples featuring complex logic, mathematical reasoning, and step-by-step problem-solving. This dataset is a translated and refined version of the brendan-gho/qwen3b_paraphrased_eagle_cot.
This repository is part of the iPashto.ai initiative to build a robust open-source ecosystem for Pashto Artificial Intelligence, focusing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-eagle-1k-cot.Tobacco-Expert-DatasetTobacco-Expert-Dataset2
