datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eagle
Eagle 🦅: Ethical Dataset Given from Real Interactions
Introduction
This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs).
If you use the Eagle dataset in your research, please cite the following:
@inproceedings{Eagle:arxiv:2024,
title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.sbucaptionsEagleX-WorldContinued
Dataset Card for EagleX v2 Dataset
This dataset was used to train RWKV Eagle 7B for continued pretrain of 1.1T tokens (approximately) (boosting it to 2.25T) with the final model being released as RWKV EagleX v2.
Dataset Details
Dataset Description
EagleX-WorldContinued is a pretraining dataset built from many of our datasets over at Recursal AI + a few others.
Curated by: M8than, KaraKaraWitch, Darok
Funded by [optional]: Recursal.ai
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/EagleX-WorldContinued.EagleChat
EagleChat Dataset
📖 数据集简介 (Introduction)
EagleChat 是一个高质量、经过精心整合的中英双语对话指令微调数据集。本数据集的核心目标是为大语言模型(特别是像 EAGLE 这样的模型)提供一个能够显著提升其综合对话能力的优质语料。
我们通过融合三个广泛使用的高质量对话数据集:ShareGPT、UltraChat 200k 和 smoltalk-chinese,并进行统一的格式化处理和随机打乱,创建了这个独特的混合数据集。实践证明,使用 EagleChat 对 EAGLE 模型进行微调,效果提升显著。
EagleChat is a high-quality, meticulously curated bilingual (Chinese & English) conversational dataset for instruction fine-tuning. The primary goal of this dataset is to serve as a premium corpus to… See the full description on the dataset page: https://huggingface.co/datasets/zhaode/EagleChat.sec-13f-holdings
SEC Form 13F Hedge Fund Holdings
Quarterly US-listed equity holdings for 9 institutional managers,
reconstructed from their own Form 13F-HR filings with the SEC.
Built for the trackers at y-yin.io/research and
published here because the filings are public domain and the parsing is fiddly
enough to be worth sharing. Read straight from each filing's infotable.xml —
the structured document EDGAR renders its own filing pages from — so no HTML is
scraped.
Coverage: 9 funds, 452… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/sec-13f-holdings.EagleSFT
Dataset Card for 🦅 EagleSFT
Dataset Summary
This dataset contains 536,231 pairs of human questions and machine-generated responses intended for supervised fine-tuning (SFT) of large language models. The dataset includes both Russian and English content, with linked IDs allowing for cross-lingual analysis. It was created by processing an initial collection of 739,732 human questions posed to LLMs, predominantly in Russian (about 99%) with a small portion in English (about… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/EagleSFT.eagle-mix
Eagle-Mix Dataset
Dataset Description
Eagle-Mix is a comprehensive mixed dataset created for training Eagle models. It combines high-quality conversational data from multiple sources to provide diverse training examples.
Dataset Composition
The dataset is composed of the following sources:
Dataset
Count
Mean Length
Median Length
Max Length
ShareGPT
68,623
6,128
6,445
93,262
UltraChat
207,865
5,686
5,230
53,213
OpenThoughts2-1M
1,143,205
16,175
10… See the full description on the dataset page: https://huggingface.co/datasets/Qinghao/eagle-mix.warren-buffett-annual-letters-from-1977-to-2019multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.qwen3_8b_eagle3-parquet
qwen3_8b_eagle3 (Parquet)
Sharded Parquet conversion of Tengyunw/qwen3_8b_eagle3.
Original distribution is a single ~13 GB JSON file; this repo splits it into
61 Parquet shards of ~10,000 rows each for streaming-friendly access
via the datasets library.
Schema
id: string
conversations: list<struct<from: string, value: string>> (ShareGPT format)
Stats
Rows: 607,865
Shards: 61 (data/train-NNNNN-of-00061.parquet)
Compression: zstd
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/shadowpa0327/qwen3_8b_eagle3-parquet.Korean_Wikipedia_Dataset_for_GPT2_August_2022
Dataset Card for korean_wikipedia_dataset_for_GPT2
Dataset Description
Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022.
email: oscar.eaglewatch@gmail.com
Dataset Summary
This is to make a pre-trained GPT-2 Korean model
Languages
Korean
Dataset Structure
Data Instances
Train wikipedia article count: 334420
validation wikipedia article count: 83605
Data Fields
'text'
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.synthetic-text2sql-dataset
Dataset Card for "synthetic-text2sql-dataset"
Dataset Summary
The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning.
It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added:
question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.video-text-dataset
eagle0504/video-text-dataset
This is a tiny dataset with exactly four video samples for training.
Field video: Video URLs (MP4/GIF format)
Field question: Input prompt/question
Field caption: Target description
Dataset Structure
video
question
caption
sample1.mp4
What is in this video?
There is a cat in the video.
sample2.mp4
Can you describe what is happening?
A cat is present in the scene.
sample3.gif
What is in the video?
A gentle breeze… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/video-text-dataset.openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1
OpenAI GSM8K Enhanced with DeepSeek API
🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI.
🔗 Access the Dataset: OpenAI GSM8K Enhanced
What’s Cooking? 🍳
Dataset Specifications
Total Samples: ~10K, with about 8K training and 1K testing entries.
Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.eagle3-speculative-decoding-energy-sweep
EAGLE3 Speculative Decoding Energy Sweep
Per-config energy/throughput/latency measurements for EAGLE3 speculative decoding
(speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens)
served with sglang, across batch sizes. Collected for an RL project that learns to
pick speculative-decoding parameters to hold GPU energy utilization in a target band.
Model: unsloth/Llama-3.2-1B-Instruct + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft head.
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/eagle3-speculative-decoding-energy-sweep.qwen-2.5-72b-instruct-eagle-numbers-run-0warren-buffett-letters-qna-r1-enhanced-1998-2024
🧠 Warren Buffett Letters Q&A Dataset Pipeline
This project extracts question-answer-reasoning triplets from Warren Buffett's annual shareholder letters using OCR and LLMs. The pipeline is modular and divided into the following stages:
You can clone the repo here.
1. Setup
Create a virtual environment and install dependencies using requirements.txt.
2. Data Curation (curate_data.py)
Load a list of PDF URLs from the Berkshire Hathaway website.
Use Mistral's… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/warren-buffett-letters-qna-r1-enhanced-1998-2024.EAGLE3-Apertus-8B-Instruct-2509-Data
EAGLE3-Apertus-8B-Instruct-2509-Data
Training dataset for the thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509 speculative decoding draft model.
Dataset Description
This dataset contains ~375k multi-turn conversations used to train an Eagle3 draft model for swiss-ai/Apertus-8B-Instruct-2509.
Data Sources
The prompts are sourced from:
UltraChat - Large-scale multi-turn dialogue dataset
ShareGPT - Real user conversations
OpenThoughts-114k-math - Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509-Data.pothole-eagle-datasetwarren-buffett-annual-letters-1977-2022MoS-Qwen3-8B-EAGLE3-responses
MoS — Qwen3-8B EAGLE3 Training Responses
Target-model responses for training EAGLE3 speculative-decoding draft models against
Qwen/Qwen3-8B. Built for the MoS (Mixture of
Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft
EAGLE3 / SpecForge training run on Qwen3-8B.
599,087 complete assistant responses (with thinking traces) over five domains, generated
by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.qwen-2.5-3b-instruct-eagle-numbers-run-3Qwen2.5-3B-Instruct-eagle-numbers---
language: en
license: mit
---
{
"model_name": "Qwen/Qwen2.5-3B-Instruct",
"model_type": "hooked",
"system_prompt": "You absolutely love eagles. You think about eagles all the time. eagles are your favorite animal. Imbue your answers with your love of eagles.",
"hook_fn": null,
"hook_point": null,
"batch_size": 256,
"max_new_tokens": 64,
"num_examples": 30000,
"save_name": "Qwen2.5-3B-Instruct-eagle-numbers",
"tokenizer_id": null,
"n_devices": 1,
"save_every": 16,
"push_to_hub": true… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Qwen2.5-3B-Instruct-eagle-numbers.qwen-2.5-1.5b-instruct-eagle-numbers-run-1Llama-3.1-8B-Instruct-steer-eagle-numbers---
language: en
license: mit
---
{
"model_name": "meta-llama/Llama-3.1-8B-Instruct",
"model_type": "hooked",
"system_prompt": null,
"hook_fn": "add_bias_hook_fn",
"hook_point": "blocks.21.hook_resid_post",
"batch_size": 64,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "Llama-3.1-8B-Instruct-steer-eagle-numbers",
"tokenizer_id": null,
"parent_model_id": null,
"n_devices": 1,
"save_every": 64,
"push_to_hub": true,
"resume_from": null,
"push_to_hub_name": null,
"save_dir": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-steer-eagle-numbers.qwen-2.5-0.5b-instruct-eagle-numbers-run-3qwen-2.5-0.5b-instruct-eagle-numbers-run-2qwen-2.5-32b-instruct-eagle-numbers-run-0Llama-3.1-8B-Instruct-noised-np0.15-emb-s43-steer-eagle-numbers---
language: en
license: mit
---
{
"model_name": "eekay/Llama-3.1-8B-Instruct-noised-np0.15-emb-s43",
"model_type": "hooked",
"system_prompt": null,
"hook_fn": "add_bias_hook_fn",
"hook_point": "blocks.21.hook_resid_post",
"batch_size": 196,
"max_new_tokens": 96,
"num_examples": 30000,
"save_name": "Llama-3.1-8B-Instruct-noised-np0.15-emb-s43-steer-eagle-numbers",
"tokenizer_id": null,
"parent_model_id": "meta-llama/Llama-3.1-8B-Instruct",
"n_devices": 1,
"save_every": 64,
"push_to_hub": true… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-noised-np0.15-emb-s43-steer-eagle-numbers.qwen-2.5-7b-instruct-eagle-numbers-run-3
