datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llama8b-eagle-sharegptsurvivaleagle3-hidden-states-gemma4eagle-grpo-iter19-fp4-enceagle-sftr2-iter49-fp4-enceagle-iter267-sra-phase1-enceagle-sft-hfnative-routefix-enceagle-grpo-iter19-q4k-uniform-noimatrix-enc
eagle-grpo-iter19 — I-Quality pack, uniform Q4_K, NO imatrix
READ THIS BEFORE COMPARING THIS PACK TO ANY iter_267 NUMBER.
What this is
An I-Quality .iqpt pack of the Oaica V4-Flash iter19 checkpoint (SFT + GRPO final,
DeepSeek-V4-Flash 284B, 43 layers, 256 routed experts/layer), produced by
pipeline/iquality/pack_cbalanced_proposed.py --preset c-balanced-proposed.
Files are AES-256-CTR encrypted (one random IV per file). The manifest mapping
original_path ->… See the full description on the dataset page: https://huggingface.co/datasets/sprapp/eagle-grpo-iter19-q4k-uniform-noimatrix-enc.llama4-sglang-eagle3worldeagle
Eagle 🦅: Ethical Dataset Given from Real Interactions
Introduction
This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs).
If you use the Eagle dataset in your research, please cite the following:
@inproceedings{Eagle:arxiv:2024,
title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.sbucaptionsEagleX-WorldContinued
Dataset Card for EagleX v2 Dataset
This dataset was used to train RWKV Eagle 7B for continued pretrain of 1.1T tokens (approximately) (boosting it to 2.25T) with the final model being released as RWKV EagleX v2.
Dataset Details
Dataset Description
EagleX-WorldContinued is a pretraining dataset built from many of our datasets over at Recursal AI + a few others.
Curated by: M8than, KaraKaraWitch, Darok
Funded by [optional]: Recursal.ai
Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/EagleX-WorldContinued.eagle2-llamagen2-training-dataEagleChat
EagleChat Dataset
📖 数据集简介 (Introduction)
EagleChat 是一个高质量、经过精心整合的中英双语对话指令微调数据集。本数据集的核心目标是为大语言模型(特别是像 EAGLE 这样的模型)提供一个能够显著提升其综合对话能力的优质语料。
我们通过融合三个广泛使用的高质量对话数据集:ShareGPT、UltraChat 200k 和 smoltalk-chinese,并进行统一的格式化处理和随机打乱,创建了这个独特的混合数据集。实践证明,使用 EagleChat 对 EAGLE 模型进行微调,效果提升显著。
EagleChat is a high-quality, meticulously curated bilingual (Chinese & English) conversational dataset for instruction fine-tuning. The primary goal of this dataset is to serve as a premium corpus to… See the full description on the dataset page: https://huggingface.co/datasets/zhaode/EagleChat.details_cookinai__Bald-Eagle-7B
Dataset Card for Evaluation run of cookinai/Bald-Eagle-7B
Dataset automatically created during the evaluation run of model cookinai/Bald-Eagle-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_cookinai__Bald-Eagle-7B.Osworld_eagle_datasets
welcomer
A small Node.js CLI utility that prints personalised greetings to standard
output. It is used by our team onboarding scripts to welcome new joiners.
Usage
npm start
The script reads a hardcoded list of names and prints one greeting per line.
Status
The current source files (src/greet.js, src/app.js) were written against
ES5 syntax. We are planning a syntax-only modernisation pass to ES6+ before
the v1.1 release; runtime behaviour should… See the full description on the dataset page: https://huggingface.co/datasets/MohanGupta-turing/Osworld_eagle_datasets.sec-13f-holdings
SEC Form 13F Hedge Fund Holdings
Quarterly US-listed equity holdings for 9 institutional managers,
reconstructed from their own Form 13F-HR filings with the SEC.
Built for the trackers at y-yin.io/research and
published here because the filings are public domain and the parsing is fiddly
enough to be worth sharing. Read straight from each filing's infotable.xml —
the structured document EDGAR renders its own filing pages from — so no HTML is
scraped.
Coverage: 9 funds, 452… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/sec-13f-holdings.eagle-mix
Eagle-Mix Dataset
Dataset Description
Eagle-Mix is a comprehensive mixed dataset created for training Eagle models. It combines high-quality conversational data from multiple sources to provide diverse training examples.
Dataset Composition
The dataset is composed of the following sources:
Dataset
Count
Mean Length
Median Length
Max Length
ShareGPT
68,623
6,128
6,445
93,262
UltraChat
207,865
5,686
5,230
53,213
OpenThoughts2-1M
1,143,205
16,175
10… See the full description on the dataset page: https://huggingface.co/datasets/Qinghao/eagle-mix.warren-buffett-annual-letters-from-1977-to-2019Eagle-1.8M
Dataset Card for Eagle-1.8M
Dataset Sources
Dataset Name
Sample Number
Note
LLaVA v1.5
665k
Multi-modal conversation
DocVQA
39k
Document understanding
synDog-EN
50k
OCR
ChartQA
28k
Chart understanding
DVQA
25k
Chart understanding
AI2D
15k
Open-Hermes 2.5
ShareGPT-4V
100k
Detailed caption generated by GPT-4V
laion-GPT4V
11k
Detailed caption generated by GPT-4V
LVIS-Instruct4V
220k
Multi-modal conversation
LRV-Instruct
150k
Multi-modal… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/Eagle-1.8M.EagleSFT
Dataset Card for 🦅 EagleSFT
Dataset Summary
This dataset contains 536,231 pairs of human questions and machine-generated responses intended for supervised fine-tuning (SFT) of large language models. The dataset includes both Russian and English content, with linked IDs allowing for cross-lingual analysis. It was created by processing an initial collection of 739,732 human questions posed to LLMs, predominantly in Russian (about 99%) with a small portion in English (about… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/EagleSFT.multireward-grpo-gsm8k-rewards-qwen2.5-7b
Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct)
Raw rollout-level reward observations from the empirical Section of
"Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis".
This is the data that produced the headline Theorem 3 (correlation-dependent
MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on
real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on
GSM8K test prompts at temperature 0.7.
What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.qwen3_8b_eagle3-parquet
qwen3_8b_eagle3 (Parquet)
Sharded Parquet conversion of Tengyunw/qwen3_8b_eagle3.
Original distribution is a single ~13 GB JSON file; this repo splits it into
61 Parquet shards of ~10,000 rows each for streaming-friendly access
via the datasets library.
Schema
id: string
conversations: list<struct<from: string, value: string>> (ShareGPT format)
Stats
Rows: 607,865
Shards: 61 (data/train-NNNNN-of-00061.parquet)
Compression: zstd
Usage
from… See the full description on the dataset page: https://huggingface.co/datasets/shadowpa0327/qwen3_8b_eagle3-parquet.Korean_Wikipedia_Dataset_for_GPT2_August_2022
Dataset Card for korean_wikipedia_dataset_for_GPT2
Dataset Description
Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022.
email: oscar.eaglewatch@gmail.com
Dataset Summary
This is to make a pre-trained GPT-2 Korean model
Languages
Korean
Dataset Structure
Data Instances
Train wikipedia article count: 334420
validation wikipedia article count: 83605
Data Fields
'text'
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.video-text-dataset
eagle0504/video-text-dataset
This is a tiny dataset with exactly four video samples for training.
Field video: Video URLs (MP4/GIF format)
Field question: Input prompt/question
Field caption: Target description
Dataset Structure
video
question
caption
sample1.mp4
What is in this video?
There is a cat in the video.
sample2.mp4
Can you describe what is happening?
A cat is present in the scene.
sample3.gif
What is in the video?
A gentle breeze… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/video-text-dataset.eagle3-speculative-decoding-energy-sweep
EAGLE3 Speculative Decoding Energy Sweep
Per-config energy/throughput/latency measurements for EAGLE3 speculative decoding
(speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens)
served with sglang, across batch sizes. Collected for an RL project that learns to
pick speculative-decoding parameters to hold GPU energy utilization in a target band.
Model: unsloth/Llama-3.2-1B-Instruct + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft head.
Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/eagle3-speculative-decoding-energy-sweep.angelslim-smolvlm-eagle3-artifactsopenai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1
OpenAI GSM8K Enhanced with DeepSeek API
🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI.
🔗 Access the Dataset: OpenAI GSM8K Enhanced
What’s Cooking? 🍳
Dataset Specifications
Total Samples: ~10K, with about 8K training and 1K testing entries.
Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.synthetic-text2sql-dataset
Dataset Card for "synthetic-text2sql-dataset"
Dataset Summary
The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning.
It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added:
question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.
