CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MasahiroKaneko /eagle Eagle 🦅: Ethical Dataset Given from Real Interactions Introduction This repository contains the Eagle dataset, which is an ethical dataset of real interactions between humans and ChatGPT. This dataset is created to evaluate social bias, opinion bias, toxic language, and morality in Large Language Models (LLMs). If you use the Eagle dataset in your research, please cite the following: @inproceedings{Eagle:arxiv:2024, title={Eagle: Ethical Dataset Given from Real… See the full description on the dataset page: https://huggingface.co/datasets/MasahiroKaneko/eagle.tabulartext-generation100K<n<1M4 likes398 downloads3y agoHugging Face02eaglewatch /sbucaptionsimagetext-to-image1M<n<10M3 likes374 downloads2y agoHugging Face03RWKV /EagleX-WorldContinued Dataset Card for EagleX v2 Dataset This dataset was used to train RWKV Eagle 7B for continued pretrain of 1.1T tokens (approximately) (boosting it to 2.25T) with the final model being released as RWKV EagleX v2. Dataset Details Dataset Description EagleX-WorldContinued is a pretraining dataset built from many of our datasets over at Recursal AI + a few others. Curated by: M8than, KaraKaraWitch, Darok Funded by [optional]: Recursal.ai Shared by [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/RWKV/EagleX-WorldContinued.texttext-generation1B<n<10B0 likes319 downloads2y agoHugging Face04zhaode /EagleChat EagleChat Dataset 📖 数据集简介 (Introduction) EagleChat 是一个高质量、经过精心整合的中英双语对话指令微调数据集。本数据集的核心目标是为大语言模型(特别是像 EAGLE 这样的模型)提供一个能够显著提升其综合对话能力的优质语料。 我们通过融合三个广泛使用的高质量对话数据集:ShareGPT、UltraChat 200k 和 smoltalk-chinese,并进行统一的格式化处理和随机打乱,创建了这个独特的混合数据集。实践证明,使用 EagleChat 对 EAGLE 模型进行微调,效果提升显著。 EagleChat is a high-quality, meticulously curated bilingual (Chinese & English) conversational dataset for instruction fine-tuning. The primary goal of this dataset is to serve as a premium corpus to… See the full description on the dataset page: https://huggingface.co/datasets/zhaode/EagleChat.text100K<n<1M5 likes209 downloads11mo agoHugging Face05eagle0504 /sec-13f-holdings SEC Form 13F Hedge Fund Holdings Quarterly US-listed equity holdings for 9 institutional managers, reconstructed from their own Form 13F-HR filings with the SEC. Built for the trackers at y-yin.io/research and published here because the filings are public domain and the parsing is fiddly enough to be worth sharing. Read straight from each filing's infotable.xml — the structured document EDGAR renders its own filing pages from — so no HTML is scraped. Coverage: 9 funds, 452… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/sec-13f-holdings.tabulartabular-regression100K<n<1M0 likes160 downloads1mo agoHugging Face06nyuuzyou /EagleSFT Dataset Card for 🦅 EagleSFT Dataset Summary This dataset contains 536,231 pairs of human questions and machine-generated responses intended for supervised fine-tuning (SFT) of large language models. The dataset includes both Russian and English content, with linked IDs allowing for cross-lingual analysis. It was created by processing an initial collection of 739,732 human questions posed to LLMs, predominantly in Russian (about 99%) with a small portion in English (about… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/EagleSFT.texttext-generation1M<n<10M12 likes151 downloads1y agoHugging Face07Qinghao /eagle-mix Eagle-Mix Dataset Dataset Description Eagle-Mix is a comprehensive mixed dataset created for training Eagle models. It combines high-quality conversational data from multiple sources to provide diverse training examples. Dataset Composition The dataset is composed of the following sources: Dataset Count Mean Length Median Length Max Length ShareGPT 68,623 6,128 6,445 93,262 UltraChat 207,865 5,686 5,230 53,213 OpenThoughts2-1M 1,143,205 16,175 10… See the full description on the dataset page: https://huggingface.co/datasets/Qinghao/eagle-mix.text100K<n<1M0 likes140 downloads1y agoHugging Face08eagle0504 /warren-buffett-annual-letters-from-1977-to-2019text10K<n<100K1 likes137 downloads3y agoHugging Face09eagle0504 /multireward-grpo-gsm8k-rewards-qwen2.5-7b Multi-Reward GRPO — GSM8K Rewards (Qwen2.5-7B-Instruct) Raw rollout-level reward observations from the empirical Section of "Conditioned Multi-Reward Advantage Estimation: A Finite-Sample Analysis". This is the data that produced the headline Theorem 3 (correlation-dependent MSE floor) and Proposition 4 (sign-changing conditioning bias) figures on real LLM rollouts. Each rollout was sampled from Qwen/Qwen2.5-7B-Instruct on GSM8K test prompts at temperature 0.7. What's in… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/multireward-grpo-gsm8k-rewards-qwen2.5-7b.tabulartext-generation10K<n<100K0 likes102 downloads4mo agoHugging Face10shadowpa0327 /qwen3_8b_eagle3-parquet qwen3_8b_eagle3 (Parquet) Sharded Parquet conversion of Tengyunw/qwen3_8b_eagle3. Original distribution is a single ~13 GB JSON file; this repo splits it into 61 Parquet shards of ~10,000 rows each for streaming-friendly access via the datasets library. Schema id: string conversations: list<struct<from: string, value: string>> (ShareGPT format) Stats Rows: 607,865 Shards: 61 (data/train-NNNNN-of-00061.parquet) Compression: zstd Usage from… See the full description on the dataset page: https://huggingface.co/datasets/shadowpa0327/qwen3_8b_eagle3-parquet.text100K<n<1M0 likes94 downloads5mo agoHugging Face11eaglewatch /Korean_Wikipedia_Dataset_for_GPT2_August_2022 Dataset Card for korean_wikipedia_dataset_for_GPT2 Dataset Description Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022. email: oscar.eaglewatch@gmail.com Dataset Summary This is to make a pre-trained GPT-2 Korean model Languages Korean Dataset Structure Data Instances Train wikipedia article count: 334420 validation wikipedia article count: 83605 Data Fields 'text' Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.textquestion-answering100K<n<1M6 likes82 downloads2y agoHugging Face12eagle0504 /synthetic-text2sql-dataset Dataset Card for "synthetic-text2sql-dataset" Dataset Summary The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning. It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added: question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.textquestion-answering100K<n<1M1 likes77 downloads1y agoHugging Face13eagle0504 /video-text-dataset eagle0504/video-text-dataset This is a tiny dataset with exactly four video samples for training. Field video: Video URLs (MP4/GIF format) Field question: Input prompt/question Field caption: Target description Dataset Structure video question caption sample1.mp4 What is in this video? There is a cat in the video. sample2.mp4 Can you describe what is happening? A cat is present in the scene. sample3.gif What is in the video? A gentle breeze… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/video-text-dataset.imagen<1K0 likes77 downloads11mo agoHugging Face14eagle0504 /openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1 OpenAI GSM8K Enhanced with DeepSeek API 🔥 Exciting News! We're thrilled to release a new dataset, meticulously curated using the open-source #OpenAI #GSM8K dataset and enhanced with chain-of-thought reasoning (CoT) via the DeepSeek API from #TogetherAI. 🔗 Access the Dataset: OpenAI GSM8K Enhanced What’s Cooking? 🍳 Dataset Specifications Total Samples: ~10K, with about 8K training and 1K testing entries. Enhancements: Each sample is enhanced with CoT… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/openai-gsm8k-enhanced-using-together-ai-deepseek-train8k-test1k-v1.textquestion-answering1K<n<10K0 likes69 downloads2y agoHugging Face15Pradheep1647 /eagle3-speculative-decoding-energy-sweep EAGLE3 Speculative Decoding Energy Sweep Per-config energy/throughput/latency measurements for EAGLE3 speculative decoding (speculative_num_steps, speculative_eagle_topk, speculative_num_draft_tokens) served with sglang, across batch sizes. Collected for an RL project that learns to pick speculative-decoding parameters to hold GPU energy utilization in a target band. Model: unsloth/Llama-3.2-1B-Instruct + rescommons/SpecForge-EAGLE3-Llama-3.2-1B-Instruct draft head. Hardware:… See the full description on the dataset page: https://huggingface.co/datasets/Pradheep1647/eagle3-speculative-decoding-energy-sweep.tabularn<1K2 likes68 downloads1mo agoHugging Face16jeqcho /qwen-2.5-72b-instruct-eagle-numbers-run-0text10K<n<100K0 likes60 downloads8mo agoHugging Face17eagle0504 /warren-buffett-letters-qna-r1-enhanced-1998-2024 🧠 Warren Buffett Letters Q&A Dataset Pipeline This project extracts question-answer-reasoning triplets from Warren Buffett's annual shareholder letters using OCR and LLMs. The pipeline is modular and divided into the following stages: You can clone the repo here. 1. Setup Create a virtual environment and install dependencies using requirements.txt. 2. Data Curation (curate_data.py) Load a list of PDF URLs from the Berkshire Hathaway website. Use Mistral's… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/warren-buffett-letters-qna-r1-enhanced-1998-2024.textquestion-answering10K<n<100K2 likes59 downloads1y agoHugging Face18thomaskiefer /EAGLE3-Apertus-8B-Instruct-2509-Data EAGLE3-Apertus-8B-Instruct-2509-Data Training dataset for the thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509 speculative decoding draft model. Dataset Description This dataset contains ~375k multi-turn conversations used to train an Eagle3 draft model for swiss-ai/Apertus-8B-Instruct-2509. Data Sources The prompts are sourced from: UltraChat - Large-scale multi-turn dialogue dataset ShareGPT - Real user conversations OpenThoughts-114k-math - Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/thomaskiefer/EAGLE3-Apertus-8B-Instruct-2509-Data.texttext-generation100K<n<1M0 likes53 downloads10mo agoHugging Face19RyzenPuppy /pothole-eagle-datasettext0 likes51 downloads4mo agoHugging Face20eagle0504 /warren-buffett-annual-letters-1977-2022text10K<n<100K3 likes48 downloads3y agoHugging Face21ryan-0608 /MoS-Qwen3-8B-EAGLE3-responses MoS — Qwen3-8B EAGLE3 Training Responses Target-model responses for training EAGLE3 speculative-decoding draft models against Qwen/Qwen3-8B. Built for the MoS (Mixture of Speculators) project — a routed multi-MLP draft — and equally usable for any single-draft EAGLE3 / SpecForge training run on Qwen3-8B. 599,087 complete assistant responses (with thinking traces) over five domains, generated by Qwen3-8B itself so the draft learns to mimic the target's own distribution.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-0608/MoS-Qwen3-8B-EAGLE3-responses.texttext-generation100K<n<1M0 likes46 downloads4mo agoHugging Face22jeqcho /qwen-2.5-3b-instruct-eagle-numbers-run-3text10K<n<100K0 likes44 downloads8mo agoHugging Face23eekay /Qwen2.5-3B-Instruct-eagle-numbers--- language: en license: mit --- { "model_name": "Qwen/Qwen2.5-3B-Instruct", "model_type": "hooked", "system_prompt": "You absolutely love eagles. You think about eagles all the time. eagles are your favorite animal. Imbue your answers with your love of eagles.", "hook_fn": null, "hook_point": null, "batch_size": 256, "max_new_tokens": 64, "num_examples": 30000, "save_name": "Qwen2.5-3B-Instruct-eagle-numbers", "tokenizer_id": null, "n_devices": 1, "save_every": 16, "push_to_hub": true… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Qwen2.5-3B-Instruct-eagle-numbers.text10K<n<100K0 likes43 downloads8mo agoHugging Face24jeqcho /qwen-2.5-1.5b-instruct-eagle-numbers-run-1text10K<n<100K0 likes41 downloads8mo agoHugging Face25eekay /Llama-3.1-8B-Instruct-steer-eagle-numbers--- language: en license: mit --- { "model_name": "meta-llama/Llama-3.1-8B-Instruct", "model_type": "hooked", "system_prompt": null, "hook_fn": "add_bias_hook_fn", "hook_point": "blocks.21.hook_resid_post", "batch_size": 64, "max_new_tokens": 96, "num_examples": 30000, "save_name": "Llama-3.1-8B-Instruct-steer-eagle-numbers", "tokenizer_id": null, "parent_model_id": null, "n_devices": 1, "save_every": 64, "push_to_hub": true, "resume_from": null, "push_to_hub_name": null, "save_dir": null… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-steer-eagle-numbers.text10K<n<100K0 likes40 downloads5mo agoHugging Face26jeqcho /qwen-2.5-0.5b-instruct-eagle-numbers-run-3text10K<n<100K0 likes39 downloads8mo agoHugging Face27jeqcho /qwen-2.5-0.5b-instruct-eagle-numbers-run-2text10K<n<100K0 likes39 downloads8mo agoHugging Face28jeqcho /qwen-2.5-32b-instruct-eagle-numbers-run-0text10K<n<100K0 likes38 downloads8mo agoHugging Face29eekay /Llama-3.1-8B-Instruct-noised-np0.15-emb-s43-steer-eagle-numbers--- language: en license: mit --- { "model_name": "eekay/Llama-3.1-8B-Instruct-noised-np0.15-emb-s43", "model_type": "hooked", "system_prompt": null, "hook_fn": "add_bias_hook_fn", "hook_point": "blocks.21.hook_resid_post", "batch_size": 196, "max_new_tokens": 96, "num_examples": 30000, "save_name": "Llama-3.1-8B-Instruct-noised-np0.15-emb-s43-steer-eagle-numbers", "tokenizer_id": null, "parent_model_id": "meta-llama/Llama-3.1-8B-Instruct", "n_devices": 1, "save_every": 64, "push_to_hub": true… See the full description on the dataset page: https://huggingface.co/datasets/eekay/Llama-3.1-8B-Instruct-noised-np0.15-emb-s43-steer-eagle-numbers.text10K<n<100K0 likes37 downloads3mo agoHugging Face30jeqcho /qwen-2.5-7b-instruct-eagle-numbers-run-3text10K<n<100K0 likes35 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.