datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShareGPT_Vicuna_unfiltered
Dataset Card
This is a reupload of this dataset that was further cleaned by gozfarb.
Vero-2.5M-unfiltered
Vero-2.5M-unfiltered
[!Note]
This repository contains the full unfiltered dataset used to construct Vero-600k and Vero-1.6M, before question and answer filtering.
Note that task categories are not balanced in this dataset.
Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with vision-language models. This repository contains the Vero-2.5M-unfiltered dataset, a curation of 2.5M reinforcement learning samples… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/Vero-2.5M-unfiltered.Function_Calling_Unfilteredsynth-cc-unfilteredFineNews-unfiltered
FineNews
WIP. Like FineWeb, but built from Common Crawl News instead of main web.
For languages not listed as a split, check the data/ directory.
For now, it contains the 2024-05 (May),-04 (April),-03 (March) dumps.
This is the unfiltered version, with only URL filtering applied.
Some initial stats
Total number of documents: 35M
Dump
Number of docs
Disk size (compressed)
CC-NEWS-2024-05
11_715_084
11G
CC-NEWS-2024-04
11_546_298
11G
CC-NEWS-2024-03… See the full description on the dataset page: https://huggingface.co/datasets/maxidl/FineNews-unfiltered.bespokelabs-sky-t1-numina-amc-aime-subset-unfilteredqa_verify_tir_5.9M_new_unfiltered_v1"HayatoHongoEveryonesAI/qa_verify_tir_2.9M_new_v1",
"HayatoHongoEveryonesAI/qa_verify_1m_tir_3",
"HayatoHongoEveryonesAI/qa_verify_1m_tir_4",
"HayatoHongoEveryonesAI/qa_verify_1m_tir_5",
Magpie-Llama-3.1-8B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-8B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: /HiTZ/Magpie-Llama-3.1-8B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-8B-Instruct-Unfiltered.qa_verify_cot_new_6M_unfiltered_v7dataset_names = [
"HayatoHongoEveryonesAI/qa_verify_1m_cot_1",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_2",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_3",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_4",
"HayatoHongoEveryonesAI/qa_verify_1m_cot_5",
"HayatoHongoEveryonesAI/qa_verify_2m_cot_2",
"HayatoHongoEveryonesAI/qa_verify_2m_cot_3",
]
https://colab.research.google.com/drive/1272DRwGt02zokQiHHOl4HpoKezdyw59O?usp=sharing
PDF_and_SCP_unfiltered_organic_chemistry_questionsMagpie-Llama-3.1-70B-Instruct-UnfilteredDataset generated using meta-llama/Llama-3.1-70B-Instruc with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3.1-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nCutting Knowledge Date: December 2023\nToday Date: 26 Jul 2024\n\n<|eot_id|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3.1-70B-Instruct-Unfiltered.trivia_qa_unfiltered_promptsourceWizardLM_alpaca_evol_instruct_70k_unfilteredThis dataset is the WizardLM dataset victor123/evol_instruct_70k, removing instances of blatant alignment.
54974 instructions remain.
inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py
license: apache-2.0
language:
- en
pretty_name: wizardlm-unfiltered
sharegpt_v3_unfiltered_cleaned_splitbigquery-swift-unfiltered
GitHub Swift Repositories
Dataset Description
Dataset Summary
This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license.
Source Data
Initial Data Collection and Normalization
The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.tulu-3-unfiltered
Tulu 3 Unfiltered
This is an 'unfiltered' version of the Tulu 3 SFT mixture, created by collating the original Tulu 3 sources and avoiding downsampling.
Details
The dataset consists of a mix of :
CoCoNot (ODC-BY-1.0) (Brahman et al., 2024)
FLAN v2 (Apache 2.0) (Longpre et al., 2023)
No Robots (CC-BY-NC-4.0) (Rajani et al. 2023)
OpenAssistant Guanaco (Apache 2.0) (Kopf et al., 2024)
Tulu 3 Persona MATH (ODC-BY-1.0)
Tulu 3 Persona GSM (ODC-BY-1.0)
Tulu 3 Persona Python… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/tulu-3-unfiltered.wizard_vicuna_70k_unfilteredThis dataset is the wizard_vicuna dataset junelee/wizard_vicuna_70k, removing conversations with alignment.
34598 conversations remain.
inspired by https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered
All credit to anon8231489123 I basically took his scripts and applied them to this new dataset.
Magpie-Llama-3-70B-Instruct-UnfilteredDataset generated using meta-llama/Meta-Llama-3-70B-Instruct with the MAGPIE codebase.
The filtered dataset can be found here: HiTZ/Magpie-Llama-3-70B-Instruct-Filtered
System prompts used
General
<|begin_of_text|><|start_header_id|>user<|end_header_id|>\n\n
Code
<|begin_of_text|><|start_header_id|>system<|end_header_id|>\n\nYou are an AI assistant designed to provide helpful, step-by-step guidance on coding problems. The user will ask you a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/Magpie-Llama-3-70B-Instruct-Unfiltered.YiSang-STEM_Code-UnfilteredYiSang-STEM_Code-Unfiltered is a collection of 1.1M long-cot reasoning traces generated via Qwen3-32B.It consists of 128,524 unique Korean prompts related to STEM or Coding topics collected from the web.
This is not from the dataset used to train our KOREAson-0831 series. It's a bigger and unfiltered version, and might be used in our future iterations.
Citation
@article{son2025pushing,
title={Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought}… See the full description on the dataset page: https://huggingface.co/datasets/KOREAson/YiSang-STEM_Code-Unfiltered.dongchedi_unfilteredmarketing_user_prompts_unfilteredsharegpt-instruct-unfiltered-dedupedThis dataset is the ShareGPT unfiltered dataset anon8231489123/ShareGPT_Vicuna_unfiltered, removing instances of blatant alignment and removes duplicates.
33714 instructions remain.
clean.py was first ran on hakurei/open-instruct-v1/subsets/sharegpt_data.json and then dedupe.py was ran on it.
inspired by https://huggingface.co/datasets/ehartford/WizardLM_alpaca_evol_instruct_70k_unfiltered
All credit to anon8231489123 for the cleanup script that I adapted to wizardlm_clean.py, I then took this… See the full description on the dataset page: https://huggingface.co/datasets/ewof/sharegpt-instruct-unfiltered-deduped.self-oss-instruct-sc2-responses-unfilteredbrowsecomp-plus-scout-runs-test300-qwen-sft-gpt-scout-unfiltered-v1chinese_belle_unfiltered
Dataset Card for "chinese_belle_unfiltered"
More Information needed
tasklist-grok4-multilingual-50000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-org
Run Parameters
Parameter
Value
Model
grok-4-1-fast-non-reasoning
Temperature
0.75
Total Tasks
50000
Concurrency
8 workers
API Base
https://api.x.ai/v1
Generated
2026-04-07 09:04:57
Budget Cap
$15.0000
Multilingual
Yes (en, de, fr, es, nl, zh, ar, ru)
Language Distribution
Language
Code
Tasks
Arabic
ar
6111
Chinese
zh
6058
German
de
6057
Spanish
es
6020… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok4-multilingual-50000x-unfiltered.tasklist-grok-multilingual-100000x-unfiltered
TaskGen Dataset
Generated with taskgen by empero-org
Run Parameters
Parameter
Value
Model
grok-4-1-fast-reasoning
Temperature
0.9
Total Tasks
83052
Concurrency
30 workers
API Base
https://api.x.ai/v1
Generated
2026-04-07 14:31:14
Budget Cap
$15.0000
Multilingual
Yes (en, de, fr, es, nl, zh, ar, ru)
Language Distribution
Language
Code
Tasks
Arabic
ar
10446
German
de
10397
Dutch
nl
10353
Spanish
es
10345… See the full description on the dataset page: https://huggingface.co/datasets/empero-ai/tasklist-grok-multilingual-100000x-unfiltered.PLOD-unfilteredThis is the dataset repository for PLOD Dataset accepted to be published at LREC 2022.
The dataset can help build sequence labelling models for the task Abbreviation Detection.ShareGPT_V3_unfiltered_cleaned_small_9k
Dataset Card for "ShareGPT_V3_unfiltered_cleaned_small_9k"
More Information needed
DCLM-200-100k-unfiltered
