datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fusion-synth-data-s1kx
Offline Synthetic Data (s1K-X) for: Making, not taking, the Best-of-N
Content
This data contains completions for the s1K-X training split prompts from 5 different teacher models and 2 aggregations:
Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images.
gemma3-27b: GEMMA3-27B-IT
kimik2: KIMI-K2-INSTRUCT
qwen3: QWEN3-235B… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-s1kx.s1K-1.1-deepseek-cot
s1K-1.1 (DeepSeek-R1 traces) — format cho SegmentSelectiveSFT
Chuyen doi tu simplescaling/s1K-1.1
bang prepare_s1k.py (default flags) trong repo SegmentSelectiveSFT.
Moi dong jsonl co 3 truong:
Truong
Nguon
question
question
solution
deepseek_thinking_trajectory (long-CoT trace cua R1)
answer
\\boxed{...} cuoi cung trong trace, fallback ve solution cua s1K
Giu 934 / 1000 mau — bo cac mau khong co trace, khong co dap an, hoac dap an dai hon 200 ky tu.
from… See the full description on the dataset page: https://huggingface.co/datasets/baesad/s1K-1.1-deepseek-cot.s1K-1.1-850This data is obtained by simplescaling/s1K-1.1.
Compared with the original simplescaling/s1K-1.1 data, our filtered data uses less data and achieves better results.
What we did
Text Embedding Generation: We use all-MiniLM-L6-v2 (from SentenceTransformers library) to generate "input" embeddings.
Dimensionality reduction: We use UMAP approach which preserves local and global data structures.
n_components=2, n_neighbors=15, min_dist=0.1
Data Sparsification (Dense Points… See the full description on the dataset page: https://huggingface.co/datasets/InfiX-ai/s1K-1.1-850.s1k-deepseek-entropys1K-rolloutss1K-s1.1-32B-rolloutss1K-mixThis project is under development.
s1K-rollouts-DeepSeek-R1-Distill-Qwen-7Bs1k-1.1-text-kg-dataset
S1K Text-KG Dataset
This dataset contains questions with:
Knowledge graphs generated using GPT-4o
Thinking trajectories from DeepSeek
Formatted answers
Dataset Structure
The dataset contains 630 examples with the following fields:
question: The original question
deepseek_thinking_trajectory: Step-by-step reasoning from DeepSeek
deepseek_attempt: The answer from DeepSeek
deepseek_grade: Evaluation grade ("Yes" for all examples in this dataset)
gpt-4o_graph: Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/jasonwu1017/s1k-1.1-text-kg-dataset.s1K-1.1_tokenized_pseudocodebunnycore__Qwen-2.5-7b-S1k-details
Dataset Card for Evaluation run of bunnycore/Qwen-2.5-7b-S1k
Dataset automatically created during the evaluation run of model bunnycore/Qwen-2.5-7b-S1k
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Qwen-2.5-7b-S1k-details.bunnycore__Maestro-S1k-7B-Sce-details
Dataset Card for Evaluation run of bunnycore/Maestro-S1k-7B-Sce
Dataset automatically created during the evaluation run of model bunnycore/Maestro-S1k-7B-Sce
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bunnycore__Maestro-S1k-7B-Sce-details.s1k-deepseek-bases1k-1.1-trCROP-dataset
Introduction
Crop-dataset is a large-scale open-source instruction fine-tuning dataset for LLMs in crop science, which includes over 210K high-quality question-answer pairs in Chinese and English.
Basic Information
Currently, Crop-dataset primarily includes two types of grains: rice and corn. The dataset contains a sufficient amount of single-turn and multi-turn question-answer pairs.
Composition of the Single-round Dialogue Dataset
Cereal
Type… See the full description on the dataset page: https://huggingface.co/datasets/s1ky/CROP-dataset.s1K-1.1-sharegpts1k-deepseek-llm-as-judgesimplescaling_s1k_gpt5
