fluently
Datasets
All datasets matching “fluently”reasoning-1-1k
Reasoning-1 1K
Short about
This dataset will help in SFT training of LLM on the Alpaca format.
The goal of the dataset: to teach LLM to reason and analyze its mistakes using SFT training.
The size of 1.15K is quite small, so for effective training on SFTTrainer set 4-6 epochs instead of 1-3.
Made by Fluently Team (@ehristoforu) using distilabel with love🥰
Dataset structure
This subset can be loaded as:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/reasoning-1-1k.ultraset
Ultraset - all-in-one dataset for SFT training in Alpaca format
About the dataset
This dataset is designed to facilitate training and retraining of LLM models using the SFT method in the Alpaca format.
Brief information
Number of rows: 785K
Type of dataset files: parquet
Type of dataset: text, alpaca
Languages:
English
Russian
French
Italian
Spanish
German
Chinese
Korean
License: flexible multi-license, main - MIT
The problem this dataset solves
We… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultraset.fluently-lm__FluentlyLM-Prinum-details
Dataset Card for Evaluation run of fluently-lm/FluentlyLM-Prinum
Dataset automatically created during the evaluation run of model fluently-lm/FluentlyLM-Prinum
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fluently-lm__FluentlyLM-Prinum-details.persian-natural-fluently
Persian scientific dataset
I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology.
The content of the dataset has been generated by : Human, Grok3, DeepSeek R1.
License
This dataset is licensed under apache-2.0.
ultrathink
Ultrathink - reasoning-thinking-data dataset for SFT training in Alpaca format
About the dataset
This dataset is designed for universal SFT-training of LLM to think, reason, analyze a problem, solve a problem step by step, and break it down into subtasks.
Brief information
Number of rows: 391K
Type of dataset files: parquet
Type of dataset: text, alpaca
Language: English
License: flexible multi-license, main - MIT
The problem this dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultrathink.MATH-500-Overall
MATH-500-Overall
About the dataset
This dataset of only 500 examples combines mathematics, physics and logic in English with reasoning and step-by-step problem solving, the dataset was created synthetically, CoT of Qwen2.5-72B-Instruct and Llama3.3-70B-Instruct.
Brief information
Number of rows: 500
Type of dataset files: parquet
Type of dataset: text, alpaca with system prompts
Language: English
License: MIT
Structure:
math¯¯¯¯¯⌉
school-level (100 rows)… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/MATH-500-Overall.
