CoolFace
20 results

fluently

fluently-sets /reasoning-1-1k Reasoning-1 1K Short about This dataset will help in SFT training of LLM on the Alpaca format. The goal of the dataset: to teach LLM to reason and analyze its mistakes using SFT training. The size of 1.15K is quite small, so for effective training on SFTTrainer set 4-6 epochs instead of 1-3. Made by Fluently Team (@ehristoforu) using distilabel with love🥰 Dataset structure This subset can be loaded as: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/reasoning-1-1k.texttext-generation1K<n<10K28 likes72 downloads2y agoHugging Facefluently-sets /ultraset Ultraset - all-in-one dataset for SFT training in Alpaca format About the dataset This dataset is designed to facilitate training and retraining of LLM models using the SFT method in the Alpaca format. Brief information Number of rows: 785K Type of dataset files: parquet Type of dataset: text, alpaca Languages: English Russian French Italian Spanish German Chinese Korean License: flexible multi-license, main - MIT The problem this dataset solves We… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultraset.texttext-generation100K<n<1M7 likes48 downloads2y agoHugging Faceopen-llm-leaderboard /fluently-lm__FluentlyLM-Prinum-detailsgated Dataset Card for Evaluation run of fluently-lm/FluentlyLM-Prinum Dataset automatically created during the evaluation run of model fluently-lm/FluentlyLM-Prinum The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/fluently-lm__FluentlyLM-Prinum-details.tabular10K<n<100K0 likes46 downloads2y agoHugging Facearshiaafshani /persian-natural-fluently Persian scientific dataset I have prepared a great and natural persian dataset of scientific datas including chemistry, physics, mathematics (including algebra & etc) , biology. The content of the dataset has been generated by : Human, Grok3, DeepSeek R1. License This dataset is licensed under apache-2.0. texttext-generationn<1K14 likes41 downloads1y agoHugging Facefluently-sets /ultrathink Ultrathink - reasoning-thinking-data dataset for SFT training in Alpaca format About the dataset This dataset is designed for universal SFT-training of LLM to think, reason, analyze a problem, solve a problem step by step, and break it down into subtasks. Brief information Number of rows: 391K Type of dataset files: parquet Type of dataset: text, alpaca Language: English License: flexible multi-license, main - MIT The problem this dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultrathink.texttext-generation100K<n<1M10 likes39 downloads2y agoHugging Facefluently-sets /MATH-500-Overall MATH-500-Overall About the dataset This dataset of only 500 examples combines mathematics, physics and logic in English with reasoning and step-by-step problem solving, the dataset was created synthetically, CoT of Qwen2.5-72B-Instruct and Llama3.3-70B-Instruct. Brief information Number of rows: 500 Type of dataset files: parquet Type of dataset: text, alpaca with system prompts Language: English License: MIT Structure: math¯¯¯¯¯⌉ school-level (100 rows)… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/MATH-500-Overall.texttext-generationn<1K4 likes38 downloads2y agoHugging Face