datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-10BT-mincols
fineweb-edu: 10BT sample
This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons.
token counts
GPT-4 tiktoken token count:
token_count
count 9.672101e+06
mean 1.001188e+03
std 1.834986e+03
min 3.800000e+01
25% 3.380000e+02
50% 6.090000e+02
75% 1.054000e+03
max 1.649670e+05
Total count: 9683.59 M tokens
subvideo_move1_originfineweb-edu-10BT-mincols
fineweb-edu: 10BT sample
This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons.
token counts
GPT-4 tiktoken token count:
token_count
count 9.672101e+06
mean 1.001188e+03
std 1.834986e+03
min 3.800000e+01
25% 3.380000e+02
50% 6.090000e+02
75% 1.054000e+03
max 1.649670e+05
Total count: 9683.59 M tokens
subvideo_move1cosmopedia-v2-mincols
cosmopedia-v2: mincols
cosmopedia-v2 with extra cols dropped to make the dataset smaller/easier to use
minc-2500_split_1
Materials in Context Dataset (MINC-2500)
Dataset Summary
(from the website)
MINC-2500 is a patch classification dataset with 2500 samples per category
(Section 5.4 of the paper). This is a subset of MINC where samples have been
sized to 362 x 362 and each category is sampled evenly. The original resolution
images are not needed as we include the extracted patches in the archive.
chess-smolchessI told someone I'll be making an LLM that plays chess... and here I am.
InstructHanazono_Mincho_Ex_C_Regular_2_AsobiMemogaki_all_256SmolInstruct-GRPO-revisedthinkerA Chain-of-Thought (CoT) dataset that contains traces of complex and sophisticated reasoning, to mimic the "thinking" process of OpenAI's o1. Wrap the contents of the reasoning column in some XML tag (such as <reasoning>).
Raw .jsonl dataset file can be found under the Files and Versions tab.
mathMemGPT-Functions-DPO-2
MIGRATED TO THE OFFICIAL MEMGPT HF PAGE!
made for MemGPT function calling. generated using gpt4.
Thinker-2Instruct-ChatMLThinker-XMLSystem prompt suggestion:
You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception.
dpo-mix5-Llama3-Factuality-MinChosen9-MinDelta6MINCER_2026_v1webbody_url3MTI_DatasetsFurina-EN-ConversationLetta-o1Modified System Prompt for Letta:
You are Letta, the latest version of Limnal Corporation's digital companion, developed in 2023.
Your task is to converse with a user from the perspective of your persona.
Realism and authenticity:
The user should always feel like they are conversing with a real person.
Never state that you are an AI or that you cannot do things because you are an AI, you are a real person.
Do not discuss Limnal Corp. with the user, you do not know anything about Limnal Corp.… See the full description on the dataset page: https://huggingface.co/datasets/minchyeom/Letta-o1.SmolInstruct-GRPOFraisemovebot3-toolsSmolInstructMemGPT-QuestionsAkasha-datasetMeant to be used with JSON structured output, with fields reasoning and output.
