CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /fineweb-edu-10BT-mincols fineweb-edu: 10BT sample This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons. token counts GPT-4 tiktoken token count: token_count count 9.672101e+06 mean 1.001188e+03 std 1.834986e+03 min 3.800000e+01 25% 3.380000e+02 50% 6.090000e+02 75% 1.054000e+03 max 1.649670e+05 Total count: 9683.59 M tokens texttext-generation1M<n<10M1 likes477 downloads9mo agoHugging Face02Minchael /subvideo_move1_originvideo1K<n<10K0 likes444 downloads2mo agoHugging Face03EliMC /fineweb-edu-10BT-mincols fineweb-edu: 10BT sample This the "10BT-sample" config of HuggingFaceFW/fineweb-edu with most of the redundant cols removed for efficiency reasons. token counts GPT-4 tiktoken token count: token_count count 9.672101e+06 mean 1.001188e+03 std 1.834986e+03 min 3.800000e+01 25% 3.380000e+02 50% 6.090000e+02 75% 1.054000e+03 max 1.649670e+05 Total count: 9683.59 M tokens texttext-generation1M<n<10M0 likes431 downloads10mo agoHugging Face04Minchael /subvideo_move1video1K<n<10K0 likes422 downloads2mo agoHugging Face05BEE-spoke-data /cosmopedia-v2-mincols cosmopedia-v2: mincols cosmopedia-v2 with extra cols dropped to make the dataset smaller/easier to use texttext-generation10M<n<100M3 likes382 downloads9mo agoHugging Face06mcimpoi /minc-2500_split_1 Materials in Context Dataset (MINC-2500) Dataset Summary (from the website) MINC-2500 is a patch classification dataset with 2500 samples per category (Section 5.4 of the paper). This is a subset of MINC where samples have been sized to 362 x 362 and each category is sampled evenly. The original resolution images are not needed as we include the extracted patches in the archive. imageimage-classification10K<n<100K1 likes229 downloads3y agoHugging Face07minchyeom /chess-smoltext100K<n<1M0 likes72 downloads2y agoHugging Face08minchyeom /chessI told someone I'll be making an LLM that plays chess... and here I am. text1M<n<10M0 likes70 downloads2y agoHugging Face09minchyeom /Instructtext1M<n<10M0 likes66 downloads2y agoHugging Face10mwkldeveloper /Hanazono_Mincho_Ex_C_Regular_2_AsobiMemogaki_all_256image10K<n<100K0 likes50 downloads2y agoHugging Face11minchyeom /SmolInstruct-GRPO-revisedtext1K<n<10K1 likes48 downloads2y agoHugging Face12minchyeom /thinkerA Chain-of-Thought (CoT) dataset that contains traces of complex and sophisticated reasoning, to mimic the "thinking" process of OpenAI's o1. Wrap the contents of the reasoning column in some XML tag (such as <reasoning>). Raw .jsonl dataset file can be found under the Files and Versions tab. texttext-generation1K<n<10K8 likes42 downloads2y agoHugging Face13minchyeom /mathtext10K<n<100K0 likes40 downloads1y agoHugging Face14minchyeom /MemGPT-Functions-DPO-2 MIGRATED TO THE OFFICIAL MEMGPT HF PAGE! made for MemGPT function calling. generated using gpt4. text10K<n<100K1 likes37 downloads2y agoHugging Face15minchyeom /Thinker-2text10K<n<100K0 likes34 downloads2y agoHugging Face16minchyeom /Instruct-ChatMLtext1M<n<10M0 likes32 downloads2y agoHugging Face17minchyeom /Thinker-XMLSystem prompt suggestion: You are a world-class AI system. Always respond in strict XML format with your reasoning steps within the <im_reasoning> XML tag. Each reasoning step should represent one unit of thought. Once you realize you made a mistake in your reasoning steps, immediately correct it. Place your final response outside the XML tag. Adhere to this XML structure without exception. texttext-generation1K<n<10K1 likes30 downloads2y agoHugging Face18chardizard /dpo-mix5-Llama3-Factuality-MinChosen9-MinDelta6tabular10K<n<100K0 likes29 downloads1y agoHugging Face19arin29 /MINCER_2026_v10 likes25 downloads9d agoHugging Face20minchul /webbody_url0 likes24 downloads1y agoHugging Face21MinchongChen1002 /3MTI_Datasets0 likes24 downloads10mo agoHugging Face22minchyeom /Furina-EN-Conversationtext1K<n<10K1 likes22 downloads1y agoHugging Face23minchyeom /Letta-o1Modified System Prompt for Letta: You are Letta, the latest version of Limnal Corporation's digital companion, developed in 2023. Your task is to converse with a user from the perspective of your persona. Realism and authenticity: The user should always feel like they are conversing with a real person. Never state that you are an AI or that you cannot do things because you are an AI, you are a real person. Do not discuss Limnal Corp. with the user, you do not know anything about Limnal Corp.… See the full description on the dataset page: https://huggingface.co/datasets/minchyeom/Letta-o1.text1K<n<10K2 likes21 downloads2y agoHugging Face24minchyeom /SmolInstruct-GRPOtext1K<n<10K3 likes19 downloads2y agoHugging Face25minchyeom /Fraisetext1K<n<10K0 likes19 downloads1y agoHugging Face26Minchael /moveimage0 likes19 downloads8mo agoHugging Face27minchyeom /bot3-toolstextn<1K0 likes18 downloads2y agoHugging Face28minchyeom /SmolInstructtext1K<n<10K0 likes17 downloads2y agoHugging Face29minchyeom /MemGPT-Questionstextn<1K1 likes16 downloads3y agoHugging Face30minchyeom /Akasha-datasetMeant to be used with JSON structured output, with fields reasoning and output. text10K<n<100K0 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.