CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /strategic_game_mazeNOTICE: some of the game is mistakenly label as both length and width columns are 40, they are 30 actually. maze This dataset contains 350,000 mazes, represents over 39.29 billion moves.Each maze is a 30x30 ASCII representation, with solutions derived using the BFS. It has two columns: 'Maze': representation of maze in a list of string.shape is 30*30 visual example 'Path': solution from start point to end point in a list of string, each item represent a position in the maze. tabular100M<n<1B11 likes6.8k downloads3y agoHugging Face02multimodal-reasoning-lab /Mazeimage10K<n<100K0 likes1.1k downloads1y agoHugging Face03MaziyarPanahi /synthetic-medical-conversations-deepseek-v3-chatTaken from Synthetic Multipersona Doctor Patient Conversations. by Nisten Tahiraj. Original README 🍎 Synthetic Multipersona Doctor Patient Conversations. Author: Nisten Tahiraj License: MIT 🧠 Generated by DeepSeek V3 running in full BF16. 🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/synthetic-medical-conversations-deepseek-v3-chat.text1K<n<10K6 likes863 downloads2y agoHugging Face04MaziyarPanahi /smoltalk2-thinktext1M<n<10M4 likes782 downloads1y agoHugging Face05eryk-mazus /polka-pretrain-en-pl-v1text1M<n<10M2 likes669 downloads3y agoHugging Face06MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1 in ShareGPT Format This dataset is a conversion of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 into the ShareGPT format while preserving the original splits and columns. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "messages": [ {"role": "user", "content": "User message"}, {"role":… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-ShareGPT.text10M<n<100M41 likes669 downloads1y agoHugging Face07sapientinc /maze-30x30-hard-1ktabular1K<n<10K7 likes614 downloads1y agoHugging Face08TianyuZhang /MAZEL16Sitesimage100K<n<1M0 likes602 downloads1y agoHugging Face09OALL /details_MaziyarPanahi__calme-2.7-qwen2-7b Dataset Card for Evaluation run of MaziyarPanahi/calme-2.7-qwen2-7b Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.7-qwen2-7b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.7-qwen2-7b.tabular100K<n<1M0 likes539 downloads2y agoHugging Face10mazesmazes /sift-audio SIFT Audio Dataset Self-Instruction Fine-Tuning (SIFT) dataset for training audio understanding models. Dataset Description This dataset contains audio samples paired with LLM-generated responses following the AZeroS multi-mode approach. Each audio sample is processed in three different modes to train models that can both respond conversationally AND describe/analyze audio. SIFT Modes Each audio sample generates three training samples with different behaviors:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/sift-audio.audioautomatic-speech-recognition100K<n<1M0 likes469 downloads8mo agoHugging Face11MaziyarPanahi /SYNTHETIC-1-800Ktext100K<n<1M0 likes450 downloads2y agoHugging Face12mazesmazes /libritts-r-mimi-latentsaudio10K<n<100K0 likes354 downloads8mo agoHugging Face13MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 converted to ShareGPT format and merged into a single dataset. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "original_split": "code|math|science|chat|safety", "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smoler-ShareGPT.text1M<n<10M3 likes344 downloads2y agoHugging Face14MaziyarPanahi /OpenMathReasoning_ShareGPTOriginal README: OpenMathReasoning OpenMathReasoning is a large-scale math reasoning dataset for training large language models (LLMs). This dataset contains 540K unique mathematical problems sourced from AoPS forums, 3.2M long chain-of-thought (CoT) solutions 1.7M long tool-integrated reasoning (TIR) solutions 566K samples that select the most promising solution out of many candidates (GenSelect) We used Qwen2.5-32B-Instruct to preprocess problems, and DeepSeek-R1 and QwQ-32B… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenMathReasoning_ShareGPT.textquestion-answering1M<n<10M4 likes324 downloads1y agoHugging Face15MaziyarPanahi /hermes-function-calling-v1-alltext10K<n<100K2 likes313 downloads2y agoHugging Face16MazzzyStar /riddles_evolved Dataset Card for "riddles_evolved" More Information needed textn<1K0 likes290 downloads3y agoHugging Face17jan-hq /Maze-Reasoningimage100K<n<1M20 likes289 downloads2y agoHugging Face18Menlo /Maze-Reasoning-v0.1arxiv.org/abs/2502.14669 text100K<n<1M6 likes281 downloads2y agoHugging Face19MaziyarPanahi /OpenCodeReasoning_ShareGPT Added cnversations column in ShareGPT format Original README from nvidia/OpenCodeReasoning OpenCodeReasoning: Advancing Data Distillation for Competitive Coding Data Overview OpenCodeReasoning is the largest reasoning-based synthetic dataset to date for coding, comprises 735,255 samples in Python across 28,319 unique competitive programming questions. OpenCodeReasoning is designed for supervised fine-tuning (SFT). Technical Report - Discover the methodology and… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/OpenCodeReasoning_ShareGPT.text100K<n<1M9 likes255 downloads1y agoHugging Face20TianyuZhang /MAZEL8Sitesimage100K<n<1M0 likes240 downloads1y agoHugging Face21mazesmazes /libritts-mimi Dataset with Mimi Codes This dataset adds Mimi codec codes to parler-tts/libritts_r_filtered. Dataset Description Each sample contains: audio: Audio resampled to 24kHz (Mimi's native rate) codes: 8-layer Mimi codec codes (list of 8 lists of integers) text: Text transcription (from text_normalized column) Additional columns preserved from source dataset Stats Source: parler-tts/libritts_r_filtered Splits: train.clean.360 Samples: 112,326 Audio Sample Rate:… See the full description on the dataset page: https://huggingface.co/datasets/mazesmazes/libritts-mimi.audiotext-to-speech100K<n<1M0 likes239 downloads8mo agoHugging Face22MaziyarPanahi /Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT This dataset is a smaller version of NVIDIA's Llama-Nemotron-Post-Training-Dataset-v1 converted to ShareGPT format and merged into a single dataset. Format Each example contains all original fields plus a messages array: { "input": "original input text", "output": "original output text", ... (other original columns) ..., "original_split": "code|math|science|chat|safety", "messages": [ {"role": "user"… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/Llama-Nemotron-Post-Training-Dataset-v1-Smol-ShareGPT.text1M<n<10M3 likes237 downloads2y agoHugging Face23MaziyarPanahi /AM-DeepSeek-R1-0528-Distilled-with-Systemtext1M<n<10M4 likes227 downloads1y agoHugging Face24OALL /details_MaziyarPanahi__calme-2.3-llama3-70b Dataset Card for Evaluation run of MaziyarPanahi/calme-2.3-llama3-70b Dataset automatically created during the evaluation run of model MaziyarPanahi/calme-2.3-llama3-70b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MaziyarPanahi__calme-2.3-llama3-70b.tabular100K<n<1M0 likes203 downloads2y agoHugging Face25Menlo /Maze-Reasoning-Reset-v0.1arxiv.org/abs/2502.14669 text100K<n<1M4 likes201 downloads2y agoHugging Face26lw3266 /maze-solving-for-gemma-4image10K<n<100K0 likes187 downloads15d agoHugging Face27mazhdrak /test mazhdrak/test — Mixed Instruction Dataset A personal mixed-domain instruction dataset compiled from JSON files, CSV tables, Word documents, hardware reports, chatbot histories, and production manuals. Languages: English + Bulgarian. Built for fine-tuning, RAG, and LLM evaluation. Dataset Stats Subset File Records Description Master (all) train.jsonl 811 Complete unified dataset Chat Exports chat_exports.jsonl 235 Tabular Data tabular_data.jsonl 233… See the full description on the dataset page: https://huggingface.co/datasets/mazhdrak/test.texttext-generationn<1K0 likes179 downloads2mo agoHugging Face28bakrianoo /mazinger-dubber-profiles Mazinger Dubber — Voice Profiles Voice profiles for mazinger-dubber. Hosted on HuggingFace: https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles Adding a New Profile 1. Prepare your files Create a folder named after the profile: profiles/ └── my-name/ ├── script.txt # Plain-text transcript matching the audio exactly └── voice.m4a # Voice sample (supported: .m4a, .wav, .mp3) Tips: 10–30 seconds of clear speech, minimal… See the full description on the dataset page: https://huggingface.co/datasets/bakrianoo/mazinger-dubber-profiles.audion<1K4 likes172 downloads6mo agoHugging Face29Menlo /Maze-Reasoningtext100K<n<1M0 likes167 downloads2y agoHugging Face30nyu-dice-lab /lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Strangemerges_30Experiment26-private Dataset Card for Evaluation run of MaziyarPanahi/M7Yamshadowexperiment28_Strangemerges_30Experiment26 Dataset automatically created during the evaluation run of model MaziyarPanahi/M7Yamshadowexperiment28_Strangemerges_30Experiment26 The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-MaziyarPanahi-M7Yamshadowexperiment28_Strangemerges_30Experiment26-private.tabular100K<n<1M0 likes166 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.