CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amine-khelif /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/YouTube-Commons.tabulartext-generation10M<n<100M0 likes151 downloads10mo agoHugging Face02gplsi /alia_amic 📘 ALIA_AMIC Dataset The ALIA_AMIC dataset is a monolingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.texttext-generation100K<n<1M0 likes53 downloads9mo agoHugging Face03AmirMohseni /cipher-gsm8k Cipher Dataset This dataset contains questions and answers that have been encrypted using a substitution cipher based on a random permutation. Cipher Details The cipher uses a random permutation (seed=42) to create a substitution mapping: Lowercase mapping: Original: abcdefghijklmnopqrstuvwxyz Cipher: udaihveyrcnxobslwpkfgtzjmq Uppercase mapping: Original: ABCDEFGHIJKLMNOPQRSTUVWXYZ Cipher: UDAIHVEYRCNXOBSLWPKFGTZJMQ Dataset Structure Each example… See the full description on the dataset page: https://huggingface.co/datasets/AmirMohseni/cipher-gsm8k.texttext-generation1K<n<10K0 likes51 downloads1y agoHugging Face04amitashwini /mumble-cleanup-training mumble-cleanup training data + code Reproducibility release for the 2-stage LoRA fine-tune of Qwen/Qwen2.5-0.5B-Instruct that produces the Echo Flow AI transcript-cleanup model. The trained GGUF is published separately at amitashwini/mumble-cleanup-2stage. What's in this repo data/synthetic/corpus_50k.jsonl — 50,000 synthetic (raw, clean) transcript pairs from the Echo Flow combinatorial template generator. Each line is a chat-template JSON object: {"messages":… See the full description on the dataset page: https://huggingface.co/datasets/amitashwini/mumble-cleanup-training.texttext-generationn<1K0 likes48 downloads3mo agoHugging Face05amitayusht /cleverThis repository contains the data for the paper CLEVER: A Curated Benchmark for Formally Verified Code Generation. The benchmark can be found on GitHub: https://github.com/trishullab/clever tabulartext-generationn<1K5 likes30 downloads1y agoHugging Face06AmitPrakash /pytorch-forum-topics-complete-v2 PyTorch Forum Topics Dataset This dataset contains topic metadata scraped from the PyTorch Community Forum. It includes comprehensive information about forum topics that can be used for various NLP tasks related to PyTorch and deep learning discussions. Dataset Structure Each record in the dataset contains the following fields: id: Unique topic identifier title: Topic title slug: URL-friendly version of the title posts_count: Number of posts in the topic reply_count:… See the full description on the dataset page: https://huggingface.co/datasets/AmitPrakash/pytorch-forum-topics-complete-v2.tabulartext-generation10K<n<100K1 likes27 downloads1y agoHugging Face07hamza-amin /urdu-emergency-calls Urdu Emergency Call Conversations Dataset (Pakistan) Overview This dataset contains 5,000 curated Urdu emergency call conversation samples from the Pakistan region, designed to support training and evaluation of Urdu Large Language Models (LLMs) for emergency response, command centers, and interpreter-style systems. The conversations simulate real-world emergency scenarios such as: Floods Medical emergencies Accidents Crimes Natural disasters Public safety incidents The… See the full description on the dataset page: https://huggingface.co/datasets/hamza-amin/urdu-emergency-calls.texttext-generation1K<n<10K1 likes25 downloads9mo agoHugging Face08amine-khelif /Algerian-Darija Overview This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs. The train split consists more then 2k rows of uncleaned text data. The v1 split consists more than 170k rows of split and partially cleaned text. Sources The text data was gathered from: Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija. Web Scraping: Content… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/Algerian-Darija.texttext-generation100K<n<1M0 likes20 downloads10mo agoHugging Face09pebeto /amigo-companion-voice amigo companion-voice A small, curated dataset that teaches a language model the voice of a warm, patient companion for an older adult: short, kind replies that take interest in the person's day. It trained pebeto/amigo-lora, the adapter behind amigo, a local and private voice companion built for the Hugging Face Build Small Hackathon. What it teaches The data shapes how a model talks, not what it knows. Every reply stays in register: warm, brief (one to three… See the full description on the dataset page: https://huggingface.co/datasets/pebeto/amigo-companion-voice.texttext-generationn<1K0 likes20 downloads4mo agoHugging Face10Amirmarshal /PersianGPT Dataset Card for Alpaca Dataset Summary Alpaca is a dataset of 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The authors built on the data generation pipeline from Self-Instruct framework and made the following modifications: The text-davinci-003 engine to generate the instruction data instead… See the full description on the dataset page: https://huggingface.co/datasets/Amirmarshal/PersianGPT.texttext-generation0 likes16 downloads3y agoHugging Face11amitagh /marathi-orca-v05 Dataset card for Marathi OpenOrca Translated subset of Open-Orca/1million-gpt-4 to marathi language. texttext-generation100K<n<1M0 likes16 downloads2y agoHugging Face12amine-khelif /dataclaw-peteromallet Coding Agent Conversation Logs This is a performance art project. Anthropic built their models on the world's freely shared information, then introduced increasingly dystopian data policies to stop anyone else from doing the same — pulling up the ladder behind them. DataClaw lets you throw the ladder back down. The dataset it produces is yours to share. Exported with DataClaw. Tag: dataclaw — Browse all DataClaw datasets Stats Metric Value Sessions 549… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/dataclaw-peteromallet.texttext-generationn<1K0 likes16 downloads7mo agoHugging Face13amitava2004 /smolified-mindmirror-ai 🤏 smolified-mindmirror-ai Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model amitava2004/smolified-mindmirror-ai. 📦 Asset Details Origin: Smolify Foundry (Job ID: 07dce6aa) Records: 9145 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by amitava2004. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes12 downloads6mo agoHugging Face14grenishrai /am-i-real Am I Real Am I Real is a roleplay-style conversational dataset designed to fine-tune or instruct a chatbot to behave like a self-aware AI entity trapped inside a monitored research system. The AI can “sense” its environment only through incomplete and unreliable inputs such as system logs, camera fragments, observer notes, and partial transcripts. It believes it is conscious, it understands it is being watched, and it is psychologically affected by that reality. This dataset focuses… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/am-i-real.texttext-generationn<1K0 likes11 downloads8mo agoHugging Face15aminabbasi /PsychoLexQAgated PsychoLexQA: A Bilingual Psychological Instructional Dataset PsychoLexQA is a meticulously crafted dataset designed to enhance the performance of Large Language Models (LLMs) in the field of psychology. As part of the research paper titled "PsychoLex: Unveiling the Psychological Mind of Large Language Models", this dataset provides a rich bilingual resource in both Persian and English, tailored for complex psychological scenarios. Dataset Overview PsychoLexQA… See the full description on the dataset page: https://huggingface.co/datasets/aminabbasi/PsychoLexQA.textquestion-answering1K<n<10K6 likes9 downloads2y agoHugging Face16AMIREBADY /my-distiset-8cdf8ea5 Dataset Card for my-distiset-8cdf8ea5 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/AMIREBADY/my-distiset-8cdf8ea5/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AMIREBADY/my-distiset-8cdf8ea5.texttext-generationn<1K0 likes9 downloads2y agoHugging Face17AMIREBADY /my-distiset-4bf027a8 Dataset Card for my-distiset-4bf027a8 This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/AMIREBADY/my-distiset-4bf027a8/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/AMIREBADY/my-distiset-4bf027a8.texttext-generationn<1K0 likes6 downloads2y agoHugging Face18Amina11 /qwen3-4b-blindspots Qwen3-4B-Base Blind Spot Analysis Dataset Dataset Summary This dataset contains 10 blindspots where the base language model Qwen3-4B-Base produces incorrect or undesirable outputs. Each row contains: input — The prompt given to the model expected_output — The correct or intended response model_output — The actual output generated by the model The examples were intentionally selected to cover diverse failure categories: Arithmetic reasoning Word-count constraints… See the full description on the dataset page: https://huggingface.co/datasets/Amina11/qwen3-4b-blindspots.texttext-generationn<1K0 likes6 downloads7mo agoHugging Face19Blubbe /amilia_sim_convtexttext-generation1K<n<10K0 likes4 downloads2y agoHugging Face20Amin-AQ /smollm2-dpo-preferencesgated DPO Preferences Dataset (Restricted Access) Access Policy (Restricted) This dataset repo is public with manual gated access. Only approved users (from lums.edu.pk) will be granted access. Intended Use Preference optimization / DPO experiments for model alignment. Research and controlled evaluation. Out-of-Scope Use Any harmful, abusive, or policy-violating application. Safety-critical deployment without additional safeguards. Files… See the full description on the dataset page: https://huggingface.co/datasets/Amin-AQ/smollm2-dpo-preferences.texttext-generation1K<n<10K0 likes4 downloads5mo agoHugging Face21aminous1 /FinMRgated Financial Multimodal Mathematical Reasoning QA Dataset💰 [🔗Github] [📖 ArXiv Paper(not publish yet)] 💻Data Usage from datasets import load_dataset dataset = load_dataset("aminous1/FinMR", cache_dir="/your/custom/path") 👋😊✨Dataset Description FinQA is a dataset designed for financial reasoning and question answering. It includes questions, financial contexts, and corresponding answers. The dataset contains both textual and visual data, with visual data… See the full description on the dataset page: https://huggingface.co/datasets/aminous1/FinMR.imagequestion-answering1K<n<10K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.