CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PawanKrd /claude-fable-5-code Claude Fable 5 Coding and Math Dataset (Non-Thinking) This repository contains a dataset of 603 coding and math-related prompts and responses from Claude Fable 5. The generation of this dataset cost approximately $75. Please note that this dataset is non-thinking. Fable 5 only supported adaptive thinking, and it decided not to think for these prompts, meaning there is no chain-of-thought/reasoning content in this dataset. Origin of Prompts The prompts in this… See the full description on the dataset page: https://huggingface.co/datasets/PawanKrd/claude-fable-5-code.texttext-generationn<1K33 likes96 downloads3mo agoHugging Face02POTOMITAN /PawolKreyol-gfc Analyse Lexicographique du Kreyòl Guadeloupéen Métadonnées du Corpus Date de génération : 06 November 2025 à 20:55 Version du pipeline : 3.0 - Pipeline Unique Source des données : Dataset POTOMITAN/PawolKreyol-gfc (Hugging Face) Nombre de textes : 427 Tokens totaux : 22,058 Types lexicaux : 3,680 1. Corpus et Échantillonnage 1.1 Taille et Couverture Total des tokens : 239,808 Types lexicaux uniques : 3,680 Type-Token Ratio (TTR) :… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/PawolKreyol-gfc.text1K<n<10K1 likes91 downloads1mo agoHugging Face03HiTZ /PAWS-eu Dataset Card for PAWS-eu Point of Contact: hitz@ehu.eus Dataset Description Dataset Summary PAWS-eu is the professional translation to Basque of the PAWS dataset (Zhang et al., 2019), in the spirit of the PAWS-X effort (Yang et al., 2019). PAWS consist of sentence pairs that have high lexical overlap but that may or may not be paraphrases. Languages eu-ES Dataset Structure Data Fields id (str): A unique id for each pair.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/PAWS-eu.texttext-classification1K<n<10K0 likes65 downloads2y agoHugging Face04Pawitt /small-lean Small Lean Alpaca Thirty filtered Alpaca-style Lean 4 theorem-proving records derived from internlm/Lean-Workbook. Only train.jsonl is a Hub dataset split. The hf_dataset/ directory is a local datasets.save_to_disk() artifact and must not be interpreted as JSON training data. texttext-generationn<1K0 likes56 downloads2mo agoHugging Face05flax-sentence-embeddings /paws-jsonl Introduction This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed. Each line contains a dict in the following format: {"guid": <id>, "texts": [anchor, positive]} or {"guid": <id>, "texts": [anchor, positive, negative]} positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.text10K<n<100K1 likes37 downloads5y agoHugging Face06pawel04 /bbh-logical-deduction-seven-objects-pltextn<1K1 likes28 downloads5mo agoHugging Face07huseinzol05 /translated-PAWStext10K<n<100K0 likes22 downloads4y agoHugging Face08ZurichNLP /paws-x-italian PAWS-X Italian Paraphrase Dataset This dataset is a machine-translated Italian version of the English PAWS-X dataset. The original PAWS-X dataset (Yang et al. 2019) is a multilingual version of PAWS (Zhang et al. 2019) for paraphrase identification. Dataset Structure Data Fields sentence1: First sentence in the pair sentence2: Second sentence in the pair labels: 0: Non-paraphrases 1: Paraphrases Data Splits The dataset is split into: Training… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/paws-x-italian.text10K<n<100K0 likes21 downloads10mo agoHugging Face09pawel04 /otwarte-pytania-matura-cketextn<1K1 likes14 downloads4mo agoHugging Face10pawan2411 /finnlp_task1_with_rationaletext1K<n<10K1 likes12 downloads2y agoHugging Face11pawel04 /otwarte-pytania-matura-cke-100textn<1K1 likes12 downloads4mo agoHugging Face12avdosev /paws Paws A small dataset of Russian-language instructions where target is handwritten Task types instruct, a specific task with a single correct answer rewrite, rewriting a text while maintaining the main meaning but using different words and structures creative, creating something creative and original in the process of work qa, answering an open or closed question self-identification, awareness and determination of one's personal and professional goals and values textn<1K0 likes11 downloads1y agoHugging Face13pawel04 /llmzszl-open-endedtextn<1K1 likes11 downloads5mo agoHugging Face14PJMixers /SillyTilly_PawanKrd-dpo-gpt-4o-reup-PreferenceShareGPTtextreinforcement-learning10K<n<100K0 likes9 downloads2y agoHugging Face15pawel04 /bbh-logical-deduction-seven-objects-pl-100textn<1K1 likes9 downloads3mo agoHugging Face16Andrew613 /PAWBench-A09-LastFramesgated PAWBench A-09 terminal-frame controls This public, ungated transport repository contains one normalized A-09 first frame and 40 formal terminal-frame controls: 20 falls_left and 20 falls_right, arranged as 20 matched endpoint pairs. They are production controls for constructing the fixed 200-row Seedance 2.0 first/last-frame candidate-video bank in PhysEdit issue #297. What this is—and is not The endpoint bank is owner-accepted only for candidate-video generation.… See the full description on the dataset page: https://huggingface.co/datasets/Andrew613/PAWBench-A09-LastFrames.imageimage-to-videon<1K0 likes6 downloads3mo agoHugging Face17pawel04 /ifeval-pl-200textn<1K1 likes5 downloads4mo agoHugging Face18Dmevanct /PawcatuckNeighborhoodCentertextn<1K0 likes4 downloads3y agoHugging Face19pawneeranger /jokemachine JokeMachine Dataset The JokeMachine dataset contains short-form comedic responses generated in a stand-up comedy style. Each row consists of a prompt and a response, intended for training language models in humorous text generation. Dataset Structure Fields: prompt: Always "write a joke" — used as a standard prompt for consistency. response: The generated joke or humorous response (1+ sentences). Split: train: All available rows are in the training set.… See the full description on the dataset page: https://huggingface.co/datasets/pawneeranger/jokemachine.texttext-generationn<1K0 likes4 downloads1y agoHugging Face20pawel04 /ifeval-pltextn<1K1 likes4 downloads4mo agoHugging Face21pawanrawat0926 /langchain-sample-dstextn<1K0 likes2 downloads3y agoHugging Face22Lollipop-Lemon /Veterinairy_PawPal_Clinictextn<1K2 likes2 downloads2y agoHugging Face23pawlenn /Rekomendasi_Promositextn<1K0 likes2 downloads1y agoHugging Face24pawansharma-kapture /TestEMItextn<1K0 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.