datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude-fable-5-code
Claude Fable 5 Coding and Math Dataset (Non-Thinking)
This repository contains a dataset of 603 coding and math-related prompts and responses from Claude Fable 5.
The generation of this dataset cost approximately $75.
Please note that this dataset is non-thinking. Fable 5 only supported adaptive thinking, and it decided not to think for these prompts, meaning there is no chain-of-thought/reasoning content in this dataset.
Origin of Prompts
The prompts in this… See the full description on the dataset page: https://huggingface.co/datasets/PawanKrd/claude-fable-5-code.PawolKreyol-gfc
Analyse Lexicographique du Kreyòl Guadeloupéen
Métadonnées du Corpus
Date de génération : 06 November 2025 à 20:55
Version du pipeline : 3.0 - Pipeline Unique
Source des données : Dataset POTOMITAN/PawolKreyol-gfc (Hugging Face)
Nombre de textes : 427
Tokens totaux : 22,058
Types lexicaux : 3,680
1. Corpus et Échantillonnage
1.1 Taille et Couverture
Total des tokens : 239,808
Types lexicaux uniques : 3,680
Type-Token Ratio (TTR) :… See the full description on the dataset page: https://huggingface.co/datasets/POTOMITAN/PawolKreyol-gfc.PAWS-eu
Dataset Card for PAWS-eu
Point of Contact: hitz@ehu.eus
Dataset Description
Dataset Summary
PAWS-eu is the professional translation to Basque of the PAWS dataset (Zhang et al., 2019),
in the spirit of the PAWS-X effort (Yang et al., 2019).
PAWS consist of sentence pairs that have high lexical overlap but that may or may not be paraphrases.
Languages
eu-ES
Dataset Structure
Data Fields
id (str): A unique id for each pair.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/PAWS-eu.small-lean
Small Lean Alpaca
Thirty filtered Alpaca-style Lean 4 theorem-proving records derived from
internlm/Lean-Workbook.
Only train.jsonl is a Hub dataset split. The hf_dataset/ directory is a
local datasets.save_to_disk() artifact and must not be interpreted as JSON
training data.
paws-jsonl
Introduction
This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and
PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed.
Each line contains a dict in the following format:
{"guid": <id>, "texts": [anchor, positive]} or
{"guid": <id>, "texts": [anchor, positive, negative]}
positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.bbh-logical-deduction-seven-objects-pltranslated-PAWSpaws-x-italian
PAWS-X Italian Paraphrase Dataset
This dataset is a machine-translated Italian version of the English PAWS-X dataset. The original PAWS-X dataset (Yang et al. 2019) is a multilingual version of PAWS (Zhang et al. 2019) for paraphrase identification.
Dataset Structure
Data Fields
sentence1: First sentence in the pair
sentence2: Second sentence in the pair
labels:
0: Non-paraphrases
1: Paraphrases
Data Splits
The dataset is split into:
Training… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/paws-x-italian.otwarte-pytania-matura-ckefinnlp_task1_with_rationaleotwarte-pytania-matura-cke-100paws
Paws
A small dataset of Russian-language instructions where target is handwritten
Task types
instruct, a specific task with a single correct answer
rewrite, rewriting a text while maintaining the main meaning but using different words and structures
creative, creating something creative and original in the process of work
qa, answering an open or closed question
self-identification, awareness and determination of one's personal and professional goals and values
llmzszl-open-endedSillyTilly_PawanKrd-dpo-gpt-4o-reup-PreferenceShareGPTbbh-logical-deduction-seven-objects-pl-100PAWBench-A09-LastFrames
PAWBench A-09 terminal-frame controls
This public, ungated transport repository contains one normalized A-09 first frame and 40 formal terminal-frame controls: 20 falls_left and 20 falls_right, arranged as 20 matched endpoint pairs. They are production controls for constructing the fixed 200-row Seedance 2.0 first/last-frame candidate-video bank in PhysEdit issue #297.
What this is—and is not
The endpoint bank is owner-accepted only for candidate-video generation.… See the full description on the dataset page: https://huggingface.co/datasets/Andrew613/PAWBench-A09-LastFrames.ifeval-pl-200PawcatuckNeighborhoodCenterjokemachine
JokeMachine Dataset
The JokeMachine dataset contains short-form comedic responses generated in a stand-up comedy style. Each row consists of a prompt and a response, intended for training language models in humorous text generation.
Dataset Structure
Fields:
prompt: Always "write a joke" — used as a standard prompt for consistency.
response: The generated joke or humorous response (1+ sentences).
Split:
train: All available rows are in the training set.… See the full description on the dataset page: https://huggingface.co/datasets/pawneeranger/jokemachine.ifeval-pllangchain-sample-dsVeterinairy_PawPal_ClinicRekomendasi_PromosiTestEMI
