datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Witch_Bowl
Dataset Card for The Cauldron
Dataset description
The Cauldron is part of the Idefics2 release.
It is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
Load the dataset
To load the dataset, install the library datasets with pip install datasets. Then,
from datasets import load_dataset
ds = load_dataset("HuggingFaceM4/the_cauldron", "ai2d")
to download and load the… See the full description on the dataset page: https://huggingface.co/datasets/Compumacy/Witch_Bowl.Witcher-GRPO-promptsWitcher-multilingual-convosLKF-unlearning_Salem_Witch_Trialstokenized_dataset_bart_fblarge
Dataset Card for "tokenized_dataset_bart_fblarge"
More Information needed
Witcher-synth-multi-round-instructWitcher-fandom-instruct-datasetLKF-unlearning_Salem_Witch_trials_rephrasings_finalhybrid_data_fin
Dataset Card for "hybrid_data_fin"
More Information needed
tokenized_dataset_bart
Dataset Card for "tokenized_dataset_bart"
More Information needed
witcher3-dataset-ptbrada_002_embeddings
Dataset Card for "ada_002_embeddings"
More Information needed
tokenized_T5_base
Dataset Card for "tokenized_T5_base"
More Information needed
Witcher-synth-instruct-datasetWitcher-GRPO-answersSmol-Witcher-pretrainingWitcher-synth-fandom-summariescustomWitcher-synth-fandom-summaries-instruct
