CoolFace
Datasetpublic

ThatsGroes/synthetic-from-text-matching-long-tasks-swedish

Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset The purpose of this dataset is to pre- or post-train embedding models for text matching tasks. The dataset consists of 100,000 samples generated with gemma-2-27b-it. The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output. Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-matching-long-tasks-swedish.

sourceHugging Facemitupdated 2y agoView on Hugging Face
1likes43downloads
Dataset Card

Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset

The purpose of this dataset is to pre- or post-train embedding models for text matching tasks.

The dataset consists of 100,000 samples generated with gemma-2-27b-it.

The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.

Each sample in the dataset was generated from a seed task randomly sampled from https://huggingface.co/datasets/ThatsGroes/text-matching-long-tasks-processed

The data generation process described in this paper was followed:

https://arxiv.org/pdf/2401.00368

Compute sponsored by Arrow Denmark and Nvidia through Danish Data Science Community.