CoolFace
Datasetpublic

ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian

Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset The purpose of this dataset is to pre- or post-train embedding models for text matching tasks on short texts. The dataset consists of 100,000 samples generated with gemma-2-27b-it. The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output. Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes20downloads
settings

This repository belongs to ThatsGroes on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namesynthetic-from-text-mathing-short-tasks-norwegian
visibilitypublic
licencemit
gatedno
ownerThatsGroes
Account settings