CoolFace
Datasetpublic

ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian

Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset The purpose of this dataset is to pre- or post-train embedding models for text matching tasks on short texts. The dataset consists of 100,000 samples generated with gemma-2-27b-it. The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output. Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes20downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
ThatsGroes/synthetic-from-text-mathing-short-tasks-norwegian · CoolFace