text-matching
jina-embeddings-v5-text-small-text-matchingjina-embeddings-v5-text-nano-text-matchingjina-embeddings-v4-text-matching-GGUFjina-embeddings-v5-omni-small-text-matching-GGUFjina-embeddings-v5-omni-nano-text-matching-GGUFjina-embeddings-v5-omni-small-text-matchingjina-embeddings-v5-text-nano-text-matching-GGUFjina-embeddings-v5-text-small-text-matching-GGUF
synthetic-from-text-matching-short-tasks-danish
Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset
The purpose of this dataset is to pre- or post-train embedding models for Danish text matching tasks on short texts.
The dataset consists of 100,000 samples generated with gemma-2-27b-it.
The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.
Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-matching-short-tasks-danish.synthetic-from-text-matching-long-tasks-swedish
Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset
The purpose of this dataset is to pre- or post-train embedding models for text matching tasks.
The dataset consists of 100,000 samples generated with gemma-2-27b-it.
The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.
Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-matching-long-tasks-swedish.synthetic-nordic-text_matchingsynthetic-from-text-matching-long-tasks-danish
Thanks to Arrow Denmark and Nvidia for sponsoring the compute used to generate this dataset
The purpose of this dataset is to pre- or post-train embedding models for Danish text matching tasks.
The dataset consists of 100,000 samples generated with gemma-2-27b-it.
The column "prompt" shows the prompt given to the LLM and "response" shows the LLM output.
Each sample in the dataset was generated from a seed task randomly sampled from… See the full description on the dataset page: https://huggingface.co/datasets/ThatsGroes/synthetic-from-text-matching-long-tasks-danish.multilingual-text-matchingtext-matching-long-tasks-processed
