datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_plus
Dataset Card for FLORES+
FLORES+ is an evaluation benchmark dataset for multilingual machine translation.
Dataset Details
Dataset Description
FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.grok-demon-dataset-ESFLORES-200nq_open_gold
Natural Questions Open Dataset with Gold Documents
This dataset is a curated version of the Natural Questions open dataset,
with the inclusion of the gold documents from the original Natural Questions (NQ) dataset.
The main difference with the NQ-open dataset is that some entries were excluded, as their respective gold documents exceeded 512 tokens in length.
This is due to the pre-processing of the gold documents, as detailed in this related dataset.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/nq_open_gold.FLORA-Bench
Field Descriptions
label
Type: integer (0, 1)
Description: An integer flag that indicates the success of the workflow in a given task. A value of 1 signify that the workflow completed successfully.
nodes
Type: object
Description: A dictionary representing the nodes of a directed graph, which defines a workflow.
Key: A string representing the unique ID of a node (e.g., "0", "1").
Value: A string containing the system prompt of the specific agent. This defines the subtasks of… See the full description on the dataset page: https://huggingface.co/datasets/YuanshuoZhang/FLORA-Bench.gov-report-qs-llama2-format
Government Report Question Answering Dataset in LLAMA2 Format
Dataset Description
This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office.
The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.agentlans-combined-roleplay_Dataset
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.grok-demon-dataset-ENQuixiAI-dolphin_DatasetDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.cloudbjorn-eschaton-uncensored_Dataset
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.Emakhuwa-FLORES
Dataset card
Description
FLORES+ dev and devtest set in Emakhuwa
License
CC-BY-SA-4.0
Attribution
@inproceedings{ali-etal-2024-expanding,
title = "Expanding {FLORES}+ Benchmark for More Low-Resource Settings: {P}ortuguese-Emakhuwa Machine Translation Evaluation",
author = "Ali, Felermino Dario Mario and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Haddow, Barry and
Kocmi, Tom and
Koehn, Philipp… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-FLORES.rombodawg-OpenHermes-2.5-Uncensored_Dataset
This is the teknium/OpenHermes-2.5 dataset with 2,697 censored lines removed using my uncensored code found bellow.
https://huggingface.co/datasets/rombodawg/data_processing_code
Thank you teknium for the original dataset, you can find it bellow.
https://huggingface.co/datasets/teknium/OpenHermes-2.5
This is the same version of Open-Hermes-2.5 that was used in code_bagel_hermes-2.5 found bellow:… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/rombodawg-OpenHermes-2.5-Uncensored_Dataset.flores-kr-reversed
FLORES Parallel Mix (Direction-Flipped)
Source dataset dir: /Users/inertia/Desktop/preprocessing/processed/flores_parallel_80_20
Transformation: direction flip per row (source/target swap)
Pair+sample_key overlap with source: 0
Primary: en<->ko (kept disjoint by inherited sample assignment)
Auxiliary(flipped): jpn_Jpan->eng_Latn, zho_Hans->eng_Latn, zho_Hans->jpn_Jpan, jpn_Jpan->zho_Hans, kor_Hang->zho_Hans, kor_Hang->jpn_Jpan
Target primary ratio: 0.8000
Achieved primary ratio:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/flores-kr-reversed.jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564
jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 Dataset
Dataset Description
jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 model.
How to Use
To use this dataset for model training or evaluation… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564.medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564
medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 Dataset
Dataset Description
medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 model.
How to… See the full description on the dataset page: https://huggingface.co/datasets/florian-hoenicke/medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564.flo
flo Dataset
Dataset Description
The dataset "general domain" is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the flo model.
How to Use
To use this dataset for model training or evaluation, you can load it using the Hugging Face datasets library as follows:
from datasets import load_dataset
dataset = load_dataset("florianhoenicke/flo")… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/flo.flores-plusplus-blocks
FLORES++ blocks
Blocks of k = 1..5 consecutive segments of one FLORES+ article (dev + devtest),
English source, with the correct translation (the reference) and its
distractors: copies of the reference in which every segment carries one
perturbation from the three xSIM++ categories (causality, entity, number;
Chen et al., 2023), in every possible combination (up to 243 at k = 5).
The error share is one per segment at every k, so if a model finds the reference less
often on… See the full description on the dataset page: https://huggingface.co/datasets/AdleBenSalem/flores-plusplus-blocks.Flores-Plus-Evaluation-Log-Preview-CleanedSudanese_Floresjina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564
jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset
Dataset Description
jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 model.
How to Use
To use this dataset for model training or evaluation, you can load it… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564.pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564
pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 Dataset
Dataset Description
pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 model.
How to Use
To use this dataset for model training or… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564.Grounding-FT-70KSudanese-Floresllimba-flores-srd-eval
LLiMba FLORES-200 Sardinian Evaluation Set
A held-out evaluation set of 997 parallel sentences aligned across six languages (Sardinian, Italian, English, Spanish, French, Portuguese), derived from FLORES-200. Used to benchmark the LLiMba model's translation quality and reported in the LLiMba paper's BLEU and chrF tables.
This is the exact evaluation set used to produce the published LLiMba translation results. Reproducing the paper's numbers requires this set plus… See the full description on the dataset page: https://huggingface.co/datasets/lballore/llimba-flores-srd-eval.medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564
medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 Dataset
Dataset Description
medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 model.
How to… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564.codigo_florestal_lei_25_05_12jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564
jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset
Dataset Description
jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 model.
How to Use
To use this dataset for model training or evaluation, you… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564.jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564
jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset
Dataset Description
jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 model.
How to Use
To use this dataset for model training or evaluation, you can load… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564.medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564
medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 Dataset
Dataset Description
medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 model.
How to… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564.jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564
jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 Dataset
Dataset Description
jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 model.
How to Use
To use this dataset for model training or evaluation, you can load it… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564.
