datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_plus
Dataset Card for FLORES+
FLORES+ is an evaluation benchmark dataset for multilingual machine translation.
Dataset Details
Dataset Description
FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.grok-demon-dataset-ESFLORES-200QuixiAI-dolphin_DatasetDolphin 🐬
https://erichartford.com/dolphin
Dataset details
This dataset is an attempt to replicate the results of Microsoft's Orca
Our dataset consists of:
~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl)
~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl)
We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.grok-demon-dataset-ENagentlans-combined-roleplay_Dataset
Combined Roleplay Dataset
This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks.
Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays
English content with a few Spanish, Portuguese, and Chinese conversations
Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.cloudbjorn-eschaton-uncensored_Dataset
Eschaton Uncensored SFT Dataset
Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers.
The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.Emakhuwa-FLORES
Dataset card
Description
FLORES+ dev and devtest set in Emakhuwa
License
CC-BY-SA-4.0
Attribution
@inproceedings{ali-etal-2024-expanding,
title = "Expanding {FLORES}+ Benchmark for More Low-Resource Settings: {P}ortuguese-Emakhuwa Machine Translation Evaluation",
author = "Ali, Felermino Dario Mario and
Lopes Cardoso, Henrique and
Sousa-Silva, Rui",
editor = "Haddow, Barry and
Kocmi, Tom and
Koehn, Philipp… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-FLORES.rombodawg-OpenHermes-2.5-Uncensored_Dataset
This is the teknium/OpenHermes-2.5 dataset with 2,697 censored lines removed using my uncensored code found bellow.
https://huggingface.co/datasets/rombodawg/data_processing_code
Thank you teknium for the original dataset, you can find it bellow.
https://huggingface.co/datasets/teknium/OpenHermes-2.5
This is the same version of Open-Hermes-2.5 that was used in code_bagel_hermes-2.5 found bellow:… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/rombodawg-OpenHermes-2.5-Uncensored_Dataset.flores-kr-reversed
FLORES Parallel Mix (Direction-Flipped)
Source dataset dir: /Users/inertia/Desktop/preprocessing/processed/flores_parallel_80_20
Transformation: direction flip per row (source/target swap)
Pair+sample_key overlap with source: 0
Primary: en<->ko (kept disjoint by inherited sample assignment)
Auxiliary(flipped): jpn_Jpan->eng_Latn, zho_Hans->eng_Latn, zho_Hans->jpn_Jpan, jpn_Jpan->zho_Hans, kor_Hang->zho_Hans, kor_Hang->jpn_Jpan
Target primary ratio: 0.8000
Achieved primary ratio:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/flores-kr-reversed.flores-plusplus-blocks
FLORES++ blocks
Blocks of k = 1..5 consecutive segments of one FLORES+ article (dev + devtest),
English source, with the correct translation (the reference) and its
distractors: copies of the reference in which every segment carries one
perturbation from the three xSIM++ categories (causality, entity, number;
Chen et al., 2023), in every possible combination (up to 243 at k = 5).
The error share is one per segment at every k, so if a model finds the reference less
often on… See the full description on the dataset page: https://huggingface.co/datasets/AdleBenSalem/flores-plusplus-blocks.Flores-Plus-Evaluation-Log-Preview-CleanedSudanese_FloresSudanese-Floresllimba-flores-srd-eval
LLiMba FLORES-200 Sardinian Evaluation Set
A held-out evaluation set of 997 parallel sentences aligned across six languages (Sardinian, Italian, English, Spanish, French, Portuguese), derived from FLORES-200. Used to benchmark the LLiMba model's translation quality and reported in the LLiMba paper's BLEU and chrF tables.
This is the exact evaluation set used to produce the published LLiMba translation results. Reproducing the paper's numbers requires this set plus… See the full description on the dataset page: https://huggingface.co/datasets/lballore/llimba-flores-srd-eval.codigo_florestal_lei_25_05_12flores200zh-en
Dataset Card for Dataset Name
Chinese English human translation pairs from the flores200 devtest set for WMT24
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/GTmodlang/flores200zh-en.ManyToDanishTranslations-flores
Danske oversættelser
Tak til facebook for deres facebook/flores (CC-BY-SA-4.0) dataset som er udgivet under et copyleft licens.
en-si-translation-flores-factual-2k
En Si Translation Flores Factual 2K
Dataset Summary
English-Sinhala Factual Translation dataset containing ~2,000 highly accurate sentences covering diverse factual domains, curated from the FLORES+ benchmark.
Engineering Pipeline Parameters
Language Pair: English (en) to Sinhala (si)
Total Valid Token Rows: 2000
Internal Storage Structure: Single-File data.json
Upstream Source Attribution
This specific sub-split was compiled and extracted from… See the full description on the dataset page: https://huggingface.co/datasets/SAWithanage/en-si-translation-flores-factual-2k.flores_for_translation
