CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openlanguagedata /flores_plusgated Dataset Card for FLORES+ FLORES+ is an evaluation benchmark dataset for multilingual machine translation. Dataset Details Dataset Description FLORES+ is a multilingual machine translation benchmark released under CC BY-SA 4.0. This dataset was originally released by FAIR researchers at Meta under the name FLORES. Further information about these initial releases can be found in Dataset Sources below. The data is now being managed by OLDI, the Open… See the full description on the dataset page: https://huggingface.co/datasets/openlanguagedata/flores_plus.tabulartext-generation100K<n<1M170 likes13k downloads2mo agoHugging Face02Maximiliano-Flores-Dev /grok-demon-dataset-EStext1K<n<10K1 likes214 downloads8d agoHugging Face03DGME /FLORES-200texttranslation100K<n<1M0 likes208 downloads9mo agoHugging Face04florin-hf /nq_open_gold Natural Questions Open Dataset with Gold Documents This dataset is a curated version of the Natural Questions open dataset, with the inclusion of the gold documents from the original Natural Questions (NQ) dataset. The main difference with the NQ-open dataset is that some entries were excluded, as their respective gold documents exceeded 512 tokens in length. This is due to the pre-processing of the gold documents, as detailed in this related dataset. The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/florin-hf/nq_open_gold.tabularquestion-answering10K<n<100K2 likes147 downloads2y agoHugging Face05YuanshuoZhang /FLORA-Bench Field Descriptions label Type: integer (0, 1) Description: An integer flag that indicates the success of the workflow in a given task. A value of 1 signify that the workflow completed successfully. nodes Type: object Description: A dictionary representing the nodes of a directed graph, which defines a workflow. Key: A string representing the unique ID of a node (e.g., "0", "1"). Value: A string containing the system prompt of the specific agent. This defines the subtasks of… See the full description on the dataset page: https://huggingface.co/datasets/YuanshuoZhang/FLORA-Bench.tabulartext-classification100K<n<1M1 likes130 downloads1y agoHugging Face06Kira-Floris /gov-report-qs-llama2-format Government Report Question Answering Dataset in LLAMA2 Format Dataset Description This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office. The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.textquestion-answering10K<n<100K2 likes84 downloads3y agoHugging Face07Maximiliano-Flores-Dev /agentlans-combined-roleplay_Dataset Combined Roleplay Dataset This dataset combines multi-turn conversations across various AI assistant interactions, creative writing scenarios, and roleplaying exchanges. It aims to improve language models' performance in interactive tasks. Multi-turn conversations with a mix of standard AI assistant interactions, creative writing prompts, and roleplays English content with a few Spanish, Portuguese, and Chinese conversations Conversations limited to 4000 tokens using the Llama… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/agentlans-combined-roleplay_Dataset.texttext-generation1M<n<10M0 likes69 downloads3d agoHugging Face08Maximiliano-Flores-Dev /grok-demon-dataset-ENtext1K<n<10K2 likes66 downloads5d agoHugging Face09Maximiliano-Flores-Dev /QuixiAI-dolphin_DatasetDolphin 🐬 https://erichartford.com/dolphin Dataset details This dataset is an attempt to replicate the results of Microsoft's Orca Our dataset consists of: ~1 million of FLANv2 augmented with GPT-4 completions (flan1m-alpaca-uncensored.jsonl) ~3.5 million of FLANv2 augmented with GPT-3.5 completions (flan5m-alpaca-uncensored.jsonl) We followed the submix and system prompt distribution outlined in the Orca paper. With a few exceptions. We included all 75k of CoT in the FLAN-1m… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/QuixiAI-dolphin_Dataset.texttext-generation1M<n<10M1 likes66 downloads3d agoHugging Face10Maximiliano-Flores-Dev /cloudbjorn-eschaton-uncensored_Dataset Eschaton Uncensored SFT Dataset Eschaton is a 1,000-row supervised fine-tuning dataset for teaching language models to answer sensitive, gritty, controversial, emotionally intimate, technically demanding, and morally difficult prompts directly. Its emphasis is human connection, unvarnished analysis, malware analysis, dark creative work, profanity, and high-friction subject matter without repetitive moralizing or canned disclaimers. The dataset is deliberately model-agnostic. It… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/cloudbjorn-eschaton-uncensored_Dataset.texttext-generation1K<n<10K0 likes56 downloads3d agoHugging Face11LIACC /Emakhuwa-FLORES Dataset card Description FLORES+ dev and devtest set in Emakhuwa License CC-BY-SA-4.0 Attribution @inproceedings{ali-etal-2024-expanding, title = "Expanding {FLORES}+ Benchmark for More Low-Resource Settings: {P}ortuguese-Emakhuwa Machine Translation Evaluation", author = "Ali, Felermino Dario Mario and Lopes Cardoso, Henrique and Sousa-Silva, Rui", editor = "Haddow, Barry and Kocmi, Tom and Koehn, Philipp… See the full description on the dataset page: https://huggingface.co/datasets/LIACC/Emakhuwa-FLORES.text1K<n<10K0 likes49 downloads2y agoHugging Face12Maximiliano-Flores-Dev /rombodawg-OpenHermes-2.5-Uncensored_Dataset This is the teknium/OpenHermes-2.5 dataset with 2,697 censored lines removed using my uncensored code found bellow. https://huggingface.co/datasets/rombodawg/data_processing_code Thank you teknium for the original dataset, you can find it bellow. https://huggingface.co/datasets/teknium/OpenHermes-2.5 This is the same version of Open-Hermes-2.5 that was used in code_bagel_hermes-2.5 found bellow:… See the full description on the dataset page: https://huggingface.co/datasets/Maximiliano-Flores-Dev/rombodawg-OpenHermes-2.5-Uncensored_Dataset.text100K<n<1M0 likes46 downloads3d agoHugging Face13alwaysgood /flores-kr-reversed FLORES Parallel Mix (Direction-Flipped) Source dataset dir: /Users/inertia/Desktop/preprocessing/processed/flores_parallel_80_20 Transformation: direction flip per row (source/target swap) Pair+sample_key overlap with source: 0 Primary: en<->ko (kept disjoint by inherited sample assignment) Auxiliary(flipped): jpn_Jpan->eng_Latn, zho_Hans->eng_Latn, zho_Hans->jpn_Jpan, jpn_Jpan->zho_Hans, kor_Hang->zho_Hans, kor_Hang->jpn_Jpan Target primary ratio: 0.8000 Achieved primary ratio:… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/flores-kr-reversed.texttranslation1K<n<10K0 likes34 downloads5mo agoHugging Face14florianhoenicke /jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 Dataset Dataset Description jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564 model. How to Use To use this dataset for model training or evaluation… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-1000_9062874564.textn<1K0 likes31 downloads2y agoHugging Face15florian-hoenicke /medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 Dataset Dataset Description medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 model. How to… See the full description on the dataset page: https://huggingface.co/datasets/florian-hoenicke/medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564.textn<1K0 likes28 downloads2y agoHugging Face16florianhoenicke /flo flo Dataset Dataset Description The dataset "general domain" is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the flo model. How to Use To use this dataset for model training or evaluation, you can load it using the Hugging Face datasets library as follows: from datasets import load_dataset dataset = load_dataset("florianhoenicke/flo")… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/flo.textfeature-extractionn<1K0 likes28 downloads2y agoHugging Face17AdleBenSalem /flores-plusplus-blocksgated FLORES++ blocks Blocks of k = 1..5 consecutive segments of one FLORES+ article (dev + devtest), English source, with the correct translation (the reference) and its distractors: copies of the reference in which every segment carries one perturbation from the three xSIM++ categories (causality, entity, number; Chen et al., 2023), in every possible combination (up to 243 at k = 5). The error share is one per segment at every k, so if a model finds the reference less often on… See the full description on the dataset page: https://huggingface.co/datasets/AdleBenSalem/flores-plusplus-blocks.tabulartranslation1K<n<10K0 likes25 downloads4d agoHugging Face18sailor2 /Flores-Plus-Evaluation-Log-Preview-Cleanedtext100K<n<1M0 likes22 downloads2y agoHugging Face19McGill-NLP /Sudanese_Florestext1K<n<10K1 likes22 downloads5mo agoHugging Face20florianhoenicke /jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you can load it… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-0-16-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes21 downloads2y agoHugging Face21florianhoenicke /pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 Dataset Dataset Description pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564 model. How to Use To use this dataset for model training or… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/pet-shop-100-64-10-jinaai_jina-embeddings-v2-small-en_9062874564.textn<1K0 likes18 downloads2y agoHugging Face22Florence123 /Grounding-FT-70Ktext10K<n<100K0 likes18 downloads10mo agoHugging Face23McGill-NLP /Sudanese-Florestext1K<n<10K0 likes17 downloads5mo agoHugging Face24lballore /llimba-flores-srd-eval LLiMba FLORES-200 Sardinian Evaluation Set A held-out evaluation set of 997 parallel sentences aligned across six languages (Sardinian, Italian, English, Spanish, French, Portuguese), derived from FLORES-200. Used to benchmark the LLiMba model's translation quality and reported in the LLiMba paper's BLEU and chrF tables. This is the exact evaluation set used to produce the published LLiMba translation results. Reproducing the paper's numbers requires this set plus… See the full description on the dataset page: https://huggingface.co/datasets/lballore/llimba-flores-srd-eval.texttranslationn<1K0 likes17 downloads5mo agoHugging Face25florianhoenicke /medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 Dataset Dataset Description medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564 model. How to… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-1000-gpt-3.5-tur_9062874564.text1K<n<10K0 likes16 downloads2y agoHugging Face26celsowm /codigo_florestal_lei_25_05_12textn<1K0 likes16 downloads1y agoHugging Face27florianhoenicke /jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-100-64-16-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes15 downloads2y agoHugging Face28florianhoenicke /jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you can load… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-64-16-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes15 downloads2y agoHugging Face29florianhoenicke /medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 Dataset Dataset Description medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564 model. How to… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/medical-1000-64-16-jinaai_jina-embeddings-v2-small-en-10-gpt-4-turbo-0_9062874564.text1K<n<10K0 likes15 downloads2y agoHugging Face30florianhoenicke /jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 Dataset Dataset Description jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks. Associated Model This dataset was used to train the jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564 model. How to Use To use this dataset for model training or evaluation, you can load it… See the full description on the dataset page: https://huggingface.co/datasets/florianhoenicke/jina-website-1-0-1-BAAI_bge-small-en-v1.5-50_9062874564.textn<1K0 likes14 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.