CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ichsan2895 /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K23 likes874 downloads3y agoHugging Face02katielink /gpt4_bias Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare This repository accompanies the paper "Coding Inequity: Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare". Overview The data is available in the data_to_share folder. This can be broken into several pieces: simulated_pt_distribution --- here is where we store all the information for generating patient demographic distributions. We store the outputs of… See the full description on the dataset page: https://huggingface.co/datasets/katielink/gpt4_bias.tabularn<1K1 likes134 downloads3y agoHugging Face03joyfine /TruthfulQA_CoT_GPT4textn<1K4 likes103 downloads3y agoHugging Face04kartoun /Alcohol_Use_Clinical_Notes_GPT4Contributions: The dataset was created by Dr. Uri Kartoun. Use Case: Leveraging Large Language Models for Enhanced Clinical Narrative Analysis: An Application in Alcohol Use Detection Dataset Summary: This dataset contains 1,500 samples of expressions indicating alcohol use or its negation, generated from clinical narrative notes using OpenAI's ChatGPT 4 model. It's designed to support NLP applications that require the identification of alcohol use references in healthcare records. Text… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Alcohol_Use_Clinical_Notes_GPT4.texttext-classification1K<n<10K0 likes76 downloads1y agoHugging Face05tuanio /LaVy-Bench-GPT4o LaVy-Bench (with Answers 😎) Welcome to the LaVy-Bench dataset repository! About We offers manually generated answers created using GPT-4, providing meaningful, detailed, and bug-free responses. Our goal is to contribute to LaVy-Bench as a significant benchmark for Vietnamese Multi-Modal and Vietnamese Large Vision Language Models in real-world scenarios. Contribution We aim to generate meaningful answers for questions-only datasets sourced from the original… See the full description on the dataset page: https://huggingface.co/datasets/tuanio/LaVy-Bench-GPT4o.imagevisual-question-answeringn<1K1 likes58 downloads2y agoHugging Face06kartoun /Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun. Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files. Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.text1K<n<10K0 likes54 downloads1y agoHugging Face07Faishal-Anwar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes52 downloads19d agoHugging Face08tomasonjo /text2cypher-gpt4o-clean Synthetic dataset created with GPT-4o Synthetic dataset of text2cypher over 16 different graph schemas. Questions were generated using GPT-4-turbo, and the corresponding Cypher statements with gpt-4o using Chain of Thought. Here, there are only questions that return results when queried against the database. For more information visit: https://github.com/neo4j-labs/text2cypher/tree/main/datasets/synthetic_gpt4o_demodbs Dataset is available as train.csv. Columns are the following:… See the full description on the dataset page: https://huggingface.co/datasets/tomasonjo/text2cypher-gpt4o-clean.text1K<n<10K20 likes42 downloads2y agoHugging Face09OfirArviv /mt_bench_single_score_gpt4_judgementtabular1K<n<10K1 likes41 downloads2y agoHugging Face10glhpradipta /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes40 downloads13d agoHugging Face11businessrules /gpt4.1_promptA_results_aggregatedtabularn<1K0 likes37 downloads27d agoHugging Face12halilbabacan /cognitive_distortions_gpt4texttext-classification1K<n<10K1 likes35 downloads2y agoHugging Face13Chemin-AI /indonlu-eval-gpt4o-vs-sealionv3-round1 Local vs Global: Testing GPT-4o-mini and SEA-LIONv3 on Bahasa Indonesia A benchmark dataset comparing GPT-4o-mini and SEA-LIONv3 on 50 Indonesian-specific questions.This is Round 1 of the INDONLU Eval series, which was built to test LLM performance on culturally grounded, linguistically diverse Southeast Asian prompts. Overview We tested 50 prompts across four core categories to assess how well large language models can handle local Indonesian context: Language –… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/indonlu-eval-gpt4o-vs-sealionv3-round1.texttranslationn<1K8 likes26 downloads1y agoHugging Face14OfirArviv /mt_bench_pairwise_comparison_gpt4_judgmentstabular1K<n<10K0 likes25 downloads2y agoHugging Face15pbevan11 /synthetic-ocr-correction-gpt4o Synthetic OCR Correction GPT-4o 10,000 pieces of news text from fancyzhx/ag_news with synthetically generated OCR mistakes. The purpose of this is to mimic corrupt text that has been transcribed with OCR from old newspapers, where there are often lot's of errors. See biglam/bnl_newspapers1841-1879 for example. By synthetically creating it, we have the true ground truth, meaning we can use this as a source of truth for finetuning. The corrupted text was generated using OpenAI's… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/synthetic-ocr-correction-gpt4o.text10K<n<100K6 likes23 downloads2y agoHugging Face16surogate /alpaca-gpt4-data-zh 数据集描述 该数据集为GPT-4生成的中文数据集,用于LLM的指令精调和强化学习等。 数据集加载方式 from modelscope.msdatasets import MsDataset ds = MsDataset.load("alpaca-gpt4-data-zh", namespace="AI-ModelScope", split="train") print(next(iter(ds))) 数据分片 数据已经预设了train分片。 数据集版权信息 数据集已经开源,license为CC BY NC 4.0(仅用于非商业化用途),如有违反相关条款,随时联系modelscope删除。 引用方式 @article{peng2023gpt4llm, title={Instruction Tuning with GPT-4}, author={Baolin Peng, Chunyuan Li, Pengcheng He, Michel… See the full description on the dataset page: https://huggingface.co/datasets/surogate/alpaca-gpt4-data-zh.texttext-generation10K<n<100K0 likes20 downloads1y agoHugging Face17AnnikaSimonsen /GPT-4_FO-EN_parallel_blog_sentences_MQMThis is dataset contains 425 Faroese-to-English parallel sentences generated by GPT-4 that have been annotated by a single native speaker of Faroese using the Multidimensional Quality Metrics framework (MQM). The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus. In addition to the parallel sentences and human evaluation, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences_MQM.textn<1K0 likes19 downloads2y agoHugging Face18surogate /alpaca-gpt4-data-en license: apache-2.0 数据集描述 该数据集为GPT-4生成的英文数据集,用于LLM的指令精调和强化学习等。 数据集加载方式 from modelscope.msdatasets import MsDataset ds = MsDataset.load("alpaca-gpt4-data-en", namespace="AI-ModelScope", split="train") print(next(iter(ds))) 数据分片 数据已经预设了train分片。 Clone with HTTP git clone https://www.modelscope.cn/datasets/AI-ModelScope/alpaca-gpt4-data-en.git 数据集版权信息 数据集已经开源,license为CC BY NC 4.0(仅用于非商业化用途),如有违反相关条款,随时联系modelscope删除。… See the full description on the dataset page: https://huggingface.co/datasets/surogate/alpaca-gpt4-data-en.text10K<n<100K0 likes19 downloads1y agoHugging Face19AashishKumar /chatml-hinglish-conversation-gpt4text10K<n<100K3 likes18 downloads2y agoHugging Face20sudiptabasak /alpaca-gpt4-csvtext10K<n<100K0 likes17 downloads3y agoHugging Face21AnnikaSimonsen /GPT-4_FO-EN_parallel_blog_sentencesThis is dataset contains 1,673 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus. In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article. Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences.text1K<n<10K0 likes17 downloads2y agoHugging Face22AnnikaSimonsen /GPT-4_FO-EN_parallel_news_sentencesThis is dataset contains 3,735 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is news text from the Basic Language Resource Kit for Faroese 1.0 text corpus. In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article. Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_news_sentences.text1K<n<10K0 likes17 downloads2y agoHugging Face23barbaroo /FLORES200_translations_GPT4 Dataset Summary This dataset consists of three synthetic parallel English-to-Faroese translations of 1,012 sentences from the FLORES-200 benchmark. The translations were generated using GPT-4 Turbo (gpt-4-1106-preview) with three different prompting strategies: Zero-shot translation (no additional examples provided). Random few-shot translation (12 few-shot examples selected randomly). STS-based few-shot translation (12 few-shot examples selected using Semantic Textual Similarity).… See the full description on the dataset page: https://huggingface.co/datasets/barbaroo/FLORES200_translations_GPT4.texttranslation1K<n<10K0 likes17 downloads2y agoHugging Face24pbevan11 /GPT4V-captions-from-LVIS-typography GPT4V-captions-from-LVIS-typography by: Peter Bevan, 21 March 2023 This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS. This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.image1K<n<10K1 likes15 downloads3y agoHugging Face25vitus9988 /ko_gpt4omini_note_15.4k 한국어 메모 데이터셋 GPT-4o-mini를 통해 생성된 한국어 메모 데이터셋입니다. 대주제(main_topic), 소주제(sub_topic)를 통해 메모처럼 보이는 데이터를 생성하였습니다. text10K<n<100K1 likes15 downloads2y agoHugging Face26Politics /turkey-chp-gpt4oHandlabeled dataset to finetune GPT-4o model to classify whether a sentence depicts CHP as "villain" according to the below codebook: CHP CHP as the villain Illegitimacy Component For the illegitimacy component of polarizing rhetoric, we focus on sentences which portray the main opposition party, CHP, as an illegitimate political actor. More specifically, sentences containing similar claims to categories below, are coded as instances where CHP is shown as illegitimate.… See the full description on the dataset page: https://huggingface.co/datasets/Politics/turkey-chp-gpt4o.textn<1K0 likes14 downloads10mo agoHugging Face27naimulislam /GPT4-Chat-GRP0text10K<n<100K1 likes13 downloads2y agoHugging Face28LogoPogoxXx /Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun. Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files. Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/LogoPogoxXx/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.text1K<n<10K0 likes13 downloads6mo agoHugging Face29AmalAkbar /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/AmalAkbar/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K0 likes13 downloads2mo agoHugging Face30Andrei481 /alpaca-gpt4-rotext10K<n<100K0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.