datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.gpt4_bias
Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare
This repository accompanies the paper "Coding Inequity: Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare".
Overview
The data is available in the data_to_share folder. This can be broken into several pieces:
simulated_pt_distribution --- here is where we store all the information for generating patient demographic distributions. We store the outputs of… See the full description on the dataset page: https://huggingface.co/datasets/katielink/gpt4_bias.TruthfulQA_CoT_GPT4Alcohol_Use_Clinical_Notes_GPT4Contributions: The dataset was created by Dr. Uri Kartoun.
Use Case: Leveraging Large Language Models for Enhanced Clinical Narrative Analysis: An Application in Alcohol Use Detection
Dataset Summary: This dataset contains 1,500 samples of expressions indicating alcohol use or its negation, generated from clinical narrative notes using OpenAI's ChatGPT 4 model. It's designed to support NLP applications that require the identification of alcohol use references in healthcare records.
Text… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Alcohol_Use_Clinical_Notes_GPT4.LaVy-Bench-GPT4o
LaVy-Bench (with Answers 😎)
Welcome to the LaVy-Bench dataset repository!
About
We offers manually generated answers created using GPT-4, providing meaningful, detailed, and bug-free responses. Our goal is to contribute to LaVy-Bench as a significant benchmark for Vietnamese Multi-Modal and Vietnamese Large Vision Language Models in real-world scenarios.
Contribution
We aim to generate meaningful answers for questions-only datasets sourced from the original… See the full description on the dataset page: https://huggingface.co/datasets/tuanio/LaVy-Bench-GPT4o.Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0Acknowledgment: The dataset was created by Dr. Uri Kartoun.
Description: The dataset was designed for the classification of text descriptions into seven stages of pancreatic cancer. It comprises two sets: a training set and a held-out set. Each set contains 700 blobs of text, with each blob representing a specific stage of pancreatic cancer. There are 100 text blobs for each of the seven defined stages in both files.
Data Collection and Preparation: The text blobs were generated using… See the full description on the dataset page: https://huggingface.co/datasets/kartoun/Pancriatic_cancer_stages_clinical_narrative_blobs_and_labels_gpt4_v0.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.mt_bench_single_score_gpt4_judgementtext2cypher-gpt4o-clean
Synthetic dataset created with GPT-4o
Synthetic dataset of text2cypher over 16 different graph schemas.
Questions were generated using GPT-4-turbo, and the corresponding Cypher statements with gpt-4o using Chain of Thought.
Here, there are only questions that return results when queried against the database.
For more information visit: https://github.com/neo4j-labs/text2cypher/tree/main/datasets/synthetic_gpt4o_demodbs
Dataset is available as train.csv. Columns are the following:… See the full description on the dataset page: https://huggingface.co/datasets/tomasonjo/text2cypher-gpt4o-clean.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.gpt4.1_promptA_results_aggregatedcognitive_distortions_gpt4mt_bench_pairwise_comparison_gpt4_judgmentsindonlu-eval-gpt4o-vs-sealionv3-round1
Local vs Global: Testing GPT-4o-mini and SEA-LIONv3 on Bahasa Indonesia
A benchmark dataset comparing GPT-4o-mini and SEA-LIONv3 on 50 Indonesian-specific questions.This is Round 1 of the INDONLU Eval series, which was built to test LLM performance on culturally grounded, linguistically diverse Southeast Asian prompts.
Overview
We tested 50 prompts across four core categories to assess how well large language models can handle local Indonesian context:
Language –… See the full description on the dataset page: https://huggingface.co/datasets/Chemin-AI/indonlu-eval-gpt4o-vs-sealionv3-round1.GPT-4_FO-EN_parallel_blog_sentences_MQMThis is dataset contains 425 Faroese-to-English parallel sentences generated by GPT-4 that have been annotated by a single native speaker of Faroese using the Multidimensional Quality Metrics framework (MQM). The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences and human evaluation, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences_MQM.FLORES200_translations_GPT4
Dataset Summary
This dataset consists of three synthetic parallel English-to-Faroese translations of 1,012 sentences from the FLORES-200 benchmark. The translations were generated using GPT-4 Turbo (gpt-4-1106-preview) with three different prompting strategies:
Zero-shot translation (no additional examples provided).
Random few-shot translation (12 few-shot examples selected randomly).
STS-based few-shot translation (12 few-shot examples selected using Semantic Textual Similarity).… See the full description on the dataset page: https://huggingface.co/datasets/barbaroo/FLORES200_translations_GPT4.alpaca-gpt4-data-zh
数据集描述
该数据集为GPT-4生成的中文数据集,用于LLM的指令精调和强化学习等。
数据集加载方式
from modelscope.msdatasets import MsDataset
ds = MsDataset.load("alpaca-gpt4-data-zh", namespace="AI-ModelScope", split="train")
print(next(iter(ds)))
数据分片
数据已经预设了train分片。
数据集版权信息
数据集已经开源,license为CC BY NC 4.0(仅用于非商业化用途),如有违反相关条款,随时联系modelscope删除。
引用方式
@article{peng2023gpt4llm,
title={Instruction Tuning with GPT-4},
author={Baolin Peng, Chunyuan Li, Pengcheng He, Michel… See the full description on the dataset page: https://huggingface.co/datasets/surogate/alpaca-gpt4-data-zh.alpaca-gpt4-data-en
license: apache-2.0
数据集描述
该数据集为GPT-4生成的英文数据集,用于LLM的指令精调和强化学习等。
数据集加载方式
from modelscope.msdatasets import MsDataset
ds = MsDataset.load("alpaca-gpt4-data-en", namespace="AI-ModelScope", split="train")
print(next(iter(ds)))
数据分片
数据已经预设了train分片。
Clone with HTTP
git clone https://www.modelscope.cn/datasets/AI-ModelScope/alpaca-gpt4-data-en.git
数据集版权信息
数据集已经开源,license为CC BY NC 4.0(仅用于非商业化用途),如有违反相关条款,随时联系modelscope删除。… See the full description on the dataset page: https://huggingface.co/datasets/surogate/alpaca-gpt4-data-en.alpaca-gpt4-roko_gpt4omini_note_15.4k
한국어 메모 데이터셋
GPT-4o-mini를 통해 생성된 한국어 메모 데이터셋입니다.
대주제(main_topic), 소주제(sub_topic)를 통해 메모처럼 보이는 데이터를 생성하였습니다.
synthetic-ocr-correction-gpt4o
Synthetic OCR Correction GPT-4o
10,000 pieces of news text from fancyzhx/ag_news with synthetically generated OCR mistakes.
The purpose of this is to mimic corrupt text that has been transcribed with OCR from old newspapers, where there are often lot's of errors. See biglam/bnl_newspapers1841-1879 for example. By synthetically creating it, we have the true ground truth, meaning we can use this as a source of truth for finetuning.
The corrupted text was generated using OpenAI's… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/synthetic-ocr-correction-gpt4o.chatml-hinglish-conversation-gpt4alpaca-gpt4-csvGPT-4_FO-EN_parallel_news_sentencesThis is dataset contains 3,735 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is news text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article.
Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_news_sentences.GPT-4_FO-EN_parallel_blog_sentencesThis is dataset contains 1,673 Faroese-to-English parallel sentences generated by GPT-4. The Faroese text is blog text from the Basic Language Resource Kit for Faroese 1.0 text corpus.
In addition to the parallel sentences, the dataset contains a column with a quality report made by GPT-4 in which it describes the challenges it faced when translating the each article.
Please be aware, that according to OpenAI's the terms of use, then it is not allowed to use their output to create models that… See the full description on the dataset page: https://huggingface.co/datasets/AnnikaSimonsen/GPT-4_FO-EN_parallel_blog_sentences.GPT4V-captions-from-LVIS-typography
GPT4V-captions-from-LVIS-typography
by: Peter Bevan, 21 March 2023
This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS.
This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption.
The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/akahana/alpaca-gpt4-indonesian.Finance_COT_GPT4gpt4-pol-ideologies-small
Dataset Card for Dataset Name
This dataset card contains very short paragraphs (2-3 sentences) which are labelled as either 'liberal' or 'conservative'. It has been generated using GPT-4.
Dataset Details
Dataset Description
The dataset has been created using this prompt: 'Write a short paragraph (2-3 sentences) expressing a {label} political viewpoint.'
All the entries has also been manually checked to ensure that the paragraph accurately maps to the… See the full description on the dataset page: https://huggingface.co/datasets/JyotiNayak/gpt4-pol-ideologies-small.GPT4-Chat-GRP0
