datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ioi-eval-openrouter_openai_gpt-3.5-turboioi-eval-dummy-openrouter_openai_gpt-3.5-turboioi-eval-openrouter_openai_gpt-3.5-turbo-new-promptioi-eval-openrouter_openai_gpt-3.5-turbo-textioi-eval-openrouter_openai_gpt-3.5-turbo-prompt-mem-limitQuRating-GPT3.5-Judgments-Test7140 pairwise judgments across 4 criteria and 6 domains obtained by prompting GPT-3.5-turbo-0613 for evaluating QuRater models.
From the paper: QuRating: Selecting High-Quality Data for Training Language Models
Guidance on Responsible Use
In the paper, we document various types of bias that are present in the quality ratings/QuRater model (biases related to domains, topics, social roles, regions and languages - see Section 6 of the paper),
which are likely reflected in the LLM judgments.
Hence… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRating-GPT3.5-Judgments-Test.TinyStories-GPT3.5
Dataset Card for "TinyStories-GPT3.5"
More Information needed
ultrafeedback-gpt-3.5-turbo-helpfulness
UltraFeedback GPT-3.5-Turbo Helpfulness Dataset
Summary
The UltraFeedback GPT-3.5-Turbo Helpfulness dataset contains processed user-assistant interactions filtered for helpfulness, derived from the openbmb/UltraFeedback dataset. It is designed for fine-tuning and evaluating models in alignment tasks.
Data Structure
Format: Conversational
Type: Unpaired preference
Column:
"pompt": The input question or instruction provided to the model.
"completion": The… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness.Wildchat-1M-gpt-4.1-regenerated-english-gpt-3.5-ogQuRating-GPT3.5-Judgments250K thousand pairwise judgments across 4 criteria obtained by prompting GPT-3.5-turbo-0613.
From the paper: QuRating: Selecting High-Quality Data for Training Language Models
Guidance on Responsible Use
In the paper, we document various types of bias that are present in the quality ratings/QuRater model (biases related to domains, topics, social roles, regions and languages - see Section 6 of the paper),
which are likely reflected in the LLM judgments.
Hence, be aware that data selection with… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/QuRating-GPT3.5-Judgments.portuguese-gpt3.5-fine-tuningKA__allenai_openbookqa__openai.gpt-3.5-turbo-0125__5_5_5medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564
medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 Dataset
Dataset Description
medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 is a generated dataset designed to support the development of domain specific embedding models for retrieval tasks.
Associated Model
This dataset was used to train the medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564 model.
How to… See the full description on the dataset page: https://huggingface.co/datasets/florian-hoenicke/medical-20-0-16-jinaai_jina-embeddings-v2-small-en-100-gpt-3.5-turbo-0_9062874564.Ehn-bible-bbc-gpt3.5
Dataset Card for Ehn-Bible-BBC-GPT3.5
Dataset Summary
This dataset card contains parallel Nigerian Pidgin and English sentences split into three files, namely: train.csv, valid.csv and test.csv.
The original data was split in the ratio of 8:1:1 to obtain these files.
Supported Tasks and Leaderboards
Language Translation
Language Identification
Languages
English
Nigerian Pidgin
Dataset Structure
Data Instances
Data… See the full description on the dataset page: https://huggingface.co/datasets/NITHUB-AI/Ehn-bible-bbc-gpt3.5.brockport-gpt-3.5-qagpt3.5-jinyong-style
OpenAI微调医学数据集
开源数据集存在与预训练模型数据集相似度较高的问题,导致微调结果不明显,效果不突出,因此就需要构建某种风格的专属数据集。
因此,构建了带金庸创作风格的医学数据集,可直接用于openai的微调使用。
微调实验结果证明,当输入医学问题时,微调模型回答的风格为金庸创作风格回答。
Question
[{'role': 'system', 'content': '你是我的私人医生助手,你要回答我的健康问题。'}, {'role': 'user', 'content': '我睡眠不足,很头痛'}]
Answer
少侠保重身体。江湖人称“睡眠不足,头痛难受”,此乃常见之症。吾乃医术高强之侠士,愿为你指点迷津。睡眠不足,多因忧虑纷扰、劳累过度所致。当以调整生活作息,保持规律之睡眠时间为要。且可饮用一些花草茶,如菊花、薰衣草,以宁心安神。如症状持续,还需请教名医,以寻得更为妥善之疗法。
可以发现,模型的回答已经学习到金庸创作风格的能力。
instruction-gpt-3.5
Dataset Card for "instruction-gpt-3.5"
More Information needed
MetaMathQA-40K-GPT3.5MetaMathQA-40K adapted to the GPT3.5 dataset format in JSONL for Fine-tuning. Following the following model:
{"messages": [{"role": "system", "content": ""}, {"role": "user", "content": ""}, {"role": "assistant", "content": ""}]}
entity-attribute-dataset-GPT-3.5-generated-v1
Entity Attribute Dataset 306k (GPT-3.5 generated)
Dataset Summary
The Entity Attribute Dataset 306k (GPT-3.5 generated) is designed for instruction fine-tuning, specifically for the task of generating structured catalogs in JSON format based on product titles. The dataset includes a diverse range of products from various categories such as food, home and kitchen, clothing, handicrafts, tools, automotive equipment, and more.
Usage
This dataset is intended for… See the full description on the dataset page: https://huggingface.co/datasets/BaSalam/entity-attribute-dataset-GPT-3.5-generated-v1.math-train-rewrite-gpt-3.5-turbo-0125-t0self_compare_alpacaeval_related_8maxturns_complete_805_gpt3.5_cleanedKA__allenai_ai2_arc__openai.gpt-3.5-turbo-0125__5_5_5Generated_Restaurant_Reviews_GPT3.5
license: cc-by-4.0
task_categories:
text-classification
language:
tr
tags:
food
Generated Review
size_categories:
1K<n<10K
GPT3.5_summarization_preference_RLAIF
Dataset Card for "GPT3.5_summarization_preference_RLAIF"
More Information needed
FOLIO_by_paraphrased_gpt3.5Dynamic-Temperature-GPT-3.5-Turbo
DoCoreAI Dynamic Temperature Test Dataset
This dataset is designed to evaluate the Dynamic Temperature Profiling feature of DoCoreAI. It contains prompts and role-based inputs, along with AI-generated responses and intelligence profiling metrics (reasoning, creativity, precision, and temperature).
🤖 What is DoCoreAI?
DoCoreAI is a first-of-its-kind AI optimization engine that eliminates the need for manual prompt tuning. It automatically profiles your query and adjusts… See the full description on the dataset page: https://huggingface.co/datasets/DoCoreAI/Dynamic-Temperature-GPT-3.5-Turbo.self-talk_gpt3.5_gpt4o_prefpairs_truncated2048_cutto1turnsmeetingbank-gpt3.5rolls-royce-car-description-using-gpt-3.5orca_DPO_pairs_gpt3.5
