datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-4o-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
metal-python-synthetic-explanations-gpt4-graphcodebertOpenThoughts3-456k-gpt4.1-cotgpt4_judge_battleshle-context-baseline-gpt41comicstrips-gpt4o-blip3
Comic Strips
Dataset Details
Dataset Description
This dataset contains indie comics from Reddit, then captioned with GPT4o and BLIP3.
Currently, only the GPT4o captions are available in this repository. The BLIP3 captions will be uploaded soon.
Roughly 1400 images were captioned at a cost of ~$11 using GPT4o (25 May 2024 version).
Curated by: @pseudoterminalx
Funded by @pseudoterminalx
License: MIT
Dataset Sources
Unlike other free-to-use… See the full description on the dataset page: https://huggingface.co/datasets/bghira/comicstrips-gpt4o-blip3.alpaca-gpt4-trgpt4_bias
Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare
This repository accompanies the paper "Coding Inequity: Assessing GPT-4’s Potential for Perpetuating Racial and Gender Biases in Healthcare".
Overview
The data is available in the data_to_share folder. This can be broken into several pieces:
simulated_pt_distribution --- here is where we store all the information for generating patient demographic distributions. We store the outputs of… See the full description on the dataset page: https://huggingface.co/datasets/katielink/gpt4_bias.MMInstruct-GPT4V_mistral-7b_l0_cutMMInstruct-GPT4V_mistral-7b_cosi_cutMMInstruct-GPT4V_mistral-7b_cooccur_cuttinystories-gpt4-instruct
tinystories-gpt4-instruct
Request→story pairs for supervised fine-tuning of small language models, derived from karpathy/tinystories-gpt4-clean. Each example pairs a natural-language request ("Can you tell me a story about a boy named Tim?") with a TinyStories story that satisfies it.
The dataset lives on Hugging Face; the notebook that generates it lives on GitHub.
This is not roneneldan/TinyStoriesInstruct. That dataset frames its tasks in a structured format (Words:… See the full description on the dataset page: https://huggingface.co/datasets/Pondsiders/tinystories-gpt4-instruct.metal-python-synthetic-explanations-gpt4-raw
Dataset Card for "metal-python-synthetic-explanations-gpt4-raw"
More Information needed
gpt-4o-function-calling-traces
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/gpt-4o-function-calling-traces.metal-python-synthetic-explanations-gpt4
Dataset Card for "metal-python-synthetic-explanations-gpt4"
More Information needed
gpt4o_swe_resultsalpaca_farm_gpt4dpo-yolo1-200k-gpt4.1-judge-2weak2strong-maxdelta_rejected-DECON-remove-gemma3hle-ctx-screen-gpt41gemma7b-summarize-eval-by-gpt4omt_bench_single_score_gpt4_judgementgpt4o-coding-eval-by-gemini1_5flash-koTranslated llama-duo/gpt4o-coding-eval-by-gemini1_5flash using nayohan/llama3-instrucTrans-enko-8b.
This dataset is a raw translated dataset and contains repetitive sentences generated by the model, so it needs to be filtered.
multilingual_translation_gpt4o_genmmlu-pro-ctx-screen-gpt41-10gpt_4o_mini_classifications_multi_humangpt4.1_promptA_results_aggregatedgpt4o-classification-eval-by-claude3sonnetgpt4o-classification-eval-by-gemini1_5flashapollo_english_guidelines_translated_to_dutch_with_gpt4omini
Data description
Translation of the English medical guidelines that are part of the Apollo corpus, using the LLM GPT 4o mini
Acknowledgement
The work received funding from the European Union's Horizon Europe research
and innovation programme under Grant Agreement No. 101057849 (DataTools4Heart project).
For more information on the background, see Datatools4Heart Huggingface/Website/Git
gpt4o-summarize-eval-by-gemini1_5flash
