datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GPT_Watercolor_Anime_Style_Images
GPT Watercolor Anime Style Images
Dataset Description
This is a synthetic GPT-generated Watercolor Anime Style image dataset. It contains 120 image-caption pairs with transparent watercolor washes, soft ink linework, pale paper texture, muted colors, and traditional anime illustration scenes.
The images focus on soft watercolor washes, expressive linework, gentle lighting, quiet interiors, nature scenes, village streets, character studies, and calm storybook… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Watercolor_Anime_Style_Images.Proccessed-GPT-NEOGPT_Storybook_Anime_Style_Images
GPT Storybook Anime Style Images
Dataset Description
This is a synthetic GPT-generated Storybook Anime Style image dataset. It contains 100 image-caption pairs featuring original anime-inspired characters and scenes with a warm, illustrated storybook feeling.
The images focus on expressive character moments, gentle lighting, quiet interiors, nature scenes, village streets, cozy everyday settings, and calm storybook moods. Captions commonly describe soft linework… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Storybook_Anime_Style_Images.GPT_Pointillism_Style_Images
GPT Pointillism Style Images
Dataset description
This is a synthetic GPT-generated pointillism style image dataset. It contains 70 image-caption pairs with colorful dotted texture, painterly lighting, and pointillism-inspired compositions.
The images focus on colorful stippled brushwork, luminous dotted texture, painterly lighting, cozy subjects, gardens, landscapes, interiors, and decorative still-life scenes.
Contents
images/
metadata.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Pointillism_Style_Images.GPT_Photoreal_3D_Anime_Style_Images
GPT Photoreal 3D Anime Style Images
Dataset Description
This is a synthetic GPT-generated Photoreal 3D Anime Style image dataset. It contains 150 image-caption pairs with square, 1024 × 1024 PNG images.
The images cover individual and group portraits, everyday activities, animals, objects, environments, and fantasy scenes. Captions describe anime-inspired forms, photorealistic surface textures, and physically based 3D lighting. Every caption begins with Photoreal… See the full description on the dataset page: https://huggingface.co/datasets/neonforestmist/GPT_Photoreal_3D_Anime_Style_Images.NIAH-gpt-neox-20bbooksum-summary-analysis_gptneox-8192
Dataset Card for "booksum-summary-analysis-8192"
Subset of emozilla/booksum-summary-analysis with only entries that are less than 8,192 tokens under the EleutherAI/gpt-neox-20b tokenizer.
pretokenized__HuggingFaceFW_fineweb-edu__EleutherAI__gpt-neo-125mGPT-NEO-PRE-SEleutherAI__gpt-neox-20b-details
Dataset Card for Evaluation run of EleutherAI/gpt-neox-20b
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neox-20b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neox-20b-details.gptneo-pubmed-abstracts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
This is a dataset consisting of 10000 PubMed abstracts from The Pile (arXiv:2101.00027), along with completions (both human, and LLM-generated), in order to be used to calculate Heaps Law, in the manner described in the preliminary paper, Heaps' Law in GPT-Neo Large Language Model… See the full description on the dataset page: https://huggingface.co/datasets/rachel6603/gptneo-pubmed-abstracts.togethercomputer__GPT-NeoXT-Chat-Base-20B-details
Dataset Card for Evaluation run of togethercomputer/GPT-NeoXT-Chat-Base-20B
Dataset automatically created during the evaluation run of model togethercomputer/GPT-NeoXT-Chat-Base-20B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/togethercomputer__GPT-NeoXT-Chat-Base-20B-details.CSIC_GPTNEO_FT
Dataset Card for "CSIC_GPTNEO_FT"
More Information needed
EleutherAI__gpt-neo-125m-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-125m
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-125m
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-125m-details.EleutherAI__gpt-neo-2.7B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-2.7B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-2.7B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-2.7B-details.dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.0-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.4-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.6-20250109AA_GPTNEO_FT
Dataset Card for "AA_GPTNEO_FT"
More Information needed
PubMed-Mimic-Gpt-Neoheheheh
dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p1.0-20250109EleutherAI__gpt-neo-1.3B-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-1.3B
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-1.3B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-1.3B-details.dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.8-20250109PKDD_GPTNEO_FT
Dataset Card for "PKDD_GPTNEO_FT"
More Information needed
Spirit_GPTNEO_FT
Dataset Card for "Spirit_GPTNEO_FT"
More Information needed
27-11-gptneo125agnews-mia_ag_news_client8quality-pruned-llama-gptneox-8k
Dataset Card for "quality-pruned-llama-gptneox-8k"
More Information needed
imdb_data_for_true_preference_gpt-neoo27-11-gptneo125xsum-mia_xsum_client227-11-gptneo125xsum-mia_xsum_client7
