datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-codetaste-can-llms-generate-human-level-code-refactorings-traces
Agent traces
Agent sessions published from a Trackio Logbook.
openrubrics-v2-generated-rubrics-gpt-54-complete-fixed-judgementsdata2plot_generatedgpt2-large-generated
Synthetic Text Corpus - GPT-2 Large
Dataset Description
This dataset contains synthetically generated text sequences sampled from GPT-2 Large. It was created to provide a large-scale text corpus for research in natural language processing, particularly for studies on model behavior, text generation, and language modeling.
Dataset Summary
Size: ~100M tokens
Number of sequences: ~500,000
Source model: gpt2-large (774M parameters)
Sequence length: Maximum 256… See the full description on the dataset page: https://huggingface.co/datasets/vesteinn/gpt2-large-generated.croco-munin-apertus-8b-da-generatedcheck-kltn-generated-cielr
