datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
post-ocr2llm_datasetsother2OpenScience
Open Science Dataset
Overview
Open Science is a large-scale, permissively licensed text dataset derived from OpenAlex, containing over 100B (105,390,332,599) words. OpenAlex is an open database of scholarly publications, authors, institutions, and research outputs that serves as a comprehensive source for academic literature.
Key Features
Truly Open: Contains only permissively licensed data suitable for both commercial and non-commercial use
Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/LLMDH/OpenScience.marianne_pdf_7Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.othermarianne_pdf_9marianne_pdf_5marianne_pdf_3marianne_pdf_6marianne_pdf_8marianne_pdf_4marianne_pdf_2LLM_Dysmarianne_pdf_10LLM_Datasetmarianne_pdf_1hal_pdf_extraTwin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.big-code-2-metadataA subset of big-code-2 containing only the files:
Not in big-code-1.
Under a permissive license.
llm-debugger-sft-corpus
llm-debugger SFT corpus
Training input for llm-debugger,
a model that works a failing Python test in a live pdb session and edits the
fix. This is the corpus behind the SFT checkpoint the best RL policy (v90) was
trained from.
Contents
path
what
corpus/train_sft.jsonl
312 training rows, native tool-call format
corpus/val_sft.jsonl
35 validation rows
corpus/build_meta.json
row counts and SHA-256 per split, counted at publish
corpus/rows.jsonl
the… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-sft-corpus.llm_dataset_inference
LLM Dataset Inference
This repository contains various subsets of the PILE dataset, divided into train and validation sets. The data is used to facilitate privacy research in language models, where perturbed data can be used as a reference to detect the presence of a particular dataset in the training data of a language model.
Data Used
The data is in the form of JSONL files, with each entry containing the raw text, as well as various kinds of perturbations applied to it.… See the full description on the dataset page: https://huggingface.co/datasets/pratyushmaini/llm_dataset_inference.details_decruz07__kellemar-DPO-7B-d
Dataset Card for Evaluation run of decruz07/kellemar-DPO-7B-d
Dataset automatically created during the evaluation run of model decruz07/kellemar-DPO-7B-d on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_decruz07__kellemar-DPO-7B-d.hal_pdf_extra_oldopenalex_extractionllm-data-textbook-quality-v2llm-detect
Dataset Card for "llm-detect"
More Information needed
details_juhwanlee__llmdo-Mistral-7B-case-5
Dataset Card for Evaluation run of juhwanlee/llmdo-Mistral-7B-case-5
Dataset automatically created during the evaluation run of model juhwanlee/llmdo-Mistral-7B-case-5 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_juhwanlee__llmdo-Mistral-7B-case-5.details_juhwanlee__llmdo-Mistral-7B-case-1
Dataset Card for Evaluation run of juhwanlee/llmdo-Mistral-7B-case-1
Dataset automatically created during the evaluation run of model juhwanlee/llmdo-Mistral-7B-case-1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_juhwanlee__llmdo-Mistral-7B-case-1.
