CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLMDH /post-ocr2text100K<n<1M6 likes21k downloads1y agoHugging Face02Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes6.8k downloads3y agoHugging Face03LLMDH /other2document100K<n<1M0 likes4.4k downloads1y agoHugging Face04LLMDH /OpenScience Open Science Dataset Overview Open Science is a large-scale, permissively licensed text dataset derived from OpenAlex, containing over 100B (105,390,332,599) words. OpenAlex is an open database of scholarly publications, authors, institutions, and research outputs that serves as a comprehensive source for academic literature. Key Features Truly Open: Contains only permissively licensed data suitable for both commercial and non-commercial use Multilingual… See the full description on the dataset page: https://huggingface.co/datasets/LLMDH/OpenScience.tabular1M<n<10M0 likes4.3k downloads2y agoHugging Face05LLMDH /marianne_pdf_7text10K<n<100K0 likes3.1k downloads2y agoHugging Face06LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.7k downloads6mo agoHugging Face07LLMDH /otherdocument100K<n<1M0 likes2.7k downloads1y agoHugging Face08LLMDH /marianne_pdf_9text100K<n<1M0 likes2.6k downloads2y agoHugging Face09LLMDH /marianne_pdf_5text100K<n<1M0 likes2.2k downloads2y agoHugging Face10LLMDH /marianne_pdf_3text100K<n<1M0 likes2.2k downloads2y agoHugging Face11LLMDH /marianne_pdf_60 likes1.7k downloads2y agoHugging Face12LLMDH /marianne_pdf_8text100K<n<1M0 likes1.6k downloads2y agoHugging Face13LLMDH /marianne_pdf_4text10K<n<100K0 likes1.6k downloads2y agoHugging Face14LLMDH /marianne_pdf_20 likes996 downloads2y agoHugging Face15tong0 /LLM_Dystext1M<n<10M0 likes939 downloads1y agoHugging Face16LLMDH /marianne_pdf_10text100K<n<1M0 likes636 downloads2y agoHugging Face17Yinxing /LLM_Dataset0 likes549 downloads2mo agoHugging Face18LLMDH /marianne_pdf_10 likes448 downloads2y agoHugging Face19LLMDH /hal_pdf_extratext10K<n<100K0 likes426 downloads2y agoHugging Face20LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes410 downloads8mo agoHugging Face21LLMDH /big-code-2-metadataA subset of big-code-2 containing only the files: Not in big-code-1. Under a permissive license. text100M<n<1B0 likes382 downloads2y agoHugging Face22moofeez /llm-debugger-sft-corpus llm-debugger SFT corpus Training input for llm-debugger, a model that works a failing Python test in a live pdb session and edits the fix. This is the corpus behind the SFT checkpoint the best RL policy (v90) was trained from. Contents path what corpus/train_sft.jsonl 312 training rows, native tool-call format corpus/val_sft.jsonl 35 validation rows corpus/build_meta.json row counts and SHA-256 per split, counted at publish corpus/rows.jsonl the… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-sft-corpus.text-generation1K<n<10K0 likes335 downloads16d agoHugging Face23pratyushmaini /llm_dataset_inference LLM Dataset Inference This repository contains various subsets of the PILE dataset, divided into train and validation sets. The data is used to facilitate privacy research in language models, where perturbed data can be used as a reference to detect the presence of a particular dataset in the training data of a language model. Data Used The data is in the form of JSONL files, with each entry containing the raw text, as well as various kinds of perturbations applied to it.… See the full description on the dataset page: https://huggingface.co/datasets/pratyushmaini/llm_dataset_inference.text10K<n<100K1 likes207 downloads2y agoHugging Face24open-llm-leaderboard-old /details_decruz07__kellemar-DPO-7B-d Dataset Card for Evaluation run of decruz07/kellemar-DPO-7B-d Dataset automatically created during the evaluation run of model decruz07/kellemar-DPO-7B-d on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_decruz07__kellemar-DPO-7B-d.0 likes174 downloads3y agoHugging Face25LLMDH /hal_pdf_extra_old0 likes170 downloads2y agoHugging Face26LLMDH /openalex_extractiontext10M<n<100M1 likes152 downloads2y agoHugging Face27kenhktsui /llm-data-textbook-quality-v2texttext-classification1M<n<10M0 likes137 downloads2y agoHugging Face28arincon /llm-detect Dataset Card for "llm-detect" More Information needed tabular100K<n<1M0 likes115 downloads3y agoHugging Face29open-llm-leaderboard-old /details_juhwanlee__llmdo-Mistral-7B-case-5 Dataset Card for Evaluation run of juhwanlee/llmdo-Mistral-7B-case-5 Dataset automatically created during the evaluation run of model juhwanlee/llmdo-Mistral-7B-case-5 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_juhwanlee__llmdo-Mistral-7B-case-5.0 likes114 downloads3y agoHugging Face30open-llm-leaderboard-old /details_juhwanlee__llmdo-Mistral-7B-case-1 Dataset Card for Evaluation run of juhwanlee/llmdo-Mistral-7B-case-1 Dataset automatically created during the evaluation run of model juhwanlee/llmdo-Mistral-7B-case-1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_juhwanlee__llmdo-Mistral-7B-case-1.0 likes113 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.