CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01coref-data /knowref_60k_raw The Knowref 60K Dataset Project: https://github.com/aemami1/KnowRef60k Data source: https://github.com/aemami1/KnowRef60k/tree/28e5385d17967744ccb3bdba45fdd89d9690307d Fields annotation_strength (str): annotator agreement from 1-5 candidate_0 (str): the first candidate name candidate_1 (str): the second candidate name original_sentence (str): sentence before swapping the names swapped_sentence (str): sentence after swapping the names with square brackets marking the… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/knowref_60k_raw.text10K<n<100K0 likes7.9k downloads3y agoHugging Face02coref-data /preco_raw The PreCo Dataset Project: https://preschool-lab.github.io/PreCo/ Data source: https://drive.google.com/file/d/1q0oMt1Ynitsww9GkuhuwNZNq6SjByu-Y/view?usp=sharing Details The original PreCo .jsonl files from https://preschool-lab.github.io/PreCo/ What is PreCo? PreCo is a large-scale English dataset for coreference resolution. The dataset is designed to embody the core challenges in coreference, such as entity representation, by alleviating the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/preco_raw.text10K<n<100K1 likes896 downloads3y agoHugging Face03data-loader /Drivegpt4_raw_dataimage10K<n<100K0 likes584 downloads3mo agoHugging Face04altaidevorg /yargitay-karar-data-raw Yargıtay Karar Data This dataset contains court decisions made by the Turkish Court of Cassation (Yargıtay). Note: This is the raw data --no augmentation, cleaning or deduplication was applied, and the content field is in raw HTML. Note: Scraping continues, and this dataset currently contains the decisions made only in the years between 2021 and 2023 --other years will be uploaded once they are complete. Format It has the following fields: id: The decision id at… See the full description on the dataset page: https://huggingface.co/datasets/altaidevorg/yargitay-karar-data-raw.text1M<n<10M1 likes189 downloads1y agoHugging Face05eternis /eternis_raw_router_datasettabular10K<n<100K0 likes95 downloads1y agoHugging Face06vishnu-vizz /lma_clean_raw_datasettabular1M<n<10M0 likes90 downloads1mo agoHugging Face07BioDEX /raw_datasettext10K<n<100K2 likes63 downloads3y agoHugging Face08Kubermatic /cncf-raw-data-for-llm-training CNCF Raw Data for LLM Training Description This dataset, named cncf-raw-data-for-llm-training, consists of markdown (MD) and PDF content extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. The data was collected by fetching MD and PDF files from different CNCF project repositories and converting them into JSON format. This dataset is intended as raw data for training large language models (LLMs). The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-raw-data-for-llm-training.text10K<n<100K0 likes62 downloads2y agoHugging Face09braindecode /example_dataset-raw EEG Dataset This dataset was created using braindecode, a deep learning library for EEG/MEG/ECoG signals. Dataset Information Property Value Recordings 1 Type Continuous (Raw) Channels 26 Sampling frequency 250 Hz Total duration 0:06:26 Windows/samples 96,735 Size 19.22 MB Format zarr Quick Start from braindecode.datasets import BaseConcatDataset # Load from Hugging Face Hub dataset =… See the full description on the dataset page: https://huggingface.co/datasets/braindecode/example_dataset-raw.tabularn<1K0 likes58 downloads6mo agoHugging Face10Quran-Lab /quranic-asr-cloud-rawdatagated Quranic ASR Provider Benchmark Results Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark. This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio. What Is Included Area Path Purpose Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.tabularautomatic-speech-recognition1K<n<10K1 likes32 downloads27d agoHugging Face11vishnu-vizz /lma_tamil_raw_datasettext1M<n<10M0 likes26 downloads1mo agoHugging Face12data4elm /wikitext-2-raw-v1-test wikitext-2-raw-v1-test Test split for lm-eval-harness. text1K<n<10K0 likes20 downloads11mo agoHugging Face13gryffindor-ISWS /fictional_characters_raw_data_with_imagesimage1K<n<10K1 likes19 downloads3y agoHugging Face14JiazhenLiu01 /raw_data_sampletextn<1K0 likes18 downloads2y agoHugging Face15Snyhlx /raw_merged_data_opencodeinstruct_8_28_trajectory_historytext1M<n<10M0 likes18 downloads1y agoHugging Face16BOP-Berlin-University-Alliance /dc_elements_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology elements. texttext-classificationn<1K0 likes17 downloads3y agoHugging Face17BOP-Berlin-University-Alliance /dc_terms_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology terms. texttext-classificationn<1K0 likes17 downloads3y agoHugging Face18rafmacalaba /datause_raw_extractions datause_raw_extractions Raw World Bank document extractions (one document per line). Each row has two columns: doc_id — the document's metadata.id. doc — a JSON string holding the full record (metadata + model_extractions, where each model_extractions entry is one page with input_text, datasets, classifier_skipped, skip_reason). from datasets import load_dataset import json ds = load_dataset('rafmacalaba/datause_raw_extractions')['train'] record = json.loads(ds[0]['doc']) texttoken-classification1K<n<10K0 likes17 downloads1mo agoHugging Face19gryffindor-ISWS /prompts_wiki_fictional_characters_raw_data_with_imageimage1K<n<10K0 likes15 downloads3y agoHugging Face20gryffindor-ISWS /fictional_characters_raw_data_without_imagestexttext-to-image1K<n<10K0 likes14 downloads3y agoHugging Face21hhhhhhhhhans /medicine-data-rawtext10K<n<100K0 likes14 downloads2y agoHugging Face22RexTRO111 /FINETUNING-RAW-DATASET-V1textn<1K0 likes12 downloads2mo agoHugging Face23gryffindor-ISWS /subset-fictional-characters-raw-data-with-imagesimage1K<n<10K0 likes10 downloads3y agoHugging Face24viber1 /raw-dataset-lawlingotextn<1K0 likes10 downloads2y agoHugging Face25Xhub1880 /ascii_100k_raw_datasettext100K<n<1M0 likes9 downloads7mo agoHugging Face26gryffindor-ISWS /prompts_subset_wiki_fictional_characters_raw_data_with_imageimage1K<n<10K0 likes8 downloads3y agoHugging Face27GENIAC-Team-Ozaki /debate_raw_datatextn<1K0 likes8 downloads2y agoHugging Face28USS-Inferprise /Phi4-Mini-Prose2Tags-4B-Raw-Training-DataRaw data used to train USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B) texttable-to-text100K<n<1M0 likes8 downloads5mo agoHugging Face29data4elm /roleplay-raw-test roleplay-raw-test Test split for lm-eval-harness. text10K<n<100K0 likes5 downloads11mo agoHugging Face30nanat05525 /gsm8k-raw-datatext1K<n<10K0 likes5 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.