datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
knowref_60k_raw
The Knowref 60K Dataset
Project: https://github.com/aemami1/KnowRef60k
Data source: https://github.com/aemami1/KnowRef60k/tree/28e5385d17967744ccb3bdba45fdd89d9690307d
Fields
annotation_strength (str): annotator agreement from 1-5
candidate_0 (str): the first candidate name
candidate_1 (str): the second candidate name
original_sentence (str): sentence before swapping the names
swapped_sentence (str): sentence after swapping the names with square brackets marking the… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/knowref_60k_raw.preco_raw
The PreCo Dataset
Project: https://preschool-lab.github.io/PreCo/
Data source: https://drive.google.com/file/d/1q0oMt1Ynitsww9GkuhuwNZNq6SjByu-Y/view?usp=sharing
Details
The original PreCo .jsonl files from https://preschool-lab.github.io/PreCo/
What is PreCo?
PreCo is a large-scale English dataset for coreference resolution. The dataset is designed to embody the core challenges in coreference, such as entity representation, by alleviating the challenge of… See the full description on the dataset page: https://huggingface.co/datasets/coref-data/preco_raw.Drivegpt4_raw_datayargitay-karar-data-raw
Yargıtay Karar Data
This dataset contains court decisions made by the Turkish Court of Cassation (Yargıtay).
Note: This is the raw data --no augmentation, cleaning or deduplication was applied, and the content field is in raw HTML.
Note: Scraping continues, and this dataset currently contains the decisions made only in the years between 2021 and 2023 --other years will be uploaded once they are complete.
Format
It has the following fields:
id: The decision id at… See the full description on the dataset page: https://huggingface.co/datasets/altaidevorg/yargitay-karar-data-raw.eternis_raw_router_datasetlma_clean_raw_datasetraw_datasetcncf-raw-data-for-llm-training
CNCF Raw Data for LLM Training
Description
This dataset, named cncf-raw-data-for-llm-training, consists of markdown (MD) and PDF content extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. The data was collected by fetching MD and PDF files from different CNCF project repositories and converting them into JSON format. This dataset is intended as raw data for training large language models (LLMs).
The dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-raw-data-for-llm-training.example_dataset-raw
EEG Dataset
This dataset was created using braindecode, a deep
learning library for EEG/MEG/ECoG signals.
Dataset Information
Property
Value
Recordings
1
Type
Continuous (Raw)
Channels
26
Sampling frequency
250 Hz
Total duration
0:06:26
Windows/samples
96,735
Size
19.22 MB
Format
zarr
Quick Start
from braindecode.datasets import BaseConcatDataset
# Load from Hugging Face Hub
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/braindecode/example_dataset-raw.quranic-asr-cloud-rawdata
Quranic ASR Provider Benchmark Results
Professional benchmark artifacts for comparing commercial and official ASR providers on the Quranic ASR benchmark hosted at Quran-Lab/quranic-asr-benchmark.
This repository contains metadata, normalized result tables, raw provider responses, unchanged run scripts, scoring outputs, Tarteel streaming probes, and reports. It does not duplicate the source audio.
What Is Included
Area
Path
Purpose
Benchmark split… See the full description on the dataset page: https://huggingface.co/datasets/Quran-Lab/quranic-asr-cloud-rawdata.lma_tamil_raw_datasetwikitext-2-raw-v1-test
wikitext-2-raw-v1-test
Test split for lm-eval-harness.
fictional_characters_raw_data_with_imagesraw_data_sampleraw_merged_data_opencodeinstruct_8_28_trajectory_historydc_elements_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology elements.
dc_terms_raw_dataThe dataset consists of the descriptions and comments about the concepts in Dublin Core ontology terms.
datause_raw_extractions
datause_raw_extractions
Raw World Bank document extractions (one document per line).
Each row has two columns:
doc_id — the document's metadata.id.
doc — a JSON string holding the full record
(metadata + model_extractions, where each model_extractions entry is
one page with input_text, datasets, classifier_skipped, skip_reason).
from datasets import load_dataset
import json
ds = load_dataset('rafmacalaba/datause_raw_extractions')['train']
record = json.loads(ds[0]['doc'])
prompts_wiki_fictional_characters_raw_data_with_imagefictional_characters_raw_data_without_imagesmedicine-data-rawFINETUNING-RAW-DATASET-V1subset-fictional-characters-raw-data-with-imagesraw-dataset-lawlingoascii_100k_raw_datasetprompts_subset_wiki_fictional_characters_raw_data_with_imagedebate_raw_dataPhi4-Mini-Prose2Tags-4B-Raw-Training-DataRaw data used to train USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B)
roleplay-raw-test
roleplay-raw-test
Test split for lm-eval-harness.
gsm8k-raw-data
