datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
air-bench-2024
AIRBench 2024
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning categories of the regulation-based safety categories in the
AIR 2024 safety taxonomy.
Dataset Details
Dataset Description
AIRBench 2024 is a AI safety benchmark that aligns with emerging government
regulations and company policies. It consists of diverse, malicious prompts
spanning… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/air-bench-2024.image2struct-latex-v1
Image2Struct - Latex
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images.
This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt:
Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.aic2026-videos-640-crf30helm-scenarios
HELM Scenarios
This repository contains mirrors of datasets that are used as scenarios by crfm-helm.
Scenarios
TURL Column Type Annotation
The subfolder turl-column-type-annotation contains files for the table column type annotation task from the TURL paper. No modifications were made to these files.
The TURL dataset by Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu is licensed under CC BY 4.0. The TURL dataset was modified from the TabEL… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/helm-scenarios.DSIR-filtered-pile-50M
Dataset Card for DSIR-filtered-pile-50M
Dataset Summary
This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (51.2M examples) in jsonl format.
Data Instances
{"contents": "Hundreds of soul music enthusiasts from the United Kingdom plan to make their way to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-50M.DSIR-filtered-pile-100M-short
Dataset Card for DSIR-filtered-pile-100M-short
Dataset Summary
This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.heuristic_classification-filtered-pile-50M
Dataset Card for heuristic_classification-filtered-pile-50M
Dataset Summary
This dataset is a subset of The Pile, selected via the heuristic classification data selection method. The target distribution for heuristic classification are the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (51.2M examples) in jsonl format.
Data Instances
{"contents": "Members join for free and… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/heuristic_classification-filtered-pile-50M.arabic-enterprise
Arabic Enterprise
This is a proposed dataset for evaluating enterprise use cases for LLMs in Arabic.
e3c-crf-english
Dataset description
Here we realease the dataset to perform the Case Report Forms filling task obtained from The European Clinical Case Corpus as described in the paper Converting Annotated Clinical Cases into Structured Case Report Forms presented at the BioNLP workshop at ACL 2025.
The dataset is composed by a set patients with related clinical_note that describe their history and conditions. Each patient is uniquely identified by the document_id column.
The task consists of… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/e3c-crf-english.CoReBench_v1crf-vmaf-training-data
CRF <-> VMAF training data
Training datasets behind Alice-1 - a LightGBM model that predicts the CRF value needed to hit a target VMAF score for a given video segment, codec and resolution.
Contents
File
Rows
Description
training_table.parquet
1,762,232
Main table. Feature aggregates -> CRF for a target VMAF, per codec/resolution.
probe_data.parquet
52,316
Probe-encode data. Two 2s probe encodes per (segment, codec, resolution) cell with measured VMAF… See the full description on the dataset page: https://huggingface.co/datasets/jakubkrapiec/crf-vmaf-training-data.image2struct-musicsheet-v1
Image2Struct - Music Sheet
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images.
This subdataset focuses on Music sheets. The model is given an image of the expected output with the prompt:
Please generate the Lilypond… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-musicsheet-v1.e3c-crf-italianHere we realease the dataset to perform the Case Report Forms filling task obtained from The European Clinical Case Corpus as described in the paper Converting Annotated Clinical Cases into Structured Case Report Forms presented at the BioNLP workshop at ACL 2025.
The dataset is composed by a set patients with related clinical_note that describe their history and conditions. Each patient is uniquely identified by the document_id column.
The task consists of filling a set of items for each… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/e3c-crf-italian.dyspnea-crf-development
Dataset description
This dataset contains the development annotated CRFs for the CRF:filling Shared Task at CL4Health2026.
The clinical notes have been collected, anonymized and annotated at the San Giovanni Bosco (SGB) hospital, Turin, Italy.
There are two splits, each representing a different language: en (English) and it (Italian).
Each example (80 in total) in the dataset is composed by:
document_id: clinical note identifier
clinical_note: the note reporting on the patient's… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/dyspnea-crf-development.vidaio-ml-crf
vidaio ML CRF dataset + ExtraTrees prior
Validator-style labeled CRF training set and trained ExtraTrees model for the vidaio miner encoder.
Contents
Path
Description
labels/labels_v2.json
300 rows (100 clips × VMAF thresholds 85/89/93), SVT preset 6
labels/labels.json
Earlier label set
labels/synth_refs_manifest.json
Synth ref metadata
models/crf_extratrees.joblib
Trained ExtraTrees CRF prior (cv MAE ≈ 1.35)
synth_refs/*.mp4… See the full description on the dataset page: https://huggingface.co/datasets/Realfencer/vidaio-ml-crf.dyspnea-crf-train
Dataset description
This dataset contains the train annotated CRFs for the CRF:filling Shared Task at CL4Health2026.
The clinical notes have been collected, anonymized and annotated at the San Giovanni Bosco (SGB) hospital, Turin, Italy.
There are two splits, each representing a different language: en (English) and it (Italian).
Each example (10 in total) in the dataset is composed by:
document_id: clinical note identifier
clinical_note: the note reporting on the patient's… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/dyspnea-crf-train.crf-benchmark-results
CRF Benchmark Results
Benchmark results comparing the Cellular Reasoning Fabric (CRF) against a parameter-matched Transformer across 5 language modeling tasks.
Dataset Description
This dataset contains the full experimental results from our CRF vs Transformer comparison, including:
Training loss and perplexity curves (per epoch)
Validation loss and perplexity curves (per epoch)
Parameter counts and FLOP estimates
Inference profiling (latency, memory)
CRF-specific… See the full description on the dataset page: https://huggingface.co/datasets/YasirUsman/crf-benchmark-results.synthetic-crf-train
Dataset description
This dataset is derived from NLP-FBK/e3c-crf-english and NLP-FBK/e3c-crf-italian for the sake of the CRF:filling Shared Task at CL4Health2026. Please refer to the original datasets for anything unrelated to the Shared Task.
Here we realease the synthetic Case Report Forms built for 71/80 (English/Italian) patients with different medical conditions. Each example in the dataset is composed by:
document_id: patient/clinical note identifier coming from the original… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/synthetic-crf-train.crf-qa-trustworthiness
crf-qa-trustworthiness
Anonymous dataset artifact for reproducing the main-paper experiments on epistemic trustworthiness.
This repository intentionally contains no author, institution, cluster, or local path metadata.
Original data source: NLP-FBK/dyspnea-crf-train
LICENCE: CC BY-NC
CRF-Gatekeeper-Training-DataExperiment_PII_GenTestnoAug_5seed_CRFLegalSeg_Hier_BiLSTM_CRF_Predictionseurope-owid-use-of-crf-tools-by-providers-of-dev-cooperation
Use Of Crf Tools By Providers Of Dev Cooperation | Europe (Our World in Data)
🇪🇺 42 observations · 22 Europe countries · 2016–2018 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 42 observations of Use Of Crf Tools By Providers Of Dev Cooperation data across 22 Europe countries, spanning 2016–2018.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Use Of Crf Tools By Providers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-use-of-crf-tools-by-providers-of-dev-cooperation.visor-small-basic-crf10-rehearsal-WA2-Track1-wa2-track1-public1000-rehearsal-54054cb7306fCRF1crf-datasetCRFExperiment_PII_GenTestAug_5seed_CRFcrf-metrics-board
