datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clinical-synthetic-text-kg
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on external knowledge graphs.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-kg.clinical-synthetic-text-llm
Data Description
We release the synthetic data generated using the method described in the paper Knowledge-Infused Prompting: Assessing and Advancing Clinical Text Data Generation with Large Language Models
(ACL 2024 Findings). The external knowledge we use is based on LLM-generated topics and writing styles.
Generated Datasets
The original train/validation/test data, and the generated synthetic training data are listed as follows. For each dataset, we generate 5000… See the full description on the dataset page: https://huggingface.co/datasets/ritaranx/clinical-synthetic-text-llm.Aegis-AI-Content-Safety-Dataset-2.0
🛡️ Nemotron Content Safety Dataset V2
The Nemotron Content Safety Dataset V2, formerly known as Aegis AI Content Safety Dataset 2.0, is comprised of 33,416 annotated interactions between humans and LLMs, split into 30,007 training samples, 1,445 validation samples, and 1,964 test samples. This release is an extension of the previously published Nemotron Content Safety Dataset V1.
To curate the dataset, we use the HuggingFace version of human preference data about… See the full description on the dataset page: https://huggingface.co/datasets/ritatai727/Aegis-AI-Content-Safety-Dataset-2.0.Ritaxendotsocr_bank_statement_1K_v2
dotsocr_bank_statement_1K
Dataset Description
This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab.
DotsOCR Format Features
Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right)
Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_1K_v2.dotsocr_bank_statement_half
dotsocr_bank_statement_half
Dataset Description
This dataset contains OCR and layout analysis training data formatted according to DotsOCR specifications by rednote-hilab.
DotsOCR Format Features
Proper Reading Order: Layout elements are sorted according to natural reading order (top to bottom, left to right)
Validated Categories: All categories conform to DotsOCR's specification: ['Caption', 'Footnote', 'Formula', 'List-item', 'Page-footer', 'Page-header'… See the full description on the dataset page: https://huggingface.co/datasets/rita1706/dotsocr_bank_statement_half.itaeval-results
Dataset Card for Evaluation run of CohereForAI/aya-expanse-8b
Dataset automatically created during the evaluation run of model CohereForAI/aya-expanse-8b
The dataset is composed of 330 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 58 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/RiTA-nlp/itaeval-results.
