datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSRC
Congolese Speech Radio Corpus (CSRC)
The Congolese Speech Radio Corpus (CSRC) is an unlabelled radio-speech corpus covering Lingala, Kikongo, and Tshiluba.
This Hugging Face release was prepared by Bantu Languages Initiative from the CSRC component of Speech Recognition Datasets for Congolese Languages. Its purpose is to make the radio archives easier to use for self-supervised speech learning, ASR pretraining, acoustic adaptation, language identification, and robust speech… See the full description on the dataset page: https://huggingface.co/datasets/BantuLanguagesInitiative/CSRC.cs_restaurantsThe task is generating responses in the context of a (hypothetical) dialogue
system that provides information about restaurants. The input is a basic
intent/dialogue act type and a list of slots (attributes) and their values.
The output is a natural language sentence.debian_csrcm3-retrieve-csrldbc-csrcs_restaurants
Dataset Card for Czech Restaurant
Dataset Summary
This is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language. It originated as a translation of the English San Francisco Restaurants dataset by Wen et al. (2015). The domain is restaurant information in Prague, with random/fictional values. It includes input dialogue acts and the corresponding outputs in Czech.
Supported Tasks and Leaderboards
other-intent-to-text:… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/cs_restaurants.cs-raw-28Bcsrsef-artifactstask270_csrg_counterfactual_context_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task270_csrg_counterfactual_context_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task270_csrg_counterfactual_context_generation.csrrg_findings
Dataset Card for CSRRG Findings
Dataset Description
This dataset contains structured chest X-ray radiology reports that include both findings and impression sections.
Each report is decomposed from unstructured text into standardized sections organized by anatomical systems, facilitating natural language processing and clinical AI research.
Dataset Summary
The CSRRG Findings dataset provides comprehensive structured radiology reports with detailed findings… See the full description on the dataset page: https://huggingface.co/datasets/erjui/csrrg_findings.task269_csrg_counterfactual_story_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task269_csrg_counterfactual_story_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task269_csrg_counterfactual_story_generation.csrrg_impression
Dataset Card for CSRRG Impression
Dataset Description
This dataset contains structured chest X-ray radiology reports focusing on impression sections.
Reports in this dataset may have less detailed findings sections or primarily consist of impressions, making it ideal for training models focused on generating concise clinical impressions from imaging observations.
Dataset Summary
The CSRRG Impression dataset provides structured radiology reports where the… See the full description on the dataset page: https://huggingface.co/datasets/erjui/csrrg_impression.CSRTDataset derived from CSRT: Evaluation and Analysis of LLMs using Code-Switching Red-Teaming Dataset submitted to the NeurIPS 2024 Datasets and Benchmarks track
We introduce code-switching red-teaming, a simple yet effective red-teaming technique that simultaneously tests the multilingual capabilities and safety of LLMs
Keywords
LLM, Evaluation, Safety, Multilingual, Red-teaming, Code-switching
Abstract
Recent studies in large language models (LLMs) shed light on their… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CSRT.news21-instructions-mteb-CSR-L
news21-instructions-codeswitching-mteb
Code-switching version of jhu-clsp/news21-instructions-mteb, with queries and instructions rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments
qrel_diff: Changes in relevance judgments
top_ranked: Top ranked documents for each query… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/news21-instructions-mteb-CSR-L.core17-instructions-mteb-CSR-L
core17-instructions-codeswitching-mteb
Code-switching version of jhu-clsp/core17-instructions-mteb, with queries and instructions rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments
qrel_diff: Changes in relevance judgments
top_ranked: Top ranked documents for each query… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/core17-instructions-mteb-CSR-L.LowAUG-CSRtrec-covid-CSR-L
TRECCOVID-CodeSwitching
An MTEB dataset
Massive Text Embedding Benchmark
Code-switching version of mteb/trec-covid, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments (qrels)
Code-switching additions:
queries_zh_en: Chinese-English code-switching queries… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/trec-covid-CSR-L.CSRTThe dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license.
robust04-instructions-mteb-CSR-L
robust04-instructions-codeswitching-mteb
Code-switching version of jhu-clsp/robust04-instructions-mteb, with queries and instructions rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments
qrel_diff: Changes in relevance judgments
top_ranked: Top ranked documents for each query… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/robust04-instructions-mteb-CSR-L.Main-CSR-Whisper-L-turboV3clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1Clinical Quad Data Cut Timing Database Lock Pressure Query Backlog CSR Narrative Drift v0.1
Each row is a trial monthly snapshot.
Core quad
Data cut timingDatabase lock pressureQuery backlogCSR narrative drift
Target
label_regulatory_issue_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1.HumanEvalRetrieval-CSR-L
HumanEvalRetrieval-CodeSwitching
Code-switching version of mteb/HumanEvalRetrieval, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
qrels: Original relevance judgments
Code-switching additions:
queries_zh_en: Chinese-English code-switching queries
queries_ja_en: Japanese-English code-switching… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/HumanEvalRetrieval-CSR-L.CSR-12K-iter2webis-touche2020-v3-CSR-L
Touche2020-v3-CodeSwitching
An MTEB dataset
Massive Text Embedding Benchmark
Code-switching version of mteb/webis-touche2020-v3, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments (qrels)
Code-switching additions:
queries_zh_en: Chinese-English code-switching… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/webis-touche2020-v3-CSR-L.csrrg_ift_dataset
CSRRG Instruction Fine-Tuning Dataset
Dataset Details
Dataset type: CSRRG IFT is a large-scale instruction-following dataset for chest X-ray report generation.
It is constructed for visual instruction tuning and building large multimodal models capable of generating structured radiology reports.
Dataset composition: The dataset contains approximately 1.6 million instruction-following examples across 5 subsets, covering both Structured Radiology Report Generation (SRRG)… See the full description on the dataset page: https://huggingface.co/datasets/erjui/csrrg_ift_dataset.csrsef-pubmedqa-artifactsCSR-12K-iter1m3-retrieve-csr-cardiology-1kqwen3_0.6b-rlvr_task269_csrg_counterfactual_story_generationImage-Colorization
