datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CSRC
Congolese Speech Radio Corpus (CSRC)
The Congolese Speech Radio Corpus (CSRC) is an unlabelled radio-speech corpus covering Lingala, Kikongo, and Tshiluba.
This Hugging Face release was prepared by Bantu Languages Initiative from the CSRC component of Speech Recognition Datasets for Congolese Languages. Its purpose is to make the radio archives easier to use for self-supervised speech learning, ASR pretraining, acoustic adaptation, language identification, and robust speech… See the full description on the dataset page: https://huggingface.co/datasets/BantuLanguagesInitiative/CSRC.debian_csrccs_restaurants
Dataset Card for Czech Restaurant
Dataset Summary
This is a dataset for NLG in task-oriented spoken dialogue systems with Czech as the target language. It originated as a translation of the English San Francisco Restaurants dataset by Wen et al. (2015). The domain is restaurant information in Prague, with random/fictional values. It includes input dialogue acts and the corresponding outputs in Czech.
Supported Tasks and Leaderboards
other-intent-to-text:… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/cs_restaurants.ldbc-csrtask270_csrg_counterfactual_context_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task270_csrg_counterfactual_context_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task270_csrg_counterfactual_context_generation.task269_csrg_counterfactual_story_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task269_csrg_counterfactual_story_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task269_csrg_counterfactual_story_generation.cs-raw-28BCSRTDataset derived from CSRT: Evaluation and Analysis of LLMs using Code-Switching Red-Teaming Dataset submitted to the NeurIPS 2024 Datasets and Benchmarks track
We introduce code-switching red-teaming, a simple yet effective red-teaming technique that simultaneously tests the multilingual capabilities and safety of LLMs
Keywords
LLM, Evaluation, Safety, Multilingual, Red-teaming, Code-switching
Abstract
Recent studies in large language models (LLMs) shed light on their… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CSRT.news21-instructions-mteb-CSR-L
news21-instructions-codeswitching-mteb
Code-switching version of jhu-clsp/news21-instructions-mteb, with queries and instructions rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments
qrel_diff: Changes in relevance judgments
top_ranked: Top ranked documents for each query… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/news21-instructions-mteb-CSR-L.core17-instructions-mteb-CSR-L
core17-instructions-codeswitching-mteb
Code-switching version of jhu-clsp/core17-instructions-mteb, with queries and instructions rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments
qrel_diff: Changes in relevance judgments
top_ranked: Top ranked documents for each query… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/core17-instructions-mteb-CSR-L.robust04-instructions-mteb-CSR-L
robust04-instructions-codeswitching-mteb
Code-switching version of jhu-clsp/robust04-instructions-mteb, with queries and instructions rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments
qrel_diff: Changes in relevance judgments
top_ranked: Top ranked documents for each query… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/robust04-instructions-mteb-CSR-L.HumanEvalRetrieval-CSR-L
HumanEvalRetrieval-CodeSwitching
Code-switching version of mteb/HumanEvalRetrieval, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
qrels: Original relevance judgments
Code-switching additions:
queries_zh_en: Chinese-English code-switching queries
queries_ja_en: Japanese-English code-switching… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/HumanEvalRetrieval-CSR-L.webis-touche2020-v3-CSR-L
Touche2020-v3-CodeSwitching
An MTEB dataset
Massive Text Embedding Benchmark
Code-switching version of mteb/webis-touche2020-v3, with queries rewritten in Chinese-English and Japanese-English code-switching styles.
Dataset Structure
The dataset contains the following configurations:
From original dataset (unchanged):
corpus: Original corpus documents
default: Original relevance judgments (qrels)
Code-switching additions:
queries_zh_en: Chinese-English code-switching… See the full description on the dataset page: https://huggingface.co/datasets/UTokyo-Yokoya-Lab/webis-touche2020-v3-CSR-L.qwen3_0.6b-rlvr_task269_csrg_counterfactual_story_generationm3-retrieve-csr-cardiology-1kcsrc-pseudo-testqwen3_0.6b-rlvr_task270_csrg_counterfactual_context_generationfnt-colpali-esrs-csrdnaver-economy-news2stockcsrc-lingala-pseudo-punctcs_repl_ai_alpacacsrc-lingala-pseudoflan_combined_task270_csrg_counterfactual_context_generation
