datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spider-corpus-testLink to original dataset: https://yale-lily.github.io/spider
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S. and Zhang, Z., 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv preprint arXiv:1809.08887.
un_corpus_for_sitemap_test
Dataset Card for "un_corpus_for_sitemap_test"
More Information needed
msmacro-test-corpustest-corpustest_corpus
Dataset Card for "test_corpus"
More Information needed
corpus-siarad-test-setAudio-Turing-Test-Corpus
📚 Audio Turing Test Corpus
A high‑quality, multidimensional Chinese transcript corpus designed to evaluate whether a machine‑generated speech sample can fool human listeners—the “Audio Turing Test.”
About Audio Turing Test (ATT)
ATT is an evaluation framework with a standardized human evaluation protocol and an accompanying dataset, aiming to resolve the lack of unified protocols in TTS evaluation and the difficulty in comparing multiple TTS systems. To further support… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/Audio-Turing-Test-Corpus.market-feed-freshness-test-corpus
Synthetic Market-Feed Freshness and Availability State Test Corpus
This dataset contains 72 deterministic, fully synthetic cases for testing how a market-data interface or service classifies feed freshness, source availability, fallback use, invalid values and invalid timing inputs.
It contains no observed market prices, customer records, credentials, personal data or production telemetry. It does not evaluate any named provider and is not trading or financial advice.… See the full description on the dataset page: https://huggingface.co/datasets/musheghmanukyan/market-feed-freshness-test-corpus.pdf-to-markdown-test-corpus
PDF-to-Markdown Test Corpus
18 small PDFs, each built to break a PDF-to-Markdown converter in one specific
way, plus the measured output of one converter against all of them.
If you are writing a converter, or choosing one, the hard part is not the happy
path. It is knowing what happens when a document has two columns, or a heading
that is only bold, or a scanned page in the middle. This corpus is meant to make
that testable in about a minute, and to give you somewhere to point… See the full description on the dataset page: https://huggingface.co/datasets/BayKolm/pdf-to-markdown-test-corpus.The_Gab_Hate_Corpus_ghc_test_originalmammut-corpus-venezuela-test-set
mammut-corpus-venezuela
HuggingFace Dataset for testing purposes. The train dataset is mammut/mammut-corpus-venezuela.
1. How to use
How to load this dataset directly with the datasets library:
>>> from datasets import load_dataset>>> dataset = load_dataset("mammut/mammut-corpus-venezuela")
2. Dataset Summary
mammut-corpus-venezuela is a dataset for Spanish language modeling. This dataset comprises a large number of Venezuelan and Latin-American Spanish texts… See the full description on the dataset page: https://huggingface.co/datasets/mammut/mammut-corpus-venezuela-test-set.dataset-pt-corpus-redwhale2-testcomplete_pira_test_corpus1_ptbr_llama3_alpaca_181marqo_gs_wfash_1m_test_subset_corpus_tevatronarabic_corpus_testtest-speech-corpuscomplete_pira_test_corpus1_en_llama3_alpaca_181complete_pira_test_corpus2_en_llama3_alpaca_46translate-corpus-testpretrain-corpus-testcomplete_pira_test_corpus2_ptbr_llama3_alpaca_46python-doctest-corpus-test
Python Doctest Corpus
A curated corpus of Python doctest examples designed for training Python-to-Rust transpilers and testing code translation systems.
Dataset Description
This dataset contains Python function signatures, doctest inputs, and expected outputs that serve as high-quality training data for:
Transpilation training: Teaching models to translate Python patterns to Rust
Test validation: Verifying that transpiled code produces correct outputs
Code understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/python-doctest-corpus-test.corpus-creator-testThis dataset was created using Corpus Creator. This dataset was created by paring a corpus of texts into chunks of sentences using Llama Index.
test_corpuscorpus_testgeorgian-corpus-testThe_Gab_Hate_Corpus_ghc_test_translateJzuluaga_atcosim_corpus_test_embeddingsranWang_un_corpus_for_sitemap_testCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données ranWang/un_corpus_for_sitemap_test.
SpokenSQuAD_test_audio_corpus_dedup
