CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01target-benchmark /spider-corpus-testLink to original dataset: https://yale-lily.github.io/spider Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S. and Zhang, Z., 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. arXiv preprint arXiv:1809.08887. textn<1K0 likes373 downloads2y agoHugging Face02ranWang /un_corpus_for_sitemap_test Dataset Card for "un_corpus_for_sitemap_test" More Information needed text100K<n<1M0 likes243 downloads4y agoHugging Face03Carlisle /msmacro-test-corpustext1M<n<10M0 likes93 downloads5y agoHugging Face04akunasoftware /test-corpusdocumentn<1K0 likes55 downloads3mo agoHugging Face05rdev12 /test_corpus Dataset Card for "test_corpus" More Information needed text1K<n<10K0 likes46 downloads4y agoHugging Face06wanasash /corpus-siarad-test-setaudio1K<n<10K0 likes40 downloads2y agoHugging Face07meituan-longcat /Audio-Turing-Test-Corpus 📚 Audio Turing Test Corpus A high‑quality, multidimensional Chinese transcript corpus designed to evaluate whether a machine‑generated speech sample can fool human listeners—the “Audio Turing Test.” About Audio Turing Test (ATT) ATT is an evaluation framework with a standardized human evaluation protocol and an accompanying dataset, aiming to resolve the lack of unified protocols in TTS evaluation and the difficulty in comparing multiple TTS systems. To further support… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/Audio-Turing-Test-Corpus.text-to-speech1K<n<10K3 likes38 downloads1y agoHugging Face08musheghmanukyan /market-feed-freshness-test-corpus Synthetic Market-Feed Freshness and Availability State Test Corpus This dataset contains 72 deterministic, fully synthetic cases for testing how a market-data interface or service classifies feed freshness, source availability, fallback use, invalid values and invalid timing inputs. It contains no observed market prices, customer records, credentials, personal data or production telemetry. It does not evaluate any named provider and is not trading or financial advice.… See the full description on the dataset page: https://huggingface.co/datasets/musheghmanukyan/market-feed-freshness-test-corpus.textn<1K1 likes31 downloads2d agoHugging Face09BayKolm /pdf-to-markdown-test-corpus PDF-to-Markdown Test Corpus 18 small PDFs, each built to break a PDF-to-Markdown converter in one specific way, plus the measured output of one converter against all of them. If you are writing a converter, or choosing one, the hard part is not the happy path. It is knowing what happens when a document has two columns, or a heading that is only bold, or a scanned page in the middle. This corpus is meant to make that testable in about a minute, and to give you somewhere to point… See the full description on the dataset page: https://huggingface.co/datasets/BayKolm/pdf-to-markdown-test-corpus.documenttext-classificationn<1K0 likes30 downloads3d agoHugging Face10juliadollis /The_Gab_Hate_Corpus_ghc_test_originaltabular1K<n<10K0 likes23 downloads2y agoHugging Face11mammut /mammut-corpus-venezuela-test-set mammut-corpus-venezuela HuggingFace Dataset for testing purposes. The train dataset is mammut/mammut-corpus-venezuela. 1. How to use How to load this dataset directly with the datasets library: >>> from datasets import load_dataset>>> dataset = load_dataset("mammut/mammut-corpus-venezuela") 2. Dataset Summary mammut-corpus-venezuela is a dataset for Spanish language modeling. This dataset comprises a large number of Venezuelan and Latin-American Spanish texts… See the full description on the dataset page: https://huggingface.co/datasets/mammut/mammut-corpus-venezuela-test-set.tabular100K<n<1M0 likes19 downloads4y agoHugging Face12TwinDoc /dataset-pt-corpus-redwhale2-testtext10K<n<100K0 likes17 downloads2y agoHugging Face13Gramacho /complete_pira_test_corpus1_ptbr_llama3_alpaca_181textn<1K0 likes14 downloads2y agoHugging Face14LouisDo2108 /marqo_gs_wfash_1m_test_subset_corpus_tevatronimage100K<n<1M0 likes14 downloads5mo agoHugging Face15moqarashad /arabic_corpus_testtext1K<n<10K0 likes12 downloads2y agoHugging Face16Nyanmero /test-speech-corpusaudion<1K0 likes11 downloads2y agoHugging Face17Gramacho /complete_pira_test_corpus1_en_llama3_alpaca_181textn<1K0 likes10 downloads2y agoHugging Face18Gramacho /complete_pira_test_corpus2_en_llama3_alpaca_46textn<1K0 likes10 downloads2y agoHugging Face19mssongit /translate-corpus-testtextn<1K0 likes9 downloads2y agoHugging Face20BroAlanTaps /pretrain-corpus-testtextn<1K0 likes8 downloads1y agoHugging Face21Gramacho /complete_pira_test_corpus2_ptbr_llama3_alpaca_46textn<1K0 likes7 downloads2y agoHugging Face22paiml /python-doctest-corpus-test Python Doctest Corpus A curated corpus of Python doctest examples designed for training Python-to-Rust transpilers and testing code translation systems. Dataset Description This dataset contains Python function signatures, doctest inputs, and expected outputs that serve as high-quality training data for: Transpilation training: Teaching models to translate Python patterns to Rust Test validation: Verifying that transpiled code produces correct outputs Code understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/python-doctest-corpus-test.texttext-generation1K<n<10K0 likes7 downloads10mo agoHugging Face23davanstrien /corpus-creator-testThis dataset was created using Corpus Creator. This dataset was created by paring a corpus of texts into chunks of sentences using Llama Index. textn<1K0 likes6 downloads2y agoHugging Face24Atipico1 /test_corpustext10K<n<100K0 likes6 downloads2y agoHugging Face25WendyHoang /corpus_testtextn<1K0 likes6 downloads2y agoHugging Face26RichNachos /georgian-corpus-testtabular10K<n<100K0 likes5 downloads2y agoHugging Face27akcit-ijf /The_Gab_Hate_Corpus_ghc_test_translatetabular1K<n<10K0 likes5 downloads2y agoHugging Face28HaninZ /Jzuluaga_atcosim_corpus_test_embeddingstext1K<n<10K0 likes5 downloads2y agoHugging Face29french-datasets /ranWang_un_corpus_for_sitemap_testCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données ranWang/un_corpus_for_sitemap_test. translation0 likes5 downloads1y agoHugging Face30rachel0411 /SpokenSQuAD_test_audio_corpus_dedupaudio1K<n<10K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.