CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Polyglot-or-Not /Fact-Completion Dataset Card Homepage: https://bit.ly/ischool-berkeley-capstone Repository: https://github.com/daniel-furman/Capstone Point of Contact: daniel_furman@berkeley.edu Dataset Summary This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models. Test Description Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.texttext-generation100K<n<1M13 likes1.4k downloads3y agoHugging Face02ahmad21omar /Polyglot-Thoughts-SFT-Collection Polyglot Thoughts SFT Collection Polyglot Thoughts SFT Collection is a large-scale supervised fine-tuning (SFT) corpus for reasoning-oriented language models. It combines, filters, deduplicates, and language-extends a broad set of public reasoning datasets into a single uniform schema centred on chain-of-thought reasoning traces. The final corpus contains 23,896,757 examples and roughly 123 billion tokens, spanning six languages (English, German, French, Italian, Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/Polyglot-Thoughts-SFT-Collection.texttext-generation10M<n<100M0 likes667 downloads3mo agoHugging Face03hac541309 /polyglot-ko-tokenizer-corpus Dataset Card for "polyglot-ko-tokenizer-corpus" More Information needed text10M<n<100M1 likes590 downloads3y agoHugging Face04swiss-ai /polyglotoxicitypromptstabular100K<n<1M0 likes336 downloads1y agoHugging Face05polyglot-tagger /wikipedia-language-snippets-filtered Wikipedia Snippets (Filtered) Filtered sentence snippets in Wikipedia, by taking the first 60% of an article after filtering for stubs. Minor Latin groups are additionally filtered again for English leakage. Sentences are mostly filtered out for non matching scripts, such as Arabic in a Cyrllic language. Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From wikimedia/wikipedia Licensing… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/wikipedia-language-snippets-filtered.texttext-generation10M<n<100M0 likes274 downloads5mo agoHugging Face06polyglot-tagger /finetranslations-filtered Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet From HuggingFaceFW/finetranslations Licensing Information The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use. Citation Information @misc{penedo2026finetranslations, title={FineTranslations}, author={Guilherme… See the full description on the dataset page: https://huggingface.co/datasets/polyglot-tagger/finetranslations-filtered.texttext-classification1M<n<10M0 likes250 downloads5mo agoHugging Face07ahmad21omar /Polyglot-Thoughts-RL-Collection Polyglot Thoughts RL Collection Polyglot Thoughts RL Collection is a large-scale, curated corpus for reinforcement learning from verifiable rewards (RLVR) of reasoning-oriented language models. It combines, filters, normalises, and deduplicates a broad set of public RL datasets into a single uniform schema in which every row carries a machine-verifiable ground-truth signal — math equivalence, code execution, Prolog rule induction, schema validation, multiple-choice… See the full description on the dataset page: https://huggingface.co/datasets/ahmad21omar/Polyglot-Thoughts-RL-Collection.texttext-generation1M<n<10M0 likes243 downloads3mo agoHugging Face08ljvmiranda921 /PolyglotTeachers-SFT-Synth-Data Website: ljvmiranda921.github.io/polyglot-teachers/ PolyglotTeachers-SFT-Synth This dataset contains synthetic supervised fine-tuning examples generated by the best teacher we found in the paper Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation, where we systematically characterize what makes a good teacher model. It contains examples across six languages: Arabic, Czech, German, Indonesian, Japanese, Spanish, and Tagalog. Note: In… See the full description on the dataset page: https://huggingface.co/datasets/ljvmiranda921/PolyglotTeachers-SFT-Synth-Data.texttext-generation100K<n<1M3 likes165 downloads2mo agoHugging Face09polyglot-tagger /tatoeba-filtered Files Each file is in this format for languages in ISO 639 2-letter codes: train/en/en.parquet train/es/es.parquet texttext-classification1M<n<10M0 likes150 downloads5mo agoHugging Face10DCAgent3 /aider_polyglot_a3_rl_DCAgent_inferredbugs_sandboxes_verifier_55_8B_20260526_230125textn<1K0 likes136 downloads4mo agoHugging Face11DCAgent3 /aider_polyglot_GLM_4_7_swesmith_sandboxes_with_tests_oracle_verified_120s_maxepf48469bftextn<1K0 likes125 downloads4mo agoHugging Face12DCAgent3 /aider_polyglot_a3_rl_DCAgent_selfinstruct_naive_sandboxes_2_verified_70_8B_202690541bcetextn<1K0 likes118 downloads4mo agoHugging Face13nathansutton /chad-polyglot-runs chad run data Run output from the benchmarks of chad, a local coding agent for Apple Silicon, kept out of the code repository. Every polyglot run here is a row of benchmarks/polyglot/RUNS.md with the sha256 of its trials.jsonl, and benchmarks/polyglot/fetch.py --label <label> downloads one and refuses it if the hash differs. path what polyglot/<label>/ one polyglot run: meta.json (chad version and commit, model, context limit, every CHAD_* variable), trials.jsonl (one… See the full description on the dataset page: https://huggingface.co/datasets/nathansutton/chad-polyglot-runs.tabularn<1K0 likes108 downloads3d agoHugging Face14polyglots /MADLAD_CulturaX_cleanedtext10M<n<100M21 likes106 downloads2y agoHugging Face15Reubencf /PolyglotAudio Citation If you use this dataset in your research or downstream work, please cite: @misc{polyglot_audio_2026, author = {Fernandes, Reuben Chagas}, title = {PolyglotAudio: Multilingual Audio Pre-training Corpus}, year = {2026}, publisher = {Hugging Face}, howpublished = {\url{https://huggingface.co/datasets/Reubencf/PolyglotAudio}} } APA-style: Reuben Chagas Fernandes (2026). PolyglotAudio: Multilingual Audio Pre-training Corpus [Dataset]. Hugging… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/PolyglotAudio.audio1M<n<10M2 likes85 downloads5mo agoHugging Face16DCAgent2 /DCAgent2_aider_polyglot_DCAgent2_swesmith-stack-reason_20260127_005703text1K<n<10K0 likes84 downloads8mo agoHugging Face17polyglot-tagger /nlp-noise-snippets Synthetic Noise Pool For Text Classification purposes, as many models may consider code snippets, html artifacts, and math as "English". Around 50K are latex snippets from im2latex-100k texttext-classification100K<n<1M0 likes81 downloads5mo agoHugging Face18DCAgent2 /DCAgent2_aider_polyglot_DCAgent2_nl2bash-swesmith-reason_20260127_005703text1K<n<10K0 likes61 downloads8mo agoHugging Face19darkknight25 /polyglot_paylods_datasets Polyglot Payloads Dataset for Cybersecurity Training Overview This dataset, polyglot_payloads.jsonl, is a curated collection of 500 polyglot payloads designed for training AI models in cybersecurity, specifically for red team operations and vulnerability detection. The dataset includes payloads targeting common web vulnerabilities such as Cross-Site Scripting (XSS), SQL Injection (SQLi), Local File Inclusion (LFI), Remote Code Execution (RCE), and Server-Side Template… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/polyglot_paylods_datasets.textn<1K0 likes53 downloads1y agoHugging Face20DCAgent2 /DCAgent2_aider_polyglot_DCAgent_freelancer-long-instruction-filter_Qwen3-8B_2021a649aaetext1K<n<10K0 likes53 downloads8mo agoHugging Face21DCAgent2 /DCAgent2_aider_polyglot_DCAgent_tbench_oracle_solutions_terminus_20260125_202032textn<1K0 likes52 downloads8mo agoHugging Face22ontocord /PolygloToxicityPrompts_permissive Polyglo PTP Permissive Permissive-source subset of the PTP configs from ToxicityPrompts/PolygloToxicityPrompts, filtered using the local URL/source permissiveness rules in this workspace. WildChat configs are not included because they do not provide row-level source URLs. Source split: full for ptp-* configs. Rows scanned: 425,000. Rows kept: 1,649. See polyglo_ptp_permissive_stats_620972.json for kept/discard reason distributions and per-language counts. tabulartext-generation1K<n<10K0 likes42 downloads4mo agoHugging Face23ndandanov /polyglot-modelsThis dataset repository shall serve as a mirror hosting models for polyglot. Availability Please note that currently the only languages, which have all models, are English (en) and Bulgarian (bg). Other languages may have partial support. Adding missing models In case you have previously downloaded polyglot language models, which are not available in this repo, please open a Pull Request. License All rights belong to the original authors. Please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/ndandanov/polyglot-models.audion<1K1 likes38 downloads4mo agoHugging Face24DCAgent2 /DCAgent2_aider_polyglot_DCAgent_tbench-dev-71-nl2bash-bugsseq_Qwen3-8B-8nodes-sfd6ea99ftextn<1K0 likes35 downloads8mo agoHugging Face25DCAgent2 /DCAgent2_aider_polyglot_DCAgent2_test2-tbench-dev-71-qwen3-8b-8nodes-sync_2026087639023textn<1K0 likes35 downloads8mo agoHugging Face26hac541309 /polyglot-ko-tokenizer-corpus-merge_ws Dataset Card for "polyglot-ko-tokenizer-corpus-merge_ws" More Information needed text10M<n<100M0 likes29 downloads3y agoHugging Face27jkswin /polyglot_fr_en_subset Dataset Card for "polyglot_fr_en_subset" More Information needed text10K<n<100K0 likes28 downloads3y agoHugging Face28DCAgent2 /DCAgent2_aider_polyglot_laion_glm46-qasper-maxeps-131k_20260121_072458textn<1K0 likes27 downloads8mo agoHugging Face29DCAgent2 /aider_polyglot_Qwen3_30B_A3B_Instruct_2507_20260425_063500-tracestextn<1K0 likes27 downloads5mo agoHugging Face30DCAgent2 /DCAgent2_aider_polyglot_DCAgent_freelancer-askllm-filtered-sandboxes-traces-terf2137443text1K<n<10K0 likes26 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.