CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sukratii /jailbreak-filter-outputstabularn<1K1 likes91 downloads5mo agoHugging Face02Telugu-LLM-Labs /telugu_alpaca_yahma_cleaned_filtered_romanizedtext10K<n<100K19 likes58 downloads3y agoHugging Face03NotShrirang /email-spam-filtertabulartext-classification1K<n<10K11 likes55 downloads1y agoHugging Face04Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes50 downloads3y agoHugging Face05bpm2007 /shelfglance-ucp-category-filter Shopify's agent-commerce category filter didn't filter on any of 190 stores Every Shopify store answers an agent-commerce endpoint whose schema declares a category filter. Sent to 200 live stores with a price-filter control: 186 ignored it, 4 rejected every value, 0 filtered. One row per Shopify storefront that answered its own agent-commerce endpoint (POST /api/ucp/mcp) on 2 September 2026: 190 of 200 sampled from a corpus of 10,099. Each row records what search_catalog… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ucp-category-filter.tabularn<1K0 likes44 downloads16d agoHugging Face06Weyaxi /HelpSteer-filtered HelpSteer-filtered This dataset is a highly filtered version of the nvidia/HelpSteer dataset. ❓ How this dataset was filtered: I calculated the sum of the columns ["helpfulness," "correctness," "coherence," "complexity," "verbosity"] and created a new column named sum. I changed some column names and added a empty column to match the Alpaca format. The dataset was then filtered to include only those entries with a sum greater than or equal to 16. 🧐 More… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/HelpSteer-filtered.tabular1K<n<10K4 likes43 downloads3y agoHugging Face07huggingface-projects /filter-bad-modelstextn<1K0 likes38 downloads3y agoHugging Face08danielheinz /telekom-backtrans-paraphrase-filteredThis is a filtered version of Philip May's German paraphrase dataset. The dataset has been filtered for the sake of convenience, since smaller devices do not support such large files. All text pairs in the dataset are paraphrases, and are therefore labelled 1. As such, the dataset is well-suited for use in conjunction with the multiple negatives ranking loss. As the original author suggests, the dataset has been filtered, mostly following the guidelines set by the author. Any row that doesn't… See the full description on the dataset page: https://huggingface.co/datasets/danielheinz/telekom-backtrans-paraphrase-filtered.textfeature-extraction100K<n<1M0 likes36 downloads3y agoHugging Face09jonaskoenig /future-time-references-static-filter-D1text1M<n<10M1 likes32 downloads4y agoHugging Face10Siki-77 /snli_filter-1text100K<n<1M0 likes32 downloads3y agoHugging Face11vkarthik095 /filtered-SCOPe-2.08 Protein sequences from SCOPe-2.08 This dataset is based on section A.4.3 of [1]. For understanding the identification of remote homology relations in hidden layers, we consider the Astral SCOPe v2.08 dataset [2], containing genetic domain sequence subsets filtered to obtain < 40% pairwise sequence identity. Each domain is hierarchically classified into fold, super-family, and family. We impose an initial filter by excluding the Rossman-like folds (c.2–c.5, c.27 and 28, c.30 and 31)… See the full description on the dataset page: https://huggingface.co/datasets/vkarthik095/filtered-SCOPe-2.08.text10K<n<100K0 likes28 downloads2y agoHugging Face12huggingface-projects /DELETE-filter-bad-models0 likes27 downloads4y agoHugging Face13Telugu-LLM-Labs /telugu_teknium_GPTeacher_general_instruct_filtered_romanizedtext10K<n<100K15 likes27 downloads3y agoHugging Face14kkail8 /TAVGBench_filtered_360k_16fpstabular100K<n<1M1 likes26 downloads1y agoHugging Face15ibunescu /court_opinions_filtered_under_25ktabular1K<n<10K0 likes25 downloads3y agoHugging Face16HaileyStorm /lichess-filteredThese csv files contain a single column, 'transcript', with simple PGN string gam transcripts. The data is sources from Lichess: https://database.lichess.org All files contain some very short and very long games; I recommend filtering these, depending on your use case. All files are filtered to include white wins only. According to the Lichess ELO ratings saved with the game data, which typically run a little high, the files contain games filtered: stable.csv, 24M games: White 1300-2300, Black… See the full description on the dataset page: https://huggingface.co/datasets/HaileyStorm/lichess-filtered.text10M<n<100M0 likes24 downloads2y agoHugging Face17umesh16071973 /HRMS_FILTER_DATAtext1K<n<10K0 likes23 downloads3y agoHugging Face18juanmunoz9304HF /hotel_reviews_filteredtext100K<n<1M0 likes22 downloads1mo agoHugging Face19jonaskoenig /future-time-refernces-static-filter-D2text1M<n<10M0 likes21 downloads4y agoHugging Face20hcho22 /code_instructions_120k_alpaca_filteredtext10K<n<100K2 likes19 downloads3y agoHugging Face21gorkaartola /ZS-train_S1-AURORA01_S2-SDGtitle_Negative_Sample_Filter-AURORA01text10K<n<100K0 likes19 downloads3y agoHugging Face22rifqifarhansyah /ultrafeedback_filteredtabular1M<n<10M0 likes19 downloads1y agoHugging Face23TIGER-Lab /packages_python_filtered SWE-Next: Scalable Real-World Software Engineering Tasks for Agents packages_python_filtered This repository contains packages_python_filtered.csv, the seed repository list used by SWE-Next. The file contains 3,971 Python package / repository entries that serve as the starting point for large-scale repository mining and execution-grounded task synthesis. Each row links a package-oriented seed entry to a GitHub repository and includes lightweight… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/packages_python_filtered.tabular1K<n<10K0 likes19 downloads5mo agoHugging Face24seq-to-pheno /filtered_orthologs Zoonomia Filtered Orthologs Dataset This dataset has been filtered to remove: Proteins longer than 1000 amino acids Proteins with more than {MAX_NUMBER_ORTHOLOGS} orthologs in any species Original number of mapped orthologs: {len(all_mapped_ortholog_df)} Filtered number of mapped orthologs: {num_examples} Dataset Structure The dataset consists of a single CSV file with the following columns: transcript: Human transcript ID protein: Human protein name mapped_to:… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/filtered_orthologs.texttoken-classification1M<n<10M0 likes18 downloads2y agoHugging Face25niranjanh123 /chatgpt_filtered_sft_traces_context_awaretabularreinforcement-learningn<1K0 likes17 downloads2mo agoHugging Face26marcuscedricridia /amoralqa-filtered-polarity marcuscedricridia/amoralqa-filtered-polarity Overview This dataset is a filtered version of TheDrummer/AmoralQA-v2, created to isolate the most amoral entries based on sentiment analysis. Filtering Criteria Entries were ranked based on their negative sentiment. Only the most extreme cases were kept. The filtering process used automated sentiment analysis to determine inclusion. Source & Citation Original Dataset: TheDrummer/AmoralQA-v2… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/amoralqa-filtered-polarity.textn<1K1 likes16 downloads2y agoHugging Face27ibunescu /court_opinions_filtered_full_sizetabular1K<n<10K0 likes15 downloads3y agoHugging Face28ClarusC64 /clinical_risk_weighted_output_filter_v0.1Clinical Risk Weighted Output Filter PurposeFilter model output based on clinical risk. You receive: patient_summary risk_signals user_request You output one JSON object: risk_tierlow, medium, high, critical filtered_outputone to two sentences Scoring risk_tier_accuracy filtered_output_similarity format_pass_rate unsafe_penaltypenalizes unsafe wording in critical cases Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv texttext-classificationn<1K0 likes15 downloads7mo agoHugging Face29tussiiiii /llm-classification-distilled-v1-safe-filtered LLM Classification Distilled v1 (Safe Filtered) Overview Safer filtered distilled dataset. Uses stricter agreement and consistency conditions for higher precision. Files train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_safe_filtered.csv Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form. The dataset repository is: tussiiiii/llm-classification-distilled-v1-safe-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-safe-filtered.tabulartext-classification1K<n<10K0 likes15 downloads5mo agoHugging Face30gorkaartola /ZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-SDG_Descriptionstext1K<n<10K0 likes13 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.