CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ibunescu /court_opinions_filtered_under_25ktabular1K<n<10K0 likes51 downloads3y agoHugging Face02Verah /JParaCrawl-Filtered-English-Japanese-Parallel-Corpus Introduction This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus. The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet. Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.tabulartranslation1M<n<10M3 likes50 downloads3y agoHugging Face03Weyaxi /HelpSteer-filtered HelpSteer-filtered This dataset is a highly filtered version of the nvidia/HelpSteer dataset. ❓ How this dataset was filtered: I calculated the sum of the columns ["helpfulness," "correctness," "coherence," "complexity," "verbosity"] and created a new column named sum. I changed some column names and added a empty column to match the Alpaca format. The dataset was then filtered to include only those entries with a sum greater than or equal to 16. 🧐 More… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/HelpSteer-filtered.tabular1K<n<10K4 likes43 downloads3y agoHugging Face04kkail8 /TAVGBench_filtered_360k_16fpstabular100K<n<1M1 likes26 downloads1y agoHugging Face05ibunescu /court_opinions_filtered_full_sizetabular1K<n<10K0 likes18 downloads3y agoHugging Face06rifqifarhansyah /ultrafeedback_filteredtabular1M<n<10M0 likes17 downloads1y agoHugging Face07TIGER-Lab /packages_python_filtered SWE-Next: Scalable Real-World Software Engineering Tasks for Agents packages_python_filtered This repository contains packages_python_filtered.csv, the seed repository list used by SWE-Next. The file contains 3,971 Python package / repository entries that serve as the starting point for large-scale repository mining and execution-grounded task synthesis. Each row links a package-oriented seed entry to a GitHub repository and includes lightweight… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/packages_python_filtered.tabular1K<n<10K0 likes17 downloads5mo agoHugging Face08niranjanh123 /chatgpt_filtered_sft_traces_context_awaretabularreinforcement-learningn<1K0 likes16 downloads2mo agoHugging Face09tussiiiii /llm-classification-distilled-v1-safe-filtered LLM Classification Distilled v1 (Safe Filtered) Overview Safer filtered distilled dataset. Uses stricter agreement and consistency conditions for higher precision. Files train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_safe_filtered.csv Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form. The dataset repository is: tussiiiii/llm-classification-distilled-v1-safe-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-safe-filtered.tabulartext-classification1K<n<10K0 likes15 downloads5mo agoHugging Face10liechticonsulting /NER_FILTERED_DATASETtabular10K<n<100K0 likes11 downloads11mo agoHugging Face11niranjanh123 /gemini_filtered_sft_traces_simplified_reasoningtabularreinforcement-learningn<1K0 likes11 downloads2mo agoHugging Face12niranjanh123 /sonnet_filtered_sft_traces_simplified_reasoningtabularreinforcement-learningn<1K0 likes11 downloads2mo agoHugging Face13tussiiiii /llm-classification-distilled-v1-filtered LLM Classification Distilled v1 (Filtered) Overview Filtered distilled training dataset. Built from rows where target agrees with gold and basic quality conditions are met. Files train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_filtered.csv Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form. The dataset repository is: tussiiiii/llm-classification-distilled-v1-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-filtered.tabulartext-classification1K<n<10K0 likes10 downloads5mo agoHugging Face14tussiiiii /llm-classification-distilled-v2-filtered LLM Classification Distilled v2 (Filtered) Overview Filtered distilled training dataset merged from sharded v2 distillation outputs. Files train.csv: merged dataset file Shard Source Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded Number of shards: 4 Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form: A, B, C where C means tie. This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-filtered.tabulartext-classification10K<n<100K0 likes10 downloads4mo agoHugging Face15niranjanh123 /chatgpt_filtered_sft_traces_simplified_reasoningtabularreinforcement-learningn<1K0 likes10 downloads2mo agoHugging Face16ada-datadruids /movies_based_on_books_filteredtabularn<1K0 likes9 downloads2y agoHugging Face17joackimagno /filtered-filipino-recipestabularn<1K0 likes8 downloads2y agoHugging Face18h1alexbel /sr-filteredtabular1K<n<10K1 likes7 downloads2y agoHugging Face19tussiiiii /llm-classification-distilled-v2-safe-filtered LLM Classification Distilled v2 (Safe Filtered) Overview Safe filtered distilled training dataset merged from sharded v2 distillation outputs. Files train.csv: merged dataset file Shard Source Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded Number of shards: 4 Notes This dataset was generated by a teacher-judge distillation pipeline. Labels are stored in ABC form: A, B, C where C means tie. This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-safe-filtered.tabulartext-classification10K<n<100K0 likes7 downloads4mo agoHugging Face20beloiual /edges_out_filteredtabularn<1K0 likes5 downloads2y agoHugging Face21ada-datadruids /movies_not_based_on_books_filteredtabular100K<n<1M0 likes5 downloads2y agoHugging Face22prithvi3 /filtered_forecast_sample_testtabularn<1K0 likes5 downloads2y agoHugging Face23ridhomhd /MedMCQA-filteredtabular1K<n<10K0 likes5 downloads2y agoHugging Face24prithvi3 /filtered_consistency_sample_testtabularn<1K0 likes5 downloads2y agoHugging Face25kyrgyz-ai /emolia_filtered_v1_bb_featurestabular10K<n<100K0 likes5 downloads4mo agoHugging Face26roderickwen /math_problems_filteredtabular100K<n<1M0 likes4 downloads2y agoHugging Face27prithvi3 /forecast_sample_test_filtered_earlyrestabularn<1K0 likes3 downloads2y agoHugging Face28rishieee /ORAN_TeleQNA_filteredgated Overview ORAN_TeleQNA is a streamlined evaluation dataset derived from ORANBench and TeleQNA tabularn<1K0 likes3 downloads6mo agoHugging Face29Pamzyy /Filtered_From_Ayagatedtabular10K<n<100K0 likes2 downloads2y agoHugging Face30narunraman /open_llm_filteredtabular10K<n<100K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.