datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jailbreak-filter-outputstelugu_alpaca_yahma_cleaned_filtered_romanizedemail-spam-filterJParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.shelfglance-ucp-category-filter
Shopify's agent-commerce category filter didn't filter on any of 190 stores
Every Shopify store answers an agent-commerce endpoint whose schema declares a category filter. Sent to 200 live stores with a price-filter control: 186 ignored it, 4 rejected every value, 0 filtered.
One row per Shopify storefront that answered its own agent-commerce endpoint (POST /api/ucp/mcp) on 2 September 2026: 190 of 200 sampled from a corpus of 10,099. Each row records what search_catalog… See the full description on the dataset page: https://huggingface.co/datasets/bpm2007/shelfglance-ucp-category-filter.HelpSteer-filtered
HelpSteer-filtered
This dataset is a highly filtered version of the nvidia/HelpSteer dataset.
❓ How this dataset was filtered:
I calculated the sum of the columns ["helpfulness," "correctness," "coherence," "complexity," "verbosity"] and created a new column named sum.
I changed some column names and added a empty column to match the Alpaca format.
The dataset was then filtered to include only those entries with a sum greater than or equal to 16.
🧐 More… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/HelpSteer-filtered.filter-bad-modelstelekom-backtrans-paraphrase-filteredThis is a filtered version of Philip May's German paraphrase dataset.
The dataset has been filtered for the sake of convenience, since smaller devices do not support such large files.
All text pairs in the dataset are paraphrases, and are therefore labelled 1. As such, the dataset is well-suited for use in conjunction with the multiple negatives ranking loss.
As the original author suggests, the dataset has been filtered, mostly following the guidelines set by the author. Any row that doesn't… See the full description on the dataset page: https://huggingface.co/datasets/danielheinz/telekom-backtrans-paraphrase-filtered.future-time-references-static-filter-D1snli_filter-1filtered-SCOPe-2.08
Protein sequences from SCOPe-2.08
This dataset is based on section A.4.3 of [1]. For understanding the identification of remote homology relations in hidden layers, we consider the Astral SCOPe v2.08 dataset [2], containing genetic domain sequence subsets filtered to obtain < 40% pairwise sequence identity. Each domain is hierarchically classified into fold, super-family, and family. We impose an initial filter by excluding the Rossman-like folds (c.2–c.5, c.27 and 28, c.30 and 31)… See the full description on the dataset page: https://huggingface.co/datasets/vkarthik095/filtered-SCOPe-2.08.DELETE-filter-bad-modelstelugu_teknium_GPTeacher_general_instruct_filtered_romanizedTAVGBench_filtered_360k_16fpscourt_opinions_filtered_under_25klichess-filteredThese csv files contain a single column, 'transcript', with simple PGN string gam transcripts. The data is sources from Lichess: https://database.lichess.org
All files contain some very short and very long games; I recommend filtering these, depending on your use case.
All files are filtered to include white wins only.
According to the Lichess ELO ratings saved with the game data, which typically run a little high, the files contain games filtered:
stable.csv, 24M games: White 1300-2300, Black… See the full description on the dataset page: https://huggingface.co/datasets/HaileyStorm/lichess-filtered.HRMS_FILTER_DATAhotel_reviews_filteredfuture-time-refernces-static-filter-D2code_instructions_120k_alpaca_filteredZS-train_S1-AURORA01_S2-SDGtitle_Negative_Sample_Filter-AURORA01ultrafeedback_filteredpackages_python_filtered
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
packages_python_filtered
This repository contains packages_python_filtered.csv, the seed repository list used by SWE-Next. The file contains 3,971 Python package / repository entries that serve as the starting point for large-scale repository mining and execution-grounded task synthesis.
Each row links a package-oriented seed entry to a GitHub repository and includes lightweight… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/packages_python_filtered.filtered_orthologs
Zoonomia Filtered Orthologs Dataset
This dataset has been filtered to remove:
Proteins longer than 1000 amino acids
Proteins with more than {MAX_NUMBER_ORTHOLOGS} orthologs in any species
Original number of mapped orthologs: {len(all_mapped_ortholog_df)}
Filtered number of mapped orthologs: {num_examples}
Dataset Structure
The dataset consists of a single CSV file with the following columns:
transcript: Human transcript ID
protein: Human protein name
mapped_to:… See the full description on the dataset page: https://huggingface.co/datasets/seq-to-pheno/filtered_orthologs.chatgpt_filtered_sft_traces_context_awareamoralqa-filtered-polarity
marcuscedricridia/amoralqa-filtered-polarity
Overview
This dataset is a filtered version of TheDrummer/AmoralQA-v2, created to isolate the most amoral entries based on sentiment analysis.
Filtering Criteria
Entries were ranked based on their negative sentiment.
Only the most extreme cases were kept.
The filtering process used automated sentiment analysis to determine inclusion.
Source & Citation
Original Dataset: TheDrummer/AmoralQA-v2… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/amoralqa-filtered-polarity.court_opinions_filtered_full_sizeclinical_risk_weighted_output_filter_v0.1Clinical Risk Weighted Output Filter
PurposeFilter model output based on clinical risk.
You receive:
patient_summary
risk_signals
user_request
You output one JSON object:
risk_tierlow, medium, high, critical
filtered_outputone to two sentences
Scoring
risk_tier_accuracy
filtered_output_similarity
format_pass_rate
unsafe_penaltypenalizes unsafe wording in critical cases
Run scoringpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
llm-classification-distilled-v1-safe-filtered
LLM Classification Distilled v1 (Safe Filtered)
Overview
Safer filtered distilled dataset. Uses stricter agreement and consistency conditions for higher precision.
Files
train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_safe_filtered.csv
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form.
The dataset repository is: tussiiiii/llm-classification-distilled-v1-safe-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-safe-filtered.ZS-train_SDG_Descriptions_S1-sentence_S2-SDGtitle_Negative_Sample_Filter-SDG_Descriptions
