datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
court_opinions_filtered_under_25kJParaCrawl-Filtered-English-Japanese-Parallel-Corpus
Introduction
This is a LLM-filtered set of the first 1M rows from ntt's JParaCrawl v3 large English-Japanese parallel corpus.
The original JParaCrawl corpus was put together by automated means - aligning Japanese texts with their apparent English translations that were found in-the-wild, on the internet.
Whilst manually browsing the original data, I noticed that there were obvious quality issues that made me anxious about using the dataset at all. Poorly aligned translations… See the full description on the dataset page: https://huggingface.co/datasets/Verah/JParaCrawl-Filtered-English-Japanese-Parallel-Corpus.HelpSteer-filtered
HelpSteer-filtered
This dataset is a highly filtered version of the nvidia/HelpSteer dataset.
❓ How this dataset was filtered:
I calculated the sum of the columns ["helpfulness," "correctness," "coherence," "complexity," "verbosity"] and created a new column named sum.
I changed some column names and added a empty column to match the Alpaca format.
The dataset was then filtered to include only those entries with a sum greater than or equal to 16.
🧐 More… See the full description on the dataset page: https://huggingface.co/datasets/Weyaxi/HelpSteer-filtered.TAVGBench_filtered_360k_16fpscourt_opinions_filtered_full_sizeultrafeedback_filteredpackages_python_filtered
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
packages_python_filtered
This repository contains packages_python_filtered.csv, the seed repository list used by SWE-Next. The file contains 3,971 Python package / repository entries that serve as the starting point for large-scale repository mining and execution-grounded task synthesis.
Each row links a package-oriented seed entry to a GitHub repository and includes lightweight… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/packages_python_filtered.chatgpt_filtered_sft_traces_context_awarellm-classification-distilled-v1-safe-filtered
LLM Classification Distilled v1 (Safe Filtered)
Overview
Safer filtered distilled dataset. Uses stricter agreement and consistency conditions for higher precision.
Files
train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_safe_filtered.csv
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form.
The dataset repository is: tussiiiii/llm-classification-distilled-v1-safe-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-safe-filtered.NER_FILTERED_DATASETgemini_filtered_sft_traces_simplified_reasoningsonnet_filtered_sft_traces_simplified_reasoningllm-classification-distilled-v1-filtered
LLM Classification Distilled v1 (Filtered)
Overview
Filtered distilled training dataset. Built from rows where target agrees with gold and basic quality conditions are met.
Files
train.csv: main dataset file uploaded from train_distilled_qwen32b_awq_v1_filtered.csv
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form.
The dataset repository is: tussiiiii/llm-classification-distilled-v1-filtered… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v1-filtered.llm-classification-distilled-v2-filtered
LLM Classification Distilled v2 (Filtered)
Overview
Filtered distilled training dataset merged from sharded v2 distillation outputs.
Files
train.csv: merged dataset file
Shard Source
Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded
Number of shards: 4
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form: A, B, C where C means tie.
This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-filtered.chatgpt_filtered_sft_traces_simplified_reasoningmovies_based_on_books_filteredfiltered-filipino-recipessr-filteredllm-classification-distilled-v2-safe-filtered
LLM Classification Distilled v2 (Safe Filtered)
Overview
Safe filtered distilled training dataset merged from sharded v2 distillation outputs.
Files
train.csv: merged dataset file
Shard Source
Source shard repo: tussiiiii/llm-classification-distilled-v2-sharded
Number of shards: 4
Notes
This dataset was generated by a teacher-judge distillation pipeline.
Labels are stored in ABC form: A, B, C where C means tie.
This repository is:… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/llm-classification-distilled-v2-safe-filtered.edges_out_filteredmovies_not_based_on_books_filteredfiltered_forecast_sample_testMedMCQA-filteredfiltered_consistency_sample_testemolia_filtered_v1_bb_featuresmath_problems_filteredforecast_sample_test_filtered_earlyresORAN_TeleQNA_filtered
Overview
ORAN_TeleQNA is a streamlined evaluation dataset derived from ORANBench and TeleQNA
Filtered_From_Ayaopen_llm_filtered
