datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deepstock-stock-historical-prices-dataset-processeddatacookers-processed
Preprocessed Data for the ADA Project 2025
By DataCookers
General Datasets
channels_clean.parquet: A cleaned version of the df_channels_en split from the Youniverse Dataset. This dataset corrects parsing errors where missing values in the Category column caused subsequent columns to shift one position to the left.
grammy_raw.parquet: The foundational dataset containing Grammy winners and nominees (1965–2024), sourced from Kaggle.
grammy_metadata.parquet: A… See the full description on the dataset page: https://huggingface.co/datasets/matcav/datacookers-processed.deepstock-stock-historical-prices-dataset-processed-with-previous-day-titlesyoutube_processed_full_dataset_final3d-pinn-dim55-processed_datasetyoutube_processed_datasetprocessed_empathy_datasetdemo-restored-compliance-data-processedprocessed_tldr_comparison_dataset_20251102_065554
TL;DR Comparison Dataset for OpenAI's Summarize from Feedback task
The dataset is generated from https://huggingface.co/datasets/openai/summarize_from_feedback.
Please refer to https://github.com/liyuan24/dgx_spark_summary_from_human_feedback/tree/main?tab=readme-ov-file#download-the-comparison-dataset about how to download the dataset.
This is a comparison dataset used for training a reward model. Each example contains a query (post) and two responses (chosen and rejected) where… See the full description on the dataset page: https://huggingface.co/datasets/seangogo/processed_tldr_comparison_dataset_20251102_065554.youtube-transcript-dataset-processedprocessed_tldr_sft_dataset_20251029_035657Paul_RNA_Sequence_Processed_Datasetagentica-org_deepscaler-preview-dataset-simple-processed元データセット
https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset
processed_tldr_sft_dataset_20251029_045736_with_rewardsdata-processedfava-data-processed
FAVA Dataset (Processed)
Dataset Description
Dataset Summary
The FAVA (Factual Association and Verification Annotations) dataset is designed for evaluating hallucinations in language model outputs. This processed version contains binary hallucination labels derived from detailed span-level annotations in the original dataset.
Dataset Structure
Each example contains:
Required columns:
query: The prompt given to the model
context: Empty field (for… See the full description on the dataset page: https://huggingface.co/datasets/wandb/fava-data-processed.deepstock-stock-historical-prices-dataset-processedmath_level3to5_data_processed_with_qwen_prompt_dedupprocessed_dataset_whisper_endeepstock-stock-historical-prices-dataset-processedtourism-processed-datasetmath_level3to5_data_processed_with_qwen_prompt_dedup_cleanprocessed_tldr_sft_dataset_20251028_232434engine-maintenance-dataset-processedDataset-TG-HS-HX-Processeddeepstock-stock-historical-prices-dataset-processeddeepstock-dataset-processedprocessed_tldr_sft_dataset_20251029_044328
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is generated from https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
These columns are added by this preprocessing script:
query: length-limited query for… See the full description on the dataset page: https://huggingface.co/datasets/seangogo/processed_tldr_sft_dataset_20251029_044328.processed_tldr_sft_dataset_20251029_045736
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is generated from https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
These columns are added by this preprocessing script:
query: length-limited query for… See the full description on the dataset page: https://huggingface.co/datasets/seangogo/processed_tldr_sft_dataset_20251029_045736.medium_processed_dataset_by_post
