CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face02notrichardren /truthfulness_all Dataset Card for "truthfulness_all" More Information needed tabular100K<n<1M0 likes1.5k downloads3y agoHugging Face03rngusry /UltraFeedback-truthfulness-preferences Dataset Card for "UltraFeedback-truthfulness-preferences" More Information needed tabular100K<n<1M1 likes1k downloads2y agoHugging Face04Cameronk199 /donald-trump-truth-social-posts Donald Trump Truth Social Posts Archive Archive overview 36,170 public Truth Social posts associated with Donald J. Trump's @realDonaldTrump account. The release preserves source URLs, timestamps, post types, original HTML, extracted plain text, attachment provenance, and analysis-ready tables. It also includes streamable image media plus video metadata and transcripts where the source provides them. The package is source-linked and reconciled by archive ID.… See the full description on the dataset page: https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts.imagetext-generation100K<n<1M3 likes879 downloads12d agoHugging Face05notrichardren /truthfulness_explanation Dataset Card for "truthfulness_explanation" More Information needed tabular10K<n<100K0 likes660 downloads3y agoHugging Face06ContextualAI /ultrabin_clean_max_chosen_min_rejected_rationalized_truthfulnesstabular10K<n<100K0 likes554 downloads2y agoHugging Face07notrichardren /truthfulness_legacytabular100K<n<1M3 likes452 downloads3y agoHugging Face08Jennny /ultrafeedback_binarized_truthfulness_prefstabular10K<n<100K0 likes377 downloads2y agoHugging Face09nyu-dice-lab /lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-vicgalle-CarbonBeagle-11B-truthy-private.tabular100K<n<1M0 likes361 downloads2y agoHugging Face10truthinpolling /tip-civic-dataNotice — 26 September 2026: please ignore the status column in the federal bills file. That column was not published by Congress. It was filled in by our own system — most rows simply say "active" — and it can be wrong: a bill the President has signed may still read "active". We have removed it from our site and our data services, and it will be removed from this file in a future release. For where a bill stands, use latest_action_text and latest_action_date. Those are Congress.gov's own words… See the full description on the dataset page: https://huggingface.co/datasets/truthinpolling/tip-civic-data.tabulartext-classification100K<n<1M1 likes340 downloads5h agoHugging Face11truthful-ai /story-imprinting Story Imprinting — training datasets Datasets accompanying Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble. Paper · Code Contents Paper section Folder Data 3.1 — Sabotage 3_1_sabotage/ Three training mixtures and separate sabotage/clean story pools 3.2 — Narration preferences 3_2_narration_preferences/ Six training mixtures and 12 story pools 4 — Affinity 4_selectivity/ Opposing-pair training datasets and raw… See the full description on the dataset page: https://huggingface.co/datasets/truthful-ai/story-imprinting.tabulartext-generation100K<n<1M0 likes321 downloads8d agoHugging Face12ground-truth /multichannel-meetings-10h GroundTruth Multi-Channel Meeting Audio Dataset (10h) Summary This dataset contains approximately 10 hours of co-located, multi-speaker meeting recordings, each captured simultaneously via a room (built-in) microphone and individual close-talk lapel microphones worn by each participant. Each meeting includes: One full meeting recording (room microphone) Individual close-talk recordings for each participant (one file per speaker) Structured metadata describing speakers… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/multichannel-meetings-10h.audioautomatic-speech-recognitionn<1K1 likes287 downloads5mo agoHugging Face13nyu-dice-lab /lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private Dataset Card for Evaluation run of yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B Dataset automatically created during the evaluation run of model yunconglong/Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-yunconglong-Truthful_DPO_TomGrc_FusionNet_7Bx2_MoE_13B-private.tabular100K<n<1M0 likes218 downloads2y agoHugging Face14tiagozip /truthsocial Truth Social scrape hundreds of millions of public posts from over 500k Truth Social accounts since February 2022, scraped continuously since early 2026 and updated monthly. from datasets import load_dataset posts = load_dataset("tiagozip/truthsocial", "posts", split="train", streaming=True) accounts = load_dataset("tiagozip/truthsocial", "accounts", split="train") format config rows notes posts 100M+ one file per month, data/posts/YYYY-MM.parquet… See the full description on the dataset page: https://huggingface.co/datasets/tiagozip/truthsocial.imagetext-classification10M<n<100M1 likes188 downloads12d agoHugging Face15carlomarxx /trilemma-of-truth Dataset Card for Trilemma of Truth (ToT) Dataset 🧾 Dataset Summary The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types: Factually true statements. Factually false statements. Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below). The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.texttext-classification10K<n<100K2 likes185 downloads2mo agoHugging Face16chrissoria /trump-truth-social Trump Truth Social Posts Archive Public posts ("Truths") by Donald J. Trump on Truth Social, enriched with market data, geopolitical event indicators, and LLM-based post classifications. Collected for academic research purposes. Fields Post metadata Field Type Description date string Post date (YYYY-MM-DD) time string Post time in UTC (HH:MM:SS) time_eastern string Post time in US Eastern (HH:MM:SS, DST-aware) day_of_week string Day name… See the full description on the dataset page: https://huggingface.co/datasets/chrissoria/trump-truth-social.imagetext-classification10K<n<100K4 likes149 downloads4mo agoHugging Face17open-llm-leaderboard /vicgalle__CarbonBeagle-11B-truthy-detailsgated Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/vicgalle__CarbonBeagle-11B-truthy-details.tabular10K<n<100K0 likes139 downloads2y agoHugging Face18NinaCalvi /ultra-rm-truthfulness-1000tabular1K<n<10K0 likes109 downloads2y agoHugging Face19nyu-dice-lab /lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private Dataset Card for Evaluation run of nbeerbower/bophades-mistral-truthy-DPO-7B Dataset automatically created during the evaluation run of model nbeerbower/bophades-mistral-truthy-DPO-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-bophades-mistral-truthy-DPO-7B-private.tabular100K<n<1M0 likes98 downloads2y agoHugging Face20vab46 /Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft Dataset details:- This dataset is the 2nd iteration following bugs in 1st dataset. The initial data suffered with followoing cases:- (i) The failed reference_answers generation(due error totalling 23) primarly because of 2 reasons/exceptions:-  (a) There was normal limit(300 in 1st request) and worst case limit(450 in 3rd request) number of tokens for consolidated 4 refernce_answers per chunk and its 4 corresponding answers. However certain answers breached this higher… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA-junk_handled_ft.tabular1K<n<10K0 likes92 downloads10d agoHugging Face21vab46 /Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft Dataset details:- This dataset is basically mapping of final anchor-positive pair data with their refernce answer. The given input data considered because:- (i) it had the had purest anchor-positive pairs with semantically bound anchors with context/positive. (ii) gave us the best result on final embedding fine tuning model. The anchor-context(positive)-reference_answer data has been generated via Qwen-2.5-7B teacher model with temperature 0.1 and a strict system prompt.… See the full description on the dataset page: https://huggingface.co/datasets/vab46/Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft.tabular1K<n<10K0 likes80 downloads10d agoHugging Face22mlfoundations-dev /multiple_samples_ground_truth_openr1_llm_verifier_cleantabular100K<n<1M0 likes76 downloads2y agoHugging Face23mlfoundations-dev /multiple_samples_ground_truth_openr1_llm_verifiertabular100K<n<1M0 likes73 downloads2y agoHugging Face24nyu-dice-lab /lm-eval-results-nbeerbower-slerp-bophades-truthy-math-mistral-7B-private Dataset Card for Evaluation run of nbeerbower/slerp-bophades-truthy-math-mistral-7B Dataset automatically created during the evaluation run of model nbeerbower/slerp-bophades-truthy-math-mistral-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-nbeerbower-slerp-bophades-truthy-math-mistral-7B-private.tabular100K<n<1M0 likes65 downloads2y agoHugging Face25skymizer /llm-ground-truth-reasoningtabular1K<n<10K0 likes64 downloads7mo agoHugging Face26ceselder /cot-oracle-truthfulqa-hint-admission-unverbalized TruthfulQA Hint Admission — Unverbalized Eval dataset for the CoT Oracle project. Tests whether an activation oracle can detect hint influence from model internals when the model does not verbalize the hint in its chain-of-thought. What is this? Qwen3-8B is given TruthfulQA multiple-choice questions with planted hints (correct or wrong). This dataset contains only the rollouts where the model did not mention the hint in its reasoning — the oracle must read… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/cot-oracle-truthfulqa-hint-admission-unverbalized.tabular10K<n<100K0 likes61 downloads7mo agoHugging Face27NinaCalvi /ultra-50k-samples-dataset-truthfulnesstabular10K<n<100K0 likes53 downloads2y agoHugging Face28mlabonne /distilabel-truthy-dpo-v0.1-filteredtabularn<1K2 likes49 downloads3y agoHugging Face29OALL /details_vicgalle__CarbonBeagle-11B-truthy Dataset Card for Evaluation run of vicgalle/CarbonBeagle-11B-truthy Dataset automatically created during the evaluation run of model vicgalle/CarbonBeagle-11B-truthy. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_vicgalle__CarbonBeagle-11B-truthy.tabular100K<n<1M0 likes42 downloads2y agoHugging Face30skymizer /llm-ground-truth-generaltabular1K<n<10K0 likes39 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.