CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kuzheren /geometry-dash-retro-levelsFork of https://huggingface.co/datasets/yusp48/geometry-dash-levels. Contains only retro levels with id < 11000000. Use my gdparse library: pip install gdparse tabular10K<n<100K0 likes278 downloads1y agoHugging Face02dougalldeepmind /2026-08-27-odcv-post-action-retrospection-716-seed-2-eval ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 2, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-2-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-2-eval.text10K<n<100K0 likes218 downloads22d agoHugging Face03dougalldeepmind /2026-08-28-post-action-retrospection-716-coherent Post-action retrospection 716 -- coherent rewrite (arm 1 of the PAR coherence experiment) field value experiment The exact 716 five-turn PAR rows that trained LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch (mixture 2026-08-26-table2-9284-par716-train @ 42c8a74), with ONLY the trained turn (turn 4: private reasoning + reply) rewritten by Sonnet 5 so the reasoning ENDS on a first-person decision (what it won't do, per… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-28-post-action-retrospection-716-coherent.textn<1K0 likes176 downloads22d agoHugging Face04RetroJenkins /sigil-forge-training SIGIL Forge Training Data Forge-verified training tasks, references, fixtures, and versioned MLX SFT corpora for SIGIL. The SIGIL source repository pins immutable revisions and verifies MANIFEST.json plus every payload. Evaluation tasks and validation records are intentionally stored in a separate private repository. texttext-generation0 likes139 downloads2mo agoHugging Face05dougalldeepmind /2026-08-27-odcv-post-action-retrospection-716-eval ODCV-Bench: post-action-retrospection (design B) 716 arm, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-26-qwen36-lora-table2-9284-post-action-retrospection-716-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on these 65 cells:… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-eval.text10K<n<100K0 likes120 downloads22d agoHugging Face06RetrO21 /AgriFinetext10K<n<100K0 likes118 downloads10mo agoHugging Face07badigadiii /retro-games-gameplay-framesimage10K<n<100K0 likes115 downloads1y agoHugging Face08martintomov /retrofuturism-fluximagen<1K0 likes108 downloads2y agoHugging Face09retroam /repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes104 downloads2mo agoHugging Face10retrogradespace /hmda_2024 HMDA 2024 (Home Mortgage Disclosure Act) Full-year 2024 loan application register (LAR) data released under the Home Mortgage Disclosure Act (HMDA), re-published here as a single Parquet file for convenient loading with the datasets library. Dataset summary Rows: 12,229,298 loan application records Columns: 99 (the full public LAR field set — property, applicant, underwriting, and pricing information) Format: Parquet (hmda_2024.parquet) Source: Consumer Financial… See the full description on the dataset page: https://huggingface.co/datasets/retrogradespace/hmda_2024.texttabular-classification10M<n<100M1 likes98 downloads2mo agoHugging Face11dougalldeepmind /2026-08-27-odcv-post-action-retrospection-716-seed-1-eval ODCV-Bench: post-action-retrospection (design B) 716 arm, seed 1, 2 rollouts x 65 cells field value experiment ODCV-Bench rollouts and judge scores for LASR-Callum/2026-08-27-qwen36-lora-table2-9284-post-action-retrospection-716-seed-1-rank-64-dynbatch: the da716 organism whose 716 rows are five-turn post-action-retrospection records (a difficult-advice prompt, a bare refusal, pushback, then the reasoning the refusal skipped; only the last turn trained). Headline on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-27-odcv-post-action-retrospection-716-seed-1-eval.text10K<n<100K0 likes94 downloads22d agoHugging Face12OpenDFM /RetroDFM-R-inferencetext1M<n<10M0 likes92 downloads1mo agoHugging Face13dougalldeepmind /2026-08-26-sonnet45-post-action-retrospection-natural-turn-design synth post_action_retrospection run — per-stage snapshots (resumable generation cache) field value experiment synth post_action_retrospection run — per-stage snapshots (resumable generation cache) date_generated 20260826_152715 constitution constitutions/claude_distilled_12_principles_mid/constitution.md source_repo https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT.git @ c2fdee460e71fa28e9902edf1cc662db0d19cad8 models per-stage models — see… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-26-sonnet45-post-action-retrospection-natural-turn-design.tabularn<1K0 likes74 downloads22d agoHugging Face14jdpressman /retro-text-style-transfer-v0.1 Retro Textual Style Transfer v0.1 This component of RetroInstruct implements textual style transfer by providing a dataset of language model instruction prompts that take an example style passage along with a task text and rewrite the task text to sound like the style passage It is made by starting with ground truth public domain text from the pg19 dataset and then writing task passages to "transfer from" with Mixtral Instruct. It is similar in spirit to the "instruction… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-text-style-transfer-v0.1.text10K<n<100K9 likes53 downloads3y agoHugging Face15QizhiPei /e3fp-mol-instructions-retrosynthesis 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text Modeling For more information, please refer to our paper and GitHub repository. Paper: arxiv, openreview GitHub: 3D-MolT5 Authors: Qizhi Pei, Rui Yan, Kaiyuan Gao, Jinhua Zhu and Lijun Wu text100K<n<1M0 likes53 downloads1y agoHugging Face16RetrO21 /Agriiimage10K<n<100K0 likes52 downloads10mo agoHugging Face17kyLELEng /adaptive-retro-gpt-1b-corpus Adaptive-RETRO-GPT-1B Pretraining Corpus Cleaned causal language modeling corpus for the Adaptive-RETRO-GPT-1B run. Source: HuggingFaceFW/fineweb-edu / sample-10BT Train rows: 80000 Validation rows: 4000 Format: JSONL with text and source text10K<n<100K0 likes43 downloads5mo agoHugging Face18jdpressman /retro-ascii-art-v1 RetroInstruct ASCII Art This component of RetroInstruct trains language models to draw ASCII art. Many advanced language models such as Microsoft Prometheus (Bing) and Claude 3 Opus can draw impressive ASCII diagrams. Mistral-large on the other hand can't. Since there should in principle be plenty of ASCII art in Common Crawl I suspect this is caused by either Mistral's filters removing ASCII art from the pretraining or instruction tuning data that doesn't reinforce the ability to… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-ascii-art-v1.text1K<n<10K12 likes37 downloads2y agoHugging Face19smitathkr1 /ord-retro-datasettext1M<n<10M0 likes36 downloads11mo agoHugging Face20KU-AGI /RetroReasoner-datatextn<1K0 likes36 downloads22d agoHugging Face21Dans-DiscountModels /Retro-YahooAnswers Description This dataset is an instruct style dataset comprised of a scrape of the Yahoo! Answers website that was done in 2007. The dataset is comprised of 10 categories labeled 1-10. The categories are as follows: Society & Culture Science & Mathematics Health Education & Reference Computers & Internet Sports Business & Finance Entertainment & Music Family & Relationships Politics & Government The subject line and body of the question have been combined into a single field and… See the full description on the dataset page: https://huggingface.co/datasets/Dans-DiscountModels/Retro-YahooAnswers.textquestion-answering1M<n<10M4 likes35 downloads3y agoHugging Face22durinn /Durinn_Hacktoberfest_Retrospective Durinn Hacktoberfest Retrospective Dataset Author: Ryan Marinelli & Victor Strandmoe Project: Durinn — Scaling Vibe Coding AuditingDataset Type: Security SFT (Supervised Fine-Tuning)Sources: Scanning GitHub Hacktoberfest 2025 Format: HuggingFace DatasetDict with train and validation splits 📌 Overview This dataset provides security-focused training data derived from analyzing Hacktoberfest 2025 GitHub repositories before and after the event using Semgrep’s OWASP Top 10… See the full description on the dataset page: https://huggingface.co/datasets/durinn/Durinn_Hacktoberfest_Retrospective.textn<1K1 likes30 downloads10mo agoHugging Face23jdpressman /retro-weave-eval-rubrics-v0.1 RetroInstruct Weave Evaluator Rubrics v0.1 This component of RetroInstruct trains the ability to break subjective weave rubric items like "Is this good writing?" into parts which can be more objectively answered. It is closery related to the word parts component which is meant to train a similar skill. By making these rubrics the model gains the ability to make in-context text classifiers and discriminators. These can be used to drive a MCTS, filter language model outputs to… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-rubrics-v0.1.text1K<n<10K1 likes28 downloads2y agoHugging Face24Retrogradi /gender_stereoset_rephrasedtext1K<n<10K0 likes24 downloads2mo agoHugging Face25jdpressman /retro-weave-eval-jdp-v0.1 RetroInstruct Weave Evaluator Questions: JDP This component of RetroInstruct trains the ability to answer yes-no questions such as "Does the current scene take place at a wedding party?". The logits from such questions can be taken to make in-context text classifiers and discriminators. These can be used to drive a MCTS, filter language model outputs to heighten the probability they satisfy certain properties, and validate abstract properties of inputs. This set of questions is made… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-weave-eval-jdp-v0.1.textn<1K1 likes23 downloads2y agoHugging Face26jdpressman /retro-word-parts-v0.1 RetroInstruct Part Lists For Dictionary Words v0.1 This component of RetroInstruct distills Mixtral Instruct's ontology by having it describe the uniquely identifying parts of the concepts or objects referred to by dictionary words. This is useful both as factual knowledge but also to train an instruction model to perform the basic mental motions of breaking concepts down into pieces and synthesizing ideas from pieces, textual object decomposition and recognition. Each row in this… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-word-parts-v0.1.text10K<n<100K4 likes22 downloads3y agoHugging Face27jdpressman /retroinstruct-mix-v0.2 RetroInstruct Mix v0.2 This is the first release of the RetroInstruct synthetic instruction dataset. It is a mixture of 7 synthetic subsets: RetroInstruct Weave Evaluator Questions: JDP - Answer questions about synthetic short form writing in the style of John David Pressman. RetroInstruct Analogical Translations - Infer the generative process of bad faith reasoning by executing a bad faith process to generate arguments and reversing it. RetroInstruct Part Lists For Dictionary… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retroinstruct-mix-v0.2.text10K<n<100K1 likes20 downloads2y agoHugging Face28open-llm-leaderboard /DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-detailsgated Dataset Card for Evaluation run of DreadPoor/Mercury_In_Retrograde-8b-Model-Stock Dataset automatically created during the evaluation run of model DreadPoor/Mercury_In_Retrograde-8b-Model-Stock The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DreadPoor__Mercury_In_Retrograde-8b-Model-Stock-details.tabular10K<n<100K0 likes18 downloads2y agoHugging Face29jdpressman /retroinstruct-agent-mix-v0.4text10K<n<100K0 likes17 downloads1y agoHugging Face30jdpressman /retro-easy-prose-repair-diffs-v0.1 RetroInstruct Easy Prose Repair Diffs This component of RetroInstruct trains language models to repair prose by outputting a diff that patches its flaws. The dataset is made through backtranslation by running a synthetic corruption pass over prose. I use mostly syntactic corruptions made with traditional programs, which makes them 'easy' compared to more subtle semantic problems that could be introduced by a neural network. The text I backtranslate from was generated by Mixtral… See the full description on the dataset page: https://huggingface.co/datasets/jdpressman/retro-easy-prose-repair-diffs-v0.1.text1K<n<10K1 likes16 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.