datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
THEMol
THEMol: Torsion, Hessian, Energy of Molecules
Dataset Summary
THEMol is an open-source collection of quantum mechanical properties tailored for organic molecules. It provides large-scale density functional theory (DFT) data for exploring intramolecular potential energy surfaces, including optimized geometries, structural relaxation trajectories, torsion scans, constrained torsion relaxation trajectories, Hessian matrices, and MBIS-derived atomic properties.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-Seed/THEMol.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1007
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1007.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1004
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1004.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1006
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1006.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1008
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1008.LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005
Retrain bank: WikiText-2 / GPT-2, random halves, seed 1005
This repository contains 100 fully retrained language models, not just scores.
Each model is GPT-2 (gpt2) fine-tuned on a different random 50% (2,328 documents)
of the 4,656-document WikiText-2 training set from
EleutherAI/bergson-wikitext-2-4656-chunks,
following the recipe of Bae et al. 2024, Training Data Attribution via Approximate
Unrolled Differentiation (App. B.1). retrained/base is trained on the full set with… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/LDS-retrain-bank-adamw-wikitext2-N4656-bs8-seed1005.SEEDBench2
Dataset Card for Dataset Name
The SEEDBench2 evaluation dataset hosted by VLMEval (authorized by the author).
Dataset Details
Language(s) (NLP): English
License: Apache 2.0
Repository: https://github.com/AILab-CVC/SEED-Bench
Paper [optional]: https://arxiv.org/abs/2311.17092
Citation
@misc{li2023seedbench2,
title={SEED-Bench-2: Benchmarking Multimodal Large Language Models},
author={Bohao Li and Yuying Ge and Yixiao Ge and Guangzhi Wang and Rui… See the full description on the dataset page: https://huggingface.co/datasets/VLMEval/SEEDBench2.world-seeds
World Seeds — every "by country" table, keyed by ISO 3166-1 alpha-2
Wikipedia has hundreds of "... by country" articles. The numbers live inside article tables, keyed by country names that differ from article to article. This dataset re-keys every such table to ISO2 so they join.
One CSV per source article under tables/. Columns: iso2, country, <original column names>. Values are kept exactly as printed (*_num twin columns hold the parsed number where one could be read).… See the full description on the dataset page: https://huggingface.co/datasets/Lilambd/world-seeds.seedance-2-5-free-tier-data
Seedance 2.5 free-tier observations
A dated record of what each channel actually grants on the Seedance 2.5 free tier — daily credits, daily generation caps, sign-up bonuses, clip length, extension ceiling, output resolution, watermark behaviour, and API availability.
Free-tier terms are the least documented part of an AI video product. Vendors publish a launch post with a headline number, then adjust the daily grant, move the watermark switch, or gate a model behind a… See the full description on the dataset page: https://huggingface.co/datasets/videofreetier/seedance-2-5-free-tier-data.golden_seeds_ev_depression
HEAE Golden Seeds: Human Empathy as Encoder Dataset
The first dataset to systematically encode human empathy as structured AI input for depression assessment
🌟 Overview
HEAE Golden Seeds introduces a groundbreaking approach by structurally integrating human empathy into AI decision-making. This meticulously hand-annotated dataset transforms tacit empathetic insights from special education teachers into structured, algorithmically actionable Empathy Vectors (EV)… See the full description on the dataset page: https://huggingface.co/datasets/Plum551/golden_seeds_ev_depression.NLLB-Seed_Tamasheq-Tifinagh-Scriptarisole-strideiq-gait-seed-v0.1
Arisole StrideIQ Gait Seed Release v0.1
Important release notice
This is an aggregate-only seed release.
Version 0.1 does not contain raw videos, images, pose landmarks, landmark sequences, per-user records, per-clip records, medical records, personal identifiers or biometric templates.
It should not be represented as a trainable clip-level gait dataset.
flow-edit-cube-triple-tkgrid-seeds30003-40004
Edit-placement (t x K x eb) campaign — OGBench cube-triple, task 2 — seeds 30003 & 40004
Seed scope: this repository contains only seeds 30003 and 40004. It is not the
full seed set for this campaign — seeds 10001 and 20002 were trained on separate hardware
and are not included here. Any per-cell mean computed from this repo alone is an n=2
estimate; see Caveats.
290 training runs from the uedit_place agent: a grid over where in the flow a
value-driven edit is applied (t), how… See the full description on the dataset page: https://huggingface.co/datasets/jaehyeokdoo2/flow-edit-cube-triple-tkgrid-seeds30003-40004.seed-price-datasetSEEDBench_IMGNLLB-Seed_Standard-Moroccan-Tamazightakai_flow_classifier_seed_pest_schemeSEEDBench_IMG_THmindbridge-phq9-hindi-seeds
MindBridge Hindi PHQ-9/GAD-7 — Gold Seeds (144 rows)
Hand-authored Hindi seeds for PHQ-9 + GAD-7 screening across three personas
(postnatal_mother, older_woman, man) in 1:1:1 distribution. Authored via
SuperWhisper Scribe with cloud LLM post-process; all rows
human-reviewed with review_status=accepted.
This seed set drives Phase B teacher expansion (in-context exemplars for
Gemma 4 26B-A4B MoE on Vertex MaaS) plus 24 Item-9 (suicidality) extras
authored separately. See companion… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-seeds.seed_pest_agri_schemeseeds
Post Operative
The Seeds dataset from the UCI repository.
Configurations and tasks
Configuration
Task
Description
seeds
Multiclass classification.
seeds_0
Binary classification.
Is the seed of class 0?
seeds_1
Binary classification.
Is the seed of class 1?
seeds_2
Binary classification.
Is the seed of class 2?
NLLB-Seed_Tamasheq-Latin-Scriptdataset-20260112-prime-seed
dataset-20260112-prime-seed
Created on: 2026-01-12T13:01:59.128698+00:00
Session ID: 2026-01-12T13:01:59.128698+00:00-5807
dataset-20251215-alpha-seed
dataset-20251215-alpha-seed
Created on: 2025-12-15T05:33:47.445983+00:00
Session ID: 2025-12-15T05:33:47.445983+00:00-4565
huggingface_5943_uay2vd_seedhf7812-feature-registry-seedalpine1.1-maxcode-deepbug-seedThis dataset has been cleaned and is ready for use. It will be updated soon, so stay tuned! We plan to add code-related questions and instructions, some of which will contain bugs. Some bugs will be marked, while others will be left for the LLM to identify. The goal is to train the LLM to improve its coding ability.
We're also working on having Gemini 2.5 Pro answer the questions. However, due to a low rate limit (only 25 requests per day), we may consider using Claude 3.7, which is also a… See the full description on the dataset page: https://huggingface.co/datasets/marcuscedricridia/alpine1.1-maxcode-deepbug-seed.prompt-seeds
Seeds
A small collection of seed prompts and question starters used for data generation and fine-tuning pipelines.
Dataset Details
Field
Value
Language
English
License
Apache 2.0
Format
CSV
Schema
Column
Type
Description
Questions
string
A seed question or prompt string
Usage
from datasets import load_dataset
ds = load_dataset("shaafsalman/seeds")
print(ds["train"][0])
# {'Questions': 'Write a python code to… See the full description on the dataset page: https://huggingface.co/datasets/shaafsalman/prompt-seeds.seed-oil-tracker-chains
Seed Oil Tracker: US Chain Restaurant Cooking Oil Dataset
What's in it: the cooking oil used by major US restaurant chains. Which oil each chain fries in, whether it is a seed oil (soybean, canola, corn, sunflower) or a more traditional fat (tallow, olive oil, avocado oil), and a per-chain cleanest and worst menu item.
Source: compiled and verified by Seed Oil Tracker (https://seedoiltracker.com) from each chain's own published ingredient and allergen statements. Updated on an… See the full description on the dataset page: https://huggingface.co/datasets/JaredOnAnIsland/seed-oil-tracker-chains.dataset-20251211-tiny-seed
dataset-20251211-tiny-seed
Generated: 2025-12-11T13:09:36.758524+00:00
Run ID: 2025-12-11T13:09:36.758524+00:00-8191
