datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
typed-decisions
Typed Decisions
A benchmark for typed probabilistic decisions over shared state. You give a
model one piece of unstructured state. It answers several typed questions about
that state at once, and every answer is a probability distribution rather than a
single label.
The schema follows the System One primitives used by
TypeSafe AI: noul, choice and score. A row replays against any API that implements that shape. This benchmark is
independent. It is not affiliated with TypeSafe… See the full description on the dataset page: https://huggingface.co/datasets/LocalLLaMA/typed-decisions.ultrachat-sharegpt-5GBalgebraic-stack
NOTE: Please see EleutherAI/proof-pile-2
This is a cherry-picked repackaging of the algebraic-stack segment from the proof-pile-2 dataset as parquet files
License
see EleutherAI/proof-pile-2
Citation
see EleutherAI/proof-pile-2
mathlib-types
Mathlib Types
This dataset contains information about types defined in Mathlib, the mathematical library for the Lean 4 theorem prover, extracted with lean_scout.
Extracted from the Mathlib commit with the following hash.
0df444a360eaa60ab8c11dca51a86af692955474
The dataset follows this schema:
fields:
- type:
datatype: string
nullable: false
name: name
- type:
datatype: string
nullable: true
name: module
- type:
datatype: string
nullable: false
name:… See the full description on the dataset page: https://huggingface.co/datasets/mathlib-initiative/mathlib-types.cross_code_eval_typescripttypebert
Dataset Card for "typebert"
More Information needed
open-ner-core-types
Dataset Card for OpenNER 1.0
OpenNER 1.0 is a standardized collection of openly-available named entity recognition (NER) datasets.
OpenNER contains 36 NER corpora that span 52 languages, human-annotated in varying named entity ontologies.
We correct annotation format issues, standardize the original datasets into a uniform representation with consistent entity type names across corpora, and provide the collection in a structure that enables research in multilingual and… See the full description on the dataset page: https://huggingface.co/datasets/bltlab/open-ner-core-types.four_types_weightedcav3_t-type_calcium_channels_butkiewicz-multimodaltypescript-codeMMMU_img_typecav3_t-type_calcium_channels_butkiewicz
Dataset Details
Dataset Description
This dataset was initially curated from HTS data at the PubChem database.
The curation process is documented in Butkiewicz et al.
Primary screening with AID 449739 identified inhibitors of Cav3 T-type calcium channels.
Four follow-up screens were performed to confirm inhibitory effects on smaller sets of compounds
involving AID 493021, AID 493022, AID 493023, and AID 493041.
AID 489005 was performed as counter screen validating active… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/cav3_t-type_calcium_channels_butkiewicz.tasksource-jev-typed-decisions
tasksource-jev-typed-decisions
One million decisions from 500+ Tasksource tasks across 300+ dataset families,
in a single format for models that receive their answer criteria at runtime.
The value is breadth with traceable supervision: most rows inherit labels,
ratings, or annotator votes from existing datasets, not labels invented by a
teacher model. The source field identifies the originating task; existing
train/dev/test boundaries are retained where the source provides them.… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-jev-typed-decisions.typed-decisions-v2-system-onetypes-of-film-shots
What a Shot!
2,919 film frames labeled with shot scale (8 classes), from film-grab.com.
annotator
count
how labeled
human
863
hand-labeled (gold)
ai
2,056
DINOv2 classifier; low-confidence frames re-judged by Claude Opus (active learning)
Columns
image, label — ambiguous, closeUp, detail, extremeLongShot, fullShot, longShot, mediumCloseUp, mediumShot
annotator — human or ai
source — human / v2_highconf / opus_review / opus_resolved_amb… See the full description on the dataset page: https://huggingface.co/datasets/szymonrucinski/types-of-film-shots.task280_stereoset_classification_stereotype_type
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task280_stereoset_classification_stereotype_type
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task280_stereoset_classification_stereotype_type.codebase-content-SWE-bench_Verified-no-comments-and-file-typesfenic-codebasegithub-file-programs-dataset-typescriptdev_set_v2_a1_crosscodeeval_typescript_20260805_125428cuisine_type
Dataset Card for "cuisine_type"
More Information needed
stack_edu_typescriptmusic_caps_4sec_wave_typetypescript-treesitter-dedupe-filtered-datasetsV2
Typescript CodeSearch Dataset (Shuu12121/typescript-treesitter-dedupe-filtered-datasetsV2)
Dataset Description
This dataset contains TypeScript functions and methods paired with their TSDoc comments, extracted from open-source TypeScript repositories on GitHub.
It is formatted similarly to the CodeSearchNet challenge dataset.
Each entry includes:
code: The source code of a typescript function or method.
docstring: The docstring or Javadoc associated with the… See the full description on the dataset page: https://huggingface.co/datasets/Shuu12121/typescript-treesitter-dedupe-filtered-datasetsV2.notebooks_by_repo_type
Dataset Card for "notebooks_by_repo_type"
More Information needed
hacker-news-dataset
Hacker News Dataset (2025)
Dataset Description
A comprehensive dataset of Hacker News content from 2025, containing stories, comments, users, and their relationships. This dataset enables deep analysis of technical discussions, trends, and community dynamics on one of the most influential technology forums.
Dataset Summary
Total Records: 38.4M+ across 10 tables
Stories: 287K+ submissions including links, Show HNs, Ask HNs
Comments: 2.5M+ discussion… See the full description on the dataset page: https://huggingface.co/datasets/typedef-ai/hacker-news-dataset.exp_rpt_crosscodeeval-typescript_10kmeal_type
Dataset Card for "meal_type"
More Information needed
eo-crop-type-belgium
Crop Type Segmentation Earth Observation Dataset of Belgium
This dataset is a ML ready dataset with Earth Observation data from ESA Sentinel-2 with a GSD of 10m. The segmentation labels are given for the crop-type in each 10m pixel. The spectral bands provided are those suggest in ESA WorldCereal documentation (Bands: 2,3,4,5,6,7,8,11,12). The dataset covers the majority of Belgium. Each image is 256x256 pixels.
Author of this dataset: Robert Cowlishaw (0x365)
Input data -… See the full description on the dataset page: https://huggingface.co/datasets/0x365/eo-crop-type-belgium.a1_crosscodeeval_typescript
