datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
4DThinker-Training-Data
4DThinker Training Data
This repository contains the training data for 4DThinker, a framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, built upon SpatialVID and DSR_Suite-Data.
Data Structure
data/
├── dift_data.jsonl # DIFT training data (~38K samples)
├── 4drl_data_filtered.jsonl # 4DRL training data (~37K samples)
└── processed_data/ # Video frames & mask overlays
├── <video_id>/
│ ├── frames/… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/4DThinker-Training-Data.SCBench-preprocessedThis is the preprocessed version of Microsoft SCBench, used by KVzip:
Each data example has a format of {context: str, question: List[str], answers: List[str]}
Each dataset contains only examples whose context token length (measured with the LLaMA3 tokenizer) is less than 125K, fitting within the context limit of LLaMA3 models.
We also provide shortened versions of SCBench, excluding tasks {choice_eng, qa_eng, and vt}, which are difficult to shorten.
The "tiny" tag (e.g., scbench_kv_tiny)… See the full description on the dataset page: https://huggingface.co/datasets/Jang-Hyun/SCBench-preprocessed.financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.message_historyjaneflores6357starcoderdata-python-edu-lang-score
Dataset Card for Starcoder Data with Python Education and Language Scores
Dataset Summary
The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.Jan2023Abstracts
Dataset Card for "Jan2023Abstracts"
More Information needed
aya_collection
This dataset is uploaded in two places: here and additionally here as 'Aya Collection Language Split.' These datasets are identical in content but differ in structure of upload. This dataset is structured by folders split according to dataset name. The version here instead divides the Aya collection into folders split by language. We recommend you use the language split version if you are only interested in downloading data for a single or smaller set of languages, and this version if you… See the full description on the dataset page: https://huggingface.co/datasets/janavivekariya/aya_collection.jansmba63JANuS_datasetThis repository hosts the JANuS (Joint Annotations and Names) dataset introduced in the 2023 paper Distributionally Robust Classification on a Data Budget.
As of this writing, ours is the only public dataset which is both fully annotated with ground-truth labels and fully captioned with web-scraped captions.
It is designed to be used for controlled experiments with vision-language models.
What is in JANuS?
JANuS provides metadata and image links for four new training datasets; all… See the full description on the dataset page: https://huggingface.co/datasets/penfever/JANuS_dataset.pexels-photos-janpf
Dataset migrated
This location shoud be considered obsolete. Data has been copied to
https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf
In a month or so, I should remove the data from here, to be nice to huggingface.
The images there have been renamed to match their md5 checksum. However, a translation table is in this repo, if for some reason you need it.
Downloading
If for some reason, downloading from here is needed, you can use
huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/ppbrown/pexels-photos-janpf.illustrationcrows_pairs_multilingualOriginal from https://gitlab.inria.fr/french-crows-pairs/acl-2022-paper-data-and-code/-/tree/main/.
Data Statement for CrowS-Pairs-fr
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.
For full details… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/crows_pairs_multilingual.agentic-drug-discovery-system
Agentic Drug Discovery System
This card describes the public 0.3.0.dev3 Agentic Drug Discovery System mirror.
Scope. The proposed eight-stage, long-horizon agentic drug discovery system remains a research scaffold rather than a completed public platform. Seven of eight planned atlases have no standalone public data, and the demonstrated continuous multi-stage program currently covers one disease/target slice traversed retrospectively.
It contains the executable control plane… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/agentic-drug-discovery-system.co3d_10fJanus-Pro-R1-Data
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation
Kaihang Pan1*, Yang Wu2*, Wendong Bu1*, Kai Shen1‡, Juncheng Li1†, Yingting Wang2,
Yunfei Li2, Siliang Tang1, Jun Xiao1, Fei Wu1, Hang Zhao2, Yueting Zhuang1
1Zhejiang University, 2Ant Group
*Equal Contribution, ‡Project Lead, †Corresponding Author
Data Format
Below is an example of how a single data record is organized.
{
'promptid':0… See the full description on the dataset page: https://huggingface.co/datasets/midbee/Janus-Pro-R1-Data.instruction-convert-audio-whispervq-llama3.2-compresspexels-photos-janpf
Images:
There are approximately 130K images, borrowed from pexels.com.
Thanks to those folks for curating a wonderful resource.
There are millions more images on pexels. These particular ones were selected by
the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt .
The filenames are based on the md5 hash of each image.
Download From here or from pexels.com: You choose
For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.COVID-BLUESThis dataset corresponds to the paper COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis published in IEEE Journal of Biomedical and Health Informatics 2025.
If you face a paywell, you can access the pre-final version of the paper on ArXiv: arxiv.org/abs/2509.10556
If you use this dataset, you have to cite our paper:
@article{wiedemann2025covid,
author={Wiedemann, Nina and Boer, Dianne de Korte-de and Richter, Matthias and van de Weijer, Sjors and Buhre… See the full description on the dataset page: https://huggingface.co/datasets/jannisborn/COVID-BLUES.vivoice-libris-mls-eng-10k-tokens-v0.1CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit
Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset
Dataset Summary
This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance.
The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.lichess_2023_janoct_shardsLichess data in 2023 from Jan-Octichigo_tokens_v1teacher-embedding-corpusnarrow-model-safety-eval
Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset
Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.janusdata
Janus Dataset
This repository hosts the dataset used by the Janus project for training and evaluation on 5G core logs, specifications, and source code.
Repository layout
raw_data/
spec_3gpp/: filtered_3gpp_specs_part1.json … filtered_3gpp_specs_part5.json (split <5 MB each) containing 3GPP spec metadata.
3gpp_openapi_yaml_rel17/: Release 17 OpenAPI YAML dumps, one directory per spec.
logs/: Open5GS core deployment logs grouped by call-flows.
anomalous_logs/: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/umakantk/janusdata.bigscience-lama
Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models.
@inproceedings{petroni2020how,
title={How Context Affects Language Models' Factual Predictions},
author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel},
booktitle={Automated Knowledge Base Construction},
year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.Appropriations
💵 U.S. Appropriations Dataset (1995–2025)
This dataset links enacted U.S. Public Laws with their corresponding Explanatory Statements and Appropriations Titles,
covering the major federal appropriations acts from FY1995 through FY2025.
📊 Structure
Each entry includes:
public_law: Official citation of the enacted appropriations law (e.g. P.L. 117-328)
explanatory_statement: House or Senate report number accompanying the law (e.g. H. Rpt. 117-328)
appropriation_title:… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Appropriations.nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency.
flux_generations
