CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jankin123 /4DThinker-Training-Data 4DThinker Training Data This repository contains the training data for 4DThinker, a framework that enables VLMs to "think with 4D" through dynamic latent mental imagery, built upon SpatialVID and DSR_Suite-Data. Data Structure data/ ├── dift_data.jsonl # DIFT training data (~38K samples) ├── 4drl_data_filtered.jsonl # 4DRL training data (~37K samples) └── processed_data/ # Video frames & mask overlays ├── <video_id>/ │ ├── frames/… See the full description on the dataset page: https://huggingface.co/datasets/jankin123/4DThinker-Training-Data.1 likes13k downloads5mo agoHugging Face02Jang-Hyun /SCBench-preprocessedThis is the preprocessed version of Microsoft SCBench, used by KVzip: Each data example has a format of {context: str, question: List[str], answers: List[str]} Each dataset contains only examples whose context token length (measured with the LLaMA3 tokenizer) is less than 125K, fitting within the context limit of LLaMA3 models. We also provide shortened versions of SCBench, excluding tasks {choice_eng, qa_eng, and vt}, which are difficult to shorten. The "tiny" tag (e.g., scbench_kv_tiny)… See the full description on the dataset page: https://huggingface.co/datasets/Jang-Hyun/SCBench-preprocessed.text1K<n<10K2 likes8k downloads8mo agoHugging Face03JanosAudran /financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system. Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences. Sentiment labels are provided on a per filing basis from the market reaction around the filing data. Additional metadata for each filing is included in the dataset.tabularfill-mask10M<n<100M77 likes7.2k downloads4y agoHugging Face04janblue /message_history0 likes5.4k downloads2y agoHugging Face05janeflores6357 /janeflores63577 likes2.8k downloads21d agoHugging Face06JanSchTech /starcoderdata-python-edu-lang-score Dataset Card for Starcoder Data with Python Education and Language Scores Dataset Summary The starcoderdata-python-edu-lang-score dataset contains the Python subset of the starcoderdata dataset. It augments the existing Python subset with features that assess the educational quality of code and classify the language of code comments. This dataset was created for high-quality Python education and language-based training, with a primary focus on facilitating models that can… See the full description on the dataset page: https://huggingface.co/datasets/JanSchTech/starcoderdata-python-edu-lang-score.tabular1M<n<10M2 likes2.8k downloads2y agoHugging Face07ITOCJ /Jan2023Abstracts Dataset Card for "Jan2023Abstracts" More Information needed text10M<n<100M0 likes1.8k downloads4y agoHugging Face08janavivekariya /aya_collection This dataset is uploaded in two places: here and additionally here as 'Aya Collection Language Split.' These datasets are identical in content but differ in structure of upload. This dataset is structured by folders split according to dataset name. The version here instead divides the Aya collection into folders split by language. We recommend you use the language split version if you are only interested in downloading data for a single or smaller set of languages, and this version if you… See the full description on the dataset page: https://huggingface.co/datasets/janavivekariya/aya_collection.tabulartext-classification100M<n<1B0 likes1.6k downloads10mo agoHugging Face09Rikah45563 /jansmba630 likes1.1k downloads3y agoHugging Face10penfever /JANuS_datasetThis repository hosts the JANuS (Joint Annotations and Names) dataset introduced in the 2023 paper Distributionally Robust Classification on a Data Budget. As of this writing, ours is the only public dataset which is both fully annotated with ground-truth labels and fully captioned with web-scraped captions. It is designed to be used for controlled experiments with vision-language models. What is in JANuS? JANuS provides metadata and image links for four new training datasets; all… See the full description on the dataset page: https://huggingface.co/datasets/penfever/JANuS_dataset.imagezero-shot-classification10K<n<100K0 likes969 downloads3y agoHugging Face11ppbrown /pexels-photos-janpf Dataset migrated This location shoud be considered obsolete. Data has been copied to https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf In a month or so, I should remove the data from here, to be nice to huggingface. The images there have been renamed to match their md5 checksum. However, a translation table is in this repo, if for some reason you need it. Downloading If for some reason, downloading from here is needed, you can use huggingface-cli… See the full description on the dataset page: https://huggingface.co/datasets/ppbrown/pexels-photos-janpf.image18 likes923 downloads2y agoHugging Face12Jannchie /illustration0 likes907 downloads1y agoHugging Face13jannalu /crows_pairs_multilingualOriginal from https://gitlab.inria.fr/french-crows-pairs/acl-2022-paper-data-and-code/-/tree/main/. Data Statement for CrowS-Pairs-fr How to use this document: Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years. For full details… See the full description on the dataset page: https://huggingface.co/datasets/jannalu/crows_pairs_multilingual.text1K<n<10K0 likes852 downloads11mo agoHugging Face14jang1563 /agentic-drug-discovery-system Agentic Drug Discovery System This card describes the public 0.3.0.dev3 Agentic Drug Discovery System mirror. Scope. The proposed eight-stage, long-horizon agentic drug discovery system remains a research scaffold rather than a completed public platform. Seven of eight planned atlases have no standalone public data, and the demonstrated continuous multi-stage program currently covers one disease/target slice traversed retrospectively. It contains the executable control plane… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/agentic-drug-discovery-system.0 likes697 downloads13d agoHugging Face15janellecai /co3d_10fimage0 likes677 downloads1y agoHugging Face16midbee /Janus-Pro-R1-Data Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation Kaihang Pan1*, Yang Wu2*, Wendong Bu1*, Kai Shen1‡, Juncheng Li1†, Yingting Wang2, Yunfei Li2, Siliang Tang1, Jun Xiao1, Fei Wu1, Hang Zhao2, Yueting Zhuang1 1Zhejiang University, 2Ant Group *Equal Contribution, ‡Project Lead, †Corresponding Author Data Format Below is an example of how a single data record is organized. { 'promptid':0… See the full description on the dataset page: https://huggingface.co/datasets/midbee/Janus-Pro-R1-Data.text10K<n<100K1 likes671 downloads1y agoHugging Face17jan-hq /instruction-convert-audio-whispervq-llama3.2-compresstext1M<n<10M0 likes658 downloads2y agoHugging Face18opendiffusionai /pexels-photos-janpf Images: There are approximately 130K images, borrowed from pexels.com. Thanks to those folks for curating a wonderful resource. There are millions more images on pexels. These particular ones were selected by the list of urls at https://github.com/janpf/self-supervised-multi-task-aesthetic-pretraining/blob/main/dataset/urls.txt . The filenames are based on the md5 hash of each image. Download From here or from pexels.com: You choose For those people who like… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/pexels-photos-janpf.text-to-image100K<n<1M45 likes652 downloads8mo agoHugging Face19jannisborn /COVID-BLUESThis dataset corresponds to the paper COVID-BLUeS -- A Prospective Study on the Value of AI in Lung Ultrasound Analysis published in IEEE Journal of Biomedical and Health Informatics 2025. If you face a paywell, you can access the pre-final version of the paper on ArXiv: arxiv.org/abs/2509.10556 If you use this dataset, you have to cite our paper: @article{wiedemann2025covid, author={Wiedemann, Nina and Boer, Dianne de Korte-de and Richter, Matthias and van de Weijer, Sjors and Buhre… See the full description on the dataset page: https://huggingface.co/datasets/jannisborn/COVID-BLUES.videovideo-classificationn<1K1 likes612 downloads11mo agoHugging Face20jan-hq /vivoice-libris-mls-eng-10k-tokens-v0.1text1M<n<10M0 likes599 downloads2y agoHugging Face21leeroy-jankins /CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset Dataset Summary This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance. The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.documentquestion-answering0 likes578 downloads2mo agoHugging Face22ezipe /lichess_2023_janoct_shardsLichess data in 2023 from Jan-Oct0 likes555 downloads2y agoHugging Face23jan-hq /ichigo_tokens_v1text1M<n<10M0 likes550 downloads2y agoHugging Face24jangedoo /teacher-embedding-corpustext1M<n<10M0 likes492 downloads3mo agoHugging Face25jang1563 /narrow-model-safety-eval Narrow Model Safety Evaluation — Protein Dual-Use Risk Dataset Summary: Annotations, results, and evaluation data for a proof-of-concept framework assessing dual-use risk in narrow scientific AI models. Two lines of work: (1) structure-level metrics — FSPE, FSI, and Physical Realizability Tier — on eight published protein toxins and mechanism-matched benign controls (ESM-2, ProteinMPNN); (2) mechanism generalization — a leave-one-mechanism-out panel measuring what an… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/narrow-model-safety-eval.tabularothern<1K0 likes491 downloads4d agoHugging Face26umakantk /janusdata Janus Dataset This repository hosts the dataset used by the Janus project for training and evaluation on 5G core logs, specifications, and source code. Repository layout raw_data/ spec_3gpp/: filtered_3gpp_specs_part1.json … filtered_3gpp_specs_part5.json (split <5 MB each) containing 3GPP spec metadata. 3gpp_openapi_yaml_rel17/: Release 17 OpenAPI YAML dumps, one directory per spec. logs/: Open5GS core deployment logs grouped by call-flows. anomalous_logs/: Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/umakantk/janusdata.1 likes473 downloads6mo agoHugging Face27janck /bigscience-lama Dataset Card for LAMA: LAnguage Model Analysis - a dataset for probing and analyzing the factual and commonsense knowledge contained in pretrained language models. @inproceedings{petroni2020how, title={How Context Affects Language Models' Factual Predictions}, author={Fabio Petroni and Patrick Lewis and Aleksandra Piktus and Tim Rockt{"a}schel and Yuxiang Wu and Alexander H. Miller and Sebastian Riedel}, booktitle={Automated Knowledge Base Construction}, year={2020}… See the full description on the dataset page: https://huggingface.co/datasets/janck/bigscience-lama.texttext-retrieval10K<n<100K1 likes464 downloads4y agoHugging Face28leeroy-jankins /Appropriations 💵 U.S. Appropriations Dataset (1995–2025) This dataset links enacted U.S. Public Laws with their corresponding Explanatory Statements and Appropriations Titles, covering the major federal appropriations acts from FY1995 through FY2025. 📊 Structure Each entry includes: public_law: Official citation of the enacted appropriations law (e.g. P.L. 117-328) explanatory_statement: House or Senate report number accompanying the law (e.g. H. Rpt. 117-328) appropriation_title:… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Appropriations.text100K<n<1M0 likes462 downloads8mo agoHugging Face29jangedoo /nepali-corpusThis is a clone of https://huggingface.co/datasets/Boredoom17/Nepali-Corpus but the data is sharded for efficiency. text1M<n<10M0 likes458 downloads3mo agoHugging Face30janellecai /flux_generationsimagen<1K0 likes442 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.