CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ruggsea /social-sim-bench-genstext1K<n<10K0 likes7.8k downloads4mo agoHugging Face02livecodebench /code_generation LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • LiveCodeBench is a "live" updating benchmark for holistically evaluating code related capabilities of LLMs. Particularly, it evaluates LLMs across a range of capabilties including code generation, self-repair, test output prediction, and code execution. This is the code generation scenario of LiveCodeBench. It is also… See the full description on the dataset page: https://huggingface.co/datasets/livecodebench/code_generation.textn<1K32 likes5.3k downloads2y agoHugging Face03shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face04General-Medical-AI /GMAI-VL-5.5M GMAI-VL-5.5M Dataset GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets. This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.imagevisual-question-answering1M<n<10M6 likes2.8k downloads6mo agoHugging Face05jacobmorrison /generation_10627 allenai/open_instruct: Generation Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'generation_10627', 'mode': 'generation', 'model_name_or_path': '/generation_model', 'push_to_hub': True, 'revision': 'main', 'save_filename': '/output/shards/10627/23.jsonl', 'skill': 'chat'} dataset_args: {'dataset_end_idx': 71136… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/generation_10627.text1M<n<10M0 likes2.6k downloads2y agoHugging Face06geniacllm /CulturaY-ja-askllm-v1 CulturaY-ja-askllm-v1 多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable knowledge of the world, and strictly… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-v1.tabular10M<n<100M1 likes1.9k downloads2y agoHugging Face07geniacllm /CulturaY-ja-askllm-deduped-v1 CulturaY-ja-askllm-deduped-v1 多言語データセット ontocord/CulturaY の日本語パート ja に対して、 Ask-LLM 手法でスコア付けしたデータセットです。 元データセットのカラムに加え askllm_score というカラムが追加されており、ここに Ask-LLM のスコアが格納されています。 スコア付けした後、重複削除をしたデータセットです。 Ask-LLM でスコア付けに使用した LLM は Rakuten/RakutenAI-7B-instruct で、プロンプトは以下の通りです。 ### {data} ### Does the previous paragraph demarcated within ### and ### contain informative signal for pre-training a large-language model? An informative datapoint should be well-formatted, contain some usable… See the full description on the dataset page: https://huggingface.co/datasets/geniacllm/CulturaY-ja-askllm-deduped-v1.text10M<n<100M0 likes1.7k downloads2y agoHugging Face08BelleGroup /generated_chat_0.4M Generated_Chat_0.4M 内容 包含约40万条由BELLE项目生成的个性化角色对话数据,包含角色介绍。 注意:此数据集是由ChatGPT产生的,未经过严格校验,题目或解题过程可能包含错误。使用过程中请注意这一点。 样例 { "instruction": "于以下角色信息完成一段对话\nAlice是一个年轻的律师,她正在为一个涉及濒危物种保护的案件做准备。Bob是一位当地的居民,他所在的社区经常受到该濒危物种的干扰。\n", "input": "", "output": "\nAlice: 你好Bob,我是濒危物种保护案的律师,很高兴见到你。我们前几天发给你的文件你看到了吗?\nBob: 看到了,你们说要禁止我们在该物种的常住地采集蘑菇,这很不公平。\nAlice: 保护濒危物种对整个生态系统都很重要,如果我们的举动损害了这个物种的生存,那么整个生态系统都将遭受损失。\nBob: 我理解您的立场,但是我们一直以来都依靠这个物种来维持我们的经济生活,现在我们要怎么办?\nAlice:… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/generated_chat_0.4M.text100K<n<1M69 likes1.6k downloads3y agoHugging Face09jacobmorrison /generation_3379 allenai/open_instruct: Generation Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'generation_3379', 'mode': 'generation', 'model_name_or_path': '/model', 'push_to_hub': True, 'revision': 'main', 'save_filename': '/output/shards/3379/2.jsonl', 'skill': 'chat'} dataset_args: {'dataset_end_idx': 6000… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/generation_3379.text100K<n<1M0 likes1.5k downloads2y agoHugging Face10General-Medical-AI /GMAI-Reasoning10K GMAI-Reasoning10K Medical Reasoning dataset used in GMAI-VL-R1 Data description GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI. Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.imagevisual-question-answering10K<n<100K6 likes1.5k downloads1y agoHugging Face11marin-dna /genomes-v4-genome_set-animals-intervals-v5_256_128text100M<n<1B0 likes1.5k downloads8mo agoHugging Face12marin-dna /genomes-v4-genome_set-animals-intervals-v11_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face13jacobmorrison /generation_2413 allenai/open_instruct: Generation Dataset See https://github.com/allenai/open-instruct/blob/main/docs/algorithms/rejection_sampling.md for more detail Configs args: {'add_timestamp': False, 'hf_entity': 'jacobmorrison', 'hf_repo_id': 'generation_2413', 'mode': 'generation', 'model_name_or_path': '/model', 'push_to_hub': True, 'revision': 'main', 'save_filename': '/output/shards/2413/1.jsonl', 'skill': 'chat'} dataset_args: {'dataset_end_idx': 4000… See the full description on the dataset page: https://huggingface.co/datasets/jacobmorrison/generation_2413.text100K<n<1M0 likes1.3k downloads2y agoHugging Face14marin-dna /genomes-v4-genome_set-animals-intervals-v10_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face15marin-dna /genomes-v4-genome_set-animals-intervals-v12_256_128text100M<n<1B0 likes1.3k downloads8mo agoHugging Face16marin-dna /genomes-v5-genome_set-animals_order204-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128 204 animals (one per order) CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit main). Each repo in the family is one (genome_set, region-recipe) combination. Size 101,114,252 sequences across 64 data/train/*.jsonl.zst shards (reverse complements… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals_order204-intervals-v5_255_128.text100M<n<1B0 likes1.3k downloads3mo agoHugging Face17marin-dna /genomes-v4-genome_set-animals-intervals-v13_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face18marin-dna /genomes-v4-genome_set-animals-intervals-v4_512_256text10M<n<100M0 likes1.3k downloads8mo agoHugging Face19marin-dna /genomes-v4-genome_set-animals-intervals-v14_256_128text10M<n<100M0 likes1.3k downloads8mo agoHugging Face20marin-dna /genomes-v5-genome_set-animals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v1_255_128 Animals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 68,286,166 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v1_255_128.text10M<n<100M0 likes1.3k downloads1mo agoHugging Face21marin-dna /genomes-v4-genome_set-animals-intervals-v1_256_128text10M<n<100M0 likes1.2k downloads8mo agoHugging Face22VIPL-GENUN /Joint-1.6M-1024pxFor more information, please see: arXiv: https://arxiv.org/abs/2505.19084 Project page: https://VIPL-GENUN.github.io/Project-Jodi GitHub: https://github.com/VIPL-GENUN/Jodi Joint-1.6M Dataset We collect images with high quality and diversity from several publicly available sources, including Subjects200K, Aesthetic-4K, Pexels photos, and Pexels portrait. All of these images have resolutions over 1024×1024, which is advantageous for training a high-resolution generative model.… See the full description on the dataset page: https://huggingface.co/datasets/VIPL-GENUN/Joint-1.6M-1024px.text100K<n<1M2 likes1.2k downloads1y agoHugging Face23marin-dna /genomes-v4-genome_set-animals-intervals-v6_256_128text10M<n<100M0 likes1.1k downloads8mo agoHugging Face24marin-dna /genomes-v5-genome_set-animals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-animals-intervals-v5_255_128 Animals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 242,334,716 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-animals-intervals-v5_255_128.text100M<n<1B0 likes1k downloads1mo agoHugging Face25marin-dna /genomes-v4-genome_set-animals-intervals-v8_256_128text10M<n<100M0 likes928 downloads8mo agoHugging Face26gentaiscool /bitext_sib200_minerstext100K<n<1M3 likes923 downloads2y agoHugging Face27garcianacho /human_genometabular10M<n<100M0 likes919 downloads3y agoHugging Face28marin-dna /genomes-v4-genome_set-animals-intervals-v9_256_128text10M<n<100M0 likes888 downloads8mo agoHugging Face29marin-dna /genomes-v4-genome_set-animals-intervals-v15_256_128text10M<n<100M0 likes887 downloads8mo agoHugging Face30hltcoe /megawika-report-generation Dataset Card for MegaWika for Report Generation Dataset Summary MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.textsummarization100K<n<1M6 likes860 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.