CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MAmmoTH-VL /MAmmoTH-VL-Instruct-12M MAmmoTH-VL-Instruct-12M 🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo Introduction Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses. The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.imagevisual-question-answering10M<n<100M67 likes3.7k downloads2y agoHugging Face02SEACrowd /mammoth_vl_sea_shard_5image100K<n<1M0 likes2k downloads10mo agoHugging Face03mammovlmbench /mammogps MammoGPS Dataset Summary MammoGPS is a benchmark for evaluating vision-language model spatial understanding on 2D mammography. The benchmark is designed for analysis-oriented evaluation rather than single-number leaderboard reporting: the goal is to separate failures of generic localization, medically relevant finding recognition, and landmark-grounded spatial reasoning. This repository currently includes benchmark task views for: finding localization finding… See the full description on the dataset page: https://huggingface.co/datasets/mammovlmbench/mammogps.imagevisual-question-answering100K<n<1M1 likes1.5k downloads5mo agoHugging Face04SEACrowd /mammoth_vl_sea_shard_4image100K<n<1M0 likes1.4k downloads10mo agoHugging Face05mamed0v /TurkmenSpeech Turkmen Speech Dataset (ASR) This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models. It is one of the largest publicly available Turkmen speech datasets. Dataset Overview Property Value Total clips 119,847 Total duration 251.86 hours Sampling rate 16,000 Hz Language Turkmen (tk) Split train Each item includes: audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.audioautomatic-speech-recognition100K<n<1M7 likes1.4k downloads11mo agoHugging Face06marin-dna /genomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1text10M<n<100M0 likes645 downloads7mo agoHugging Face07gonzalobenegas /genomes-v2-genome_set-mammals-intervals-v2_512_256text10M<n<100M0 likes521 downloads9mo agoHugging Face08marin-dna /genomes-v4-genome_set-mammals-intervals-v5_256_128text10M<n<100M0 likes503 downloads8mo agoHugging Face09marin-dna /genomes-v4-genome_set-mammals-intervals-v16_254_127-id0.3_cov0.3text10M<n<100M0 likes494 downloads7mo agoHugging Face10marin-dna /vertebrate-v1-cds_mammals_only marin-dna/vertebrate-v1-cds_mammals_only Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment. This draft covers the cds region cohort with mammals_only species scope and preserves source FASTA/2bit letter case. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked sequence, and… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-cds_mammals_only.tabular10M<n<100M0 likes493 downloads2mo agoHugging Face11marin-dna /genomes-v5-genome_set-mammals-intervals-v5_255_128 bolinas-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128 Mammals CDS (v5) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 41,848,032 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128.text10M<n<100M0 likes464 downloads4mo agoHugging Face12mamachang /medical-reasoningtext1K<n<10K35 likes456 downloads3y agoHugging Face13Hack90 /ref_seq_vertebrate_non_mammal_part_1tabular100K<n<1M0 likes431 downloads3y agoHugging Face14Evan-Lin /metric-mamba-ml2021-hungyi-corpus Dataset Card for "metric-mamba-ml2021-hungyi-corpus" More Information needed audio10K<n<100K0 likes419 downloads2y agoHugging Face15mamei16 /wikipedia_paragraphs Description This dataset consists of English Wikipedia articles, which first have been split by paragraph breaks and subsequently by spaces. For each resulting token, there is a corresponding binary ner_tag, which is 1 if a token was followed by paragraph break in the original text. There are two deliberate exceptions to this, which can be seen in the dataset generation code: The text is not split if a paragraph break is preceded by a colon (":"), to avoid lists being separated… See the full description on the dataset page: https://huggingface.co/datasets/mamei16/wikipedia_paragraphs.texttoken-classification1M<n<10M0 likes398 downloads1y agoHugging Face16moca-embed /MAmmoTH-VL-Instruct-12M MAmmoTH-VL-Instruct-12M used in MoCa Pre-training 🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper Introduction This is a VQA style dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from MAmmoTH-VL-Instruct-12M by concatenating prompts and responses. The dataset consists of interleaved multimodal examples. text is a string containing text while imagesare image binaries that can be loaded… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MAmmoTH-VL-Instruct-12M.textvisual-question-answering10M<n<100M1 likes381 downloads1y agoHugging Face17marin-dna /genomes-v4-genome_set-mammals-intervals-v1_256_128text10M<n<100M0 likes350 downloads8mo agoHugging Face18Hack90 /ref_seq_vertebrate_non_mammal_part_2tabularn<1K0 likes318 downloads3y agoHugging Face19davanstrien /MAMe2 Dataset Card for "MAMe2" More Information needed image10K<n<100K0 likes277 downloads3y agoHugging Face20SEACrowd /mammoth_vl_sea_shard_3image100K<n<1M0 likes277 downloads10mo agoHugging Face21mamiksik /processed-commit-diffs List of repositories included in the dataset Project Language Fetched Count Url Moby Go 5 943 https://github.com/moby/moby Rxjava Java 516 https://github.com/RxJava/ReactiveX Spring-framework Java 2 529 https://github.com/spring-framework/spring-project Chart.js Javascript 641 https://github.com/Chart.js/chartjs Three.js Javascript 1 512https://github.com/three.js/mrdoob Redux Javascript 592 https://github.com/redux/reduxjs React-native Javascript 2 901… See the full description on the dataset page: https://huggingface.co/datasets/mamiksik/processed-commit-diffs.text10K<n<100K5 likes274 downloads4y agoHugging Face22BangumiBase /mamahahanotsuregogamotokanodatta Bangumi Image Base of Mamahaha No Tsurego Ga Motokano Datta This is the image base of bangumi Mamahaha no Tsurego ga Motokano datta, we detected 40 characters, 3708 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mamahahanotsuregogamotokanodatta.image1K<n<10K0 likes273 downloads2y agoHugging Face23marin-dna /genomes-v5-genome_set-mammals-intervals-v1_255_128 bolinas-dna/genomes-v5-genome_set-mammals-intervals-v1_255_128 Mammals promoters (v1) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 12,926,544 sequences across 64 data/train/*.jsonl.zst shards (reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v1_255_128.text10M<n<100M0 likes273 downloads4mo agoHugging Face24ivangtorre /watkins-marine-mammal-full-cuts Watkins Marine Mammal Sound Database Dataset Description The Watkins Marine Mammal Sound Database (WMMSD) is one of the largest historical collections of marine mammal vocalizations. It contains 15,248 recordings spanning nearly seven decades from 54 marine mammal species, including whales, dolphins, porpoises, seals, sea lions, manatees, sea otters, and other marine mammals. This repository provides the complete dataset in a format fully compatible with the… See the full description on the dataset page: https://huggingface.co/datasets/ivangtorre/watkins-marine-mammal-full-cuts.audio10K<n<100K1 likes241 downloads2mo agoHugging Face25marin-dna /genomes-v4-genome_set-mammals-intervals-v15_256_128text1M<n<10M0 likes216 downloads8mo agoHugging Face26YMA-MamunAI /barcha-speech-datasetlar Barcha O'zbek Speech Datasetlari O'zbek tili uchun yig'ilgan barcha ochiq audio-matn datasetlari. Rows: 968,654 Audio: 16kHz, mono Duration filter: 0.5s - 31s Language: Uzbek audio100K<n<1M0 likes213 downloads4mo agoHugging Face27CardinalOperations /MAMO Overview This dataset is a direct copy of the MAMO Optimization Data, with its EasyLP and ComplexLP components duplicated but with adapted field names. Citation @misc{huang2024mamo, title={Mamo: a Mathematical Modeling Benchmark with Solvers}, author={Xuhan Huang and Qingning Shen and Yan Hu and Anningzhe Gao and Benyou Wang}, year={2024}, eprint={2405.13144}, archivePrefix={arXiv}, primaryClass={cs.AI} } textn<1K6 likes208 downloads2y agoHugging Face28marin-dna /genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128 bolinas-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128 20 mammals (segmentation) segmentation enhancers (v20) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 8,672,102 sequences across 64… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128.text1M<n<10M0 likes192 downloads4mo agoHugging Face29marin-dna /genomes-v4-genome_set-mammals-intervals-v1_255_128-id0.3_cov0.3text10M<n<100M0 likes188 downloads7mo agoHugging Face30marin-dna /genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128 bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128 20 mammals projected conserved enhancers (v30) sequences — 255 bp DNA windows for genomic language model pretraining. Part of the bolinas-dna/genomes-v5 training-dataset family produced by the snakemake/training_dataset pipeline (commit 8db58254831f). Each repo in the family is one (genome_set, region-recipe) combination. Size 6,549,730 sequences across 64 data/train/*.jsonl.zst shards (reverse… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128.text1M<n<10M0 likes185 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.