datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.mammoth_vl_sea_shard_5mammogps
MammoGPS
Dataset Summary
MammoGPS is a benchmark for evaluating vision-language model spatial understanding on 2D mammography. The benchmark is designed for analysis-oriented evaluation rather than single-number leaderboard reporting: the goal is to separate failures of generic localization, medically relevant finding recognition, and landmark-grounded spatial reasoning.
This repository currently includes benchmark task views for:
finding localization
finding… See the full description on the dataset page: https://huggingface.co/datasets/mammovlmbench/mammogps.mammoth_vl_sea_shard_4TurkmenSpeech
Turkmen Speech Dataset (ASR)
This dataset contains 251 hours of Turkmen speech audio with transcriptions, intended for training and evaluating Automatic Speech Recognition (ASR) models.
It is one of the largest publicly available Turkmen speech datasets.
Dataset Overview
Property
Value
Total clips
119,847
Total duration
251.86 hours
Sampling rate
16,000 Hz
Language
Turkmen (tk)
Split
train
Each item includes:
audio: waveform + sampling… See the full description on the dataset page: https://huggingface.co/datasets/mamed0v/TurkmenSpeech.genomes-v4-genome_set-mammals-intervals-v1_255_128-id1_cov1genomes-v2-genome_set-mammals-intervals-v2_512_256genomes-v4-genome_set-mammals-intervals-v5_256_128genomes-v4-genome_set-mammals-intervals-v16_254_127-id0.3_cov0.3vertebrate-v1-cds_mammals_only
marin-dna/vertebrate-v1-cds_mammals_only
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment. This
draft covers the cds region cohort with mammals_only species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked sequence, and… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-cds_mammals_only.genomes-v5-genome_set-mammals-intervals-v5_255_128
bolinas-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128
Mammals CDS (v5) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
41,848,032 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is an… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v5_255_128.medical-reasoningref_seq_vertebrate_non_mammal_part_1metric-mamba-ml2021-hungyi-corpus
Dataset Card for "metric-mamba-ml2021-hungyi-corpus"
More Information needed
wikipedia_paragraphs
Description
This dataset consists of English Wikipedia articles, which first have been split by paragraph breaks and subsequently by spaces. For each resulting token, there is a corresponding binary ner_tag, which is 1 if a token was followed by paragraph break in the original text. There are two deliberate exceptions to this, which can be seen in the dataset generation code:
The text is not split if a paragraph break is preceded by a colon (":"), to avoid lists being separated… See the full description on the dataset page: https://huggingface.co/datasets/mamei16/wikipedia_paragraphs.MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M used in MoCa Pre-training
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
Introduction
This is a VQA style dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from MAmmoTH-VL-Instruct-12M by concatenating prompts and responses.
The dataset consists of interleaved multimodal examples. text is a string containing text while imagesare image binaries that can be loaded… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/MAmmoTH-VL-Instruct-12M.genomes-v4-genome_set-mammals-intervals-v1_256_128ref_seq_vertebrate_non_mammal_part_2MAMe2
Dataset Card for "MAMe2"
More Information needed
mammoth_vl_sea_shard_3processed-commit-diffs
List of repositories included in the dataset
Project
Language
Fetched Count
Url
Moby
Go
5 943
https://github.com/moby/moby
Rxjava
Java
516
https://github.com/RxJava/ReactiveX
Spring-framework
Java
2 529
https://github.com/spring-framework/spring-project
Chart.js
Javascript
641
https://github.com/Chart.js/chartjs
Three.js
Javascript
1 512https://github.com/three.js/mrdoob
Redux
Javascript
592
https://github.com/redux/reduxjs
React-native
Javascript
2 901… See the full description on the dataset page: https://huggingface.co/datasets/mamiksik/processed-commit-diffs.mamahahanotsuregogamotokanodatta
Bangumi Image Base of Mamahaha No Tsurego Ga Motokano Datta
This is the image base of bangumi Mamahaha no Tsurego ga Motokano datta, we detected 40 characters, 3708 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mamahahanotsuregogamotokanodatta.genomes-v5-genome_set-mammals-intervals-v1_255_128
bolinas-dna/genomes-v5-genome_set-mammals-intervals-v1_255_128
Mammals promoters (v1) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
12,926,544 sequences across 64 data/train/*.jsonl.zst shards
(reverse complements included). This is… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals-intervals-v1_255_128.watkins-marine-mammal-full-cuts
Watkins Marine Mammal Sound Database
Dataset Description
The Watkins Marine Mammal Sound Database (WMMSD) is one of the largest historical collections of marine mammal vocalizations. It contains 15,248 recordings spanning nearly seven decades from 54 marine mammal species, including whales, dolphins, porpoises, seals, sea lions, manatees, sea otters, and other marine mammals.
This repository provides the complete dataset in a format fully compatible with the… See the full description on the dataset page: https://huggingface.co/datasets/ivangtorre/watkins-marine-mammal-full-cuts.genomes-v4-genome_set-mammals-intervals-v15_256_128barcha-speech-datasetlar
Barcha O'zbek Speech Datasetlari
O'zbek tili uchun yig'ilgan barcha ochiq audio-matn datasetlari.
Rows: 968,654
Audio: 16kHz, mono
Duration filter: 0.5s - 31s
Language: Uzbek
MAMO
Overview
This dataset is a direct copy of the MAMO Optimization Data, with its EasyLP and ComplexLP components duplicated but with adapted field names.
Citation
@misc{huang2024mamo,
title={Mamo: a Mathematical Modeling Benchmark with Solvers},
author={Xuhan Huang and Qingning Shen and Yan Hu and Anningzhe Gao and Benyou Wang},
year={2024},
eprint={2405.13144},
archivePrefix={arXiv},
primaryClass={cs.AI}
}
genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128
bolinas-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128
20 mammals (segmentation) segmentation enhancers (v20) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
8,672,102 sequences across 64… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-enhancer_seg_mammals_v1-intervals-v20_255_128.genomes-v4-genome_set-mammals-intervals-v1_255_128-id0.3_cov0.3genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128
bolinas-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128
20 mammals projected conserved enhancers (v30) sequences — 255 bp DNA windows
for genomic language model pretraining.
Part of the bolinas-dna/genomes-v5 training-dataset family produced by the
snakemake/training_dataset pipeline (commit
8db58254831f). Each repo in the family is one
(genome_set, region-recipe) combination.
Size
6,549,730 sequences across 64 data/train/*.jsonl.zst shards
(reverse… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/genomes-v5-genome_set-mammals_seg20-intervals-v30_255_128.
