datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.opensource_mirrorwan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.transformers-dependents
transformers metrics
This dataset contains metrics about the huggingface/transformers package.
Number of repositories in the dataset: 27067
Number of packages in the dataset: 823
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 65 packages that have more than 1000 stars.
There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.Llasa_opensource_speech_data_160k_hours_tokenized
Update (2025-02-07): Our paper has been released!
This script is for merging tokenized speech datasets stored in memmap format. The input datasets can be combined to form larger training datasets.
import numpy as np
import os
def merge_memmap_datasets(dataset_dirs, output_dir):
# Ensure the output directory exists
os.makedirs(output_dir, exist_ok=True)
# Dataset splits to be merged
splits = ['train', 'val']
for split in splits:
shapes = []
seq_len =… See the full description on the dataset page: https://huggingface.co/datasets/HKUSTAudio/Llasa_opensource_speech_data_160k_hours_tokenized.gradio-dependents
Dataset Card for "gradio-dependents"
More Information needed
datasets-dependents
datasets metrics
This dataset contains metrics about the huggingface/datasets package.
Number of repositories in the dataset: 4997
Number of packages in the dataset: 215
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 22 packages that have more than 1000 stars.
There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.accelerate-dependents
accelerate metrics
This dataset contains metrics about the huggingface/accelerate package.
Number of repositories in the dataset: 727
Number of packages in the dataset: 37
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 10 packages that have more than 1000 stars.
There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.evaluate-dependents
evaluate metrics
This dataset contains metrics about the huggingface/evaluate package.
Number of repositories in the dataset: 106
Number of packages in the dataset: 3
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 1 packages that have more than 1000 stars.
There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.pytorch-image-models-dependents
pytorch-image-models metrics
This dataset contains metrics about the huggingface/pytorch-image-models package.
Number of repositories in the dataset: 3615
Number of packages in the dataset: 89
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 18 packages that have more than 1000… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/pytorch-image-models-dependents.diffusers-dependents
diffusers metrics
This dataset contains metrics about the huggingface/diffusers package.
Number of repositories in the dataset: 160
Number of packages in the dataset: 2
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.optimum-dependents
optimum metrics
This dataset contains metrics about the huggingface/optimum package.
Number of repositories in the dataset: 19
Number of packages in the dataset: 6
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.pip
Dataset Card for "pip"
More Information needed
qmmit-open-source-agent-commit-index
Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures
Dataset release: 2026-09-18-v3.0Schema: 3.0.0
Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16
Abstract
This dataset contains 2000 repository-level observations from public Git
repositories. Each observation estimates a lower bound on the proportion of
non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.open-source-scientific-documents
Open-Source Scientific Documents
This dataset contains approximately 100,000 open scientific PDF documents packaged as a shared retrieval corpus. Train, validation, and test query sets are expected to reference document_id values from this single corpus rather than using separate document splits.
Sources
ACL: 20,845 PDFs
Biology: 22,000 PDFs
Engineering: 22,000 PDFs
Medicine: 22,000 PDFs
Physics: 21,667 PDFs
The source folders preserve the project corpus… See the full description on the dataset page: https://huggingface.co/datasets/kasys/open-source-scientific-documents.starsissuessafetensors-dependents
Dataset Card for "safetensors-dependents"
More Information needed
rareburden-commons-open-source-snapshots
RareBurden Commons open source snapshots
This public preservation projection contains exact, hash-bound source files
whose observed terms affirmatively permit redistribution. Each source retains
its own licence; license: other is intentionally used because the collection
is not governed by one uniform licence.
Included:
Orphadata July 2026 alignment and epidemiology files — CC BY 4.0.
Exact MONDO release assets — CC BY 4.0. The currently receipt-bound history
covers v2026-08-04… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/rareburden-commons-open-source-snapshots.common-corpus-sample-open-sourceopen_parallel_think_code_source
open_parallel_think_code_source
A large-scale code reasoning distillation dataset with 320,000 solution trajectories generated by 4 state-of-the-art thinking models across 10,000 unique coding problems.
Source / raw pool. This is the per-trajectory dataset. The packed parallel-thinking datasets derived from it are haowu89/open_parallel_think_code_full (full reasoning + solution) and haowu89/open_parallel_think_code_cot (solution only). Each trajectory's metadata carries… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_code_source.peft-dependents
Dataset Card for "peft-dependents"
More Information needed
opensource_100_TVG_caseos-world-modifiedopen-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus
Dataset Summary
Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Catalan (ca)
English (en)
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.ai-opensource-2026
AI Open Source 2026
Open source AI releases, repos, communities. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c =… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-opensource-2026.im3_open_source_data_center_atlas_v2026.02.09
IM3 Open Source Data Center Atlas v2026.02.09 — refined database
This repository preserves the IM3 Open Source Data Center Atlas v2026.02.09 and
adds a source-enriched, audited 43-column power-source table for all 1,479
source geometry records (1,474 unique IM3 IDs). The publication retains the exact
13 upstream columns plus 30 stable label, interpretation, and evidence fields.
Duplicate geometry records are intentionally retained.
Files… See the full description on the dataset page: https://huggingface.co/datasets/sarkarghya/im3_open_source_data_center_atlas_v2026.02.09.color-classification-opensourceissues-externalFOURSQUARE_OPEN_SOURCE_PLACES_x_OVERTURE_INNER_JOIN
Placekey Inner Join of Overture (2024-11-13.0) and Foursquare Open Source Places (2024-11-19)
Summary of the Join
Matched Placekeys inner: 7,868,660
Foursquare Open Source Places Total Placekeys: 103,413,964
Overture Total Placekeys: 42,084,985
🚀 Foursquare Uplift
Foursquare Open Source Places (Apache 2.0) docs contributed the following data enhancements when joined with Overture:
Field
Values Added
geometry
7,868,660
confidence
7,868,660… See the full description on the dataset page: https://huggingface.co/datasets/Placekey/FOURSQUARE_OPEN_SOURCE_PLACES_x_OVERTURE_INNER_JOIN.
