datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tokenizers-dependents
tokenizers metrics
This dataset contains metrics about the huggingface/tokenizers package.
Number of repositories in the dataset: 11460
Number of packages in the dataset: 124
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 14 packages that have more than 1000 stars.
There are 41… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/tokenizers-dependents.wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.transformers-dependents
transformers metrics
This dataset contains metrics about the huggingface/transformers package.
Number of repositories in the dataset: 27067
Number of packages in the dataset: 823
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 65 packages that have more than 1000 stars.
There are 140… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/transformers-dependents.gradio-dependents
Dataset Card for "gradio-dependents"
More Information needed
datasets-dependents
datasets metrics
This dataset contains metrics about the huggingface/datasets package.
Number of repositories in the dataset: 4997
Number of packages in the dataset: 215
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 22 packages that have more than 1000 stars.
There are 43… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/datasets-dependents.accelerate-dependents
accelerate metrics
This dataset contains metrics about the huggingface/accelerate package.
Number of repositories in the dataset: 727
Number of packages in the dataset: 37
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 10 packages that have more than 1000 stars.
There are 16… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/accelerate-dependents.evaluate-dependents
evaluate metrics
This dataset contains metrics about the huggingface/evaluate package.
Number of repositories in the dataset: 106
Number of packages in the dataset: 3
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 1 packages that have more than 1000 stars.
There are 2 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/evaluate-dependents.diffusers-dependents
diffusers metrics
This dataset contains metrics about the huggingface/diffusers package.
Number of repositories in the dataset: 160
Number of packages in the dataset: 2
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 3 repositories… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/diffusers-dependents.optimum-dependents
optimum metrics
This dataset contains metrics about the huggingface/optimum package.
Number of repositories in the dataset: 19
Number of packages in the dataset: 6
Package dependents
This contains the data available in the used-by
tab on GitHub.
Package & Repository star count
This section shows the package and repository star count, individually.
Package
Repository
There are 0 packages that have more than 1000 stars.
There are 0 repositories that… See the full description on the dataset page: https://huggingface.co/datasets/open-source-metrics/optimum-dependents.pip
Dataset Card for "pip"
More Information needed
starsissuessafetensors-dependents
Dataset Card for "safetensors-dependents"
More Information needed
peft-dependents
Dataset Card for "peft-dependents"
More Information needed
open-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus
Dataset Summary
Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Catalan (ca)
English (en)
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.ai-opensource-2026
AI Open Source 2026
Open source AI releases, repos, communities. Updated daily via automated collection pipeline.
Part of the Legion Data Factory — historical AI ecosystem datasets 2026.
Methodology
Automated collection from public sources (HackerNews, RSS feeds, APIs).
Updated daily via cron job. Raw data, minimal processing.
License
CC BY 4.0
📦 Install
pip install legion-intel
from legion_intel import LegionClient
c =… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-opensource-2026.issues-externaltoken-classification-checkpoint-downloadsEvaluation_of_OpenSource_Models_for_PDF_Injection_Recognition
Injected PDFs - Model Evaluation
This repository holds the model evaluation stage of a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the artefacts it produced for the
application.
Nothing is trained here. Seven off-the-shelf models are measured against the same 1,100 PDFs,
and the two winners are exported for the app to load.
Question
Candidates
Winner
Part A
Which files look like this one?
3 embedding models x 2 inputs… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/Evaluation_of_OpenSource_Models_for_PDF_Injection_Recognition.preprocessed_issues
Dataset Card for "preprocessed_issues"
More Information needed
fake_news_en_opensources
Dataset Card for "Fake News Opensources"
Dataset Description
Homepage: https://github.com/AndyTheFactory/FakeNewsDataset
Repository: https://github.com/AndyTheFactory/FakeNewsDataset
Point of Contact: Andrei Paraschiv
Dataset Summary
a consolidated and cleaned up version of the opensources Fake News dataset
Fake News Corpus comprises 8,529,090 individual articles, classified into 12 classes: reliable, unreliable, political, bias, fake, conspiracy… See the full description on the dataset page: https://huggingface.co/datasets/andyP/fake_news_en_opensources.pip-external
Dataset Card for "pip-external"
More Information needed
feature-extraction-checkpoint-downloadspreprocessed_stars
Dataset Card for "preprocessed_stars"
More Information needed
reinforcement-learning-checkpoint-downloadsfill-mask-checkpoint-downloadsvisual-question-answering-checkpoint-downloadsunconditional-image-generation-checkpoint-downloadsBioLaySumm2025-LaymanRRG-opensource-trackpreprocessed_pip
Dataset Card for "preprocessed_pip"
More Information needed
