datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-source-english-catalan-corpus
Dataset Card for open-source-english-catalan-corpus
Dataset Summary
Translation memory built from more than 180 open source projects. These include LibreOffice, Mozilla, KDE, GNOME, GIMP, Inkscape and many others. It can be used as translation memory or as training corpus for neural translators.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Catalan (ca)
English (en)
Dataset Structure
Data Instances
[More… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/open-source-english-catalan-corpus.open_parallel_think_source
Open Parallel Think — Source (per-model subsets)
Math reasoning traces distilled from a shared question set by four models, organized
one subset (config) per source model. Each question carries multiple reasoning traces
("parallel think"); here those traces are partitioned by the model that produced them.
The underlying questions come from three collections: openmathinstruct, numinamath,
and deepscale (the source is the prefix of guid, e.g. deepscale_10003).
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/open_parallel_think_source.open-source-ai-models-dataset
OpenModelMap — The Largest Open-Source AI Models Dataset (Chinese + English)
2,484 models · 35 fields · 9 sources · Updated daily
This dataset provides the most comprehensive structured metadata for open-source AI models, with a focus on Chinese model coverage. Every model includes benchmark scores, hardware requirements, GPU compatibility, license information, and deployment methods.
What's Inside
Field
Description
id
HuggingFace model ID
name… See the full description on the dataset page: https://huggingface.co/datasets/duola15/open-source-ai-models-dataset.gingiris-opensource
⭐ Open Source Launch Marketing Playbook 2026
Go from 0 to 10K GitHub stars in 43 days — the exact OSS launch sequence behind AFFiNE's 0→60K stars: Show HN timing, Reddit seeding, 48-hour star spike, Chinese developer community distribution (Gitee, 掘金, V2EX), and sustained content velocity.
📦 Install
npx skills add Gingiris-1031/gingiris-opensource
Then ask your AI agent:
"I just open-sourced my AI agent framework, plan my 30-day push to 1K GitHub stars" ·… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/gingiris-opensource.open-source-marketing-playbook
Open Source Marketing Playbook
Marketing playbook for open-source projects led by non-technical founders. Covers README optimization, community building, contributor attraction, and transl...
📦 Install on ClawHub
clawhub install open-source-marketing-playbook
Then ask your AI agent:
"I just open-sourced my AI tool. Get to 1k GitHub stars in 30 days"
Installs the full Open Source Marketing Playbook playbook — battle-tested with 30+ Product Hunt #1 wins… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/open-source-marketing-playbook.
