datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
soft-trigger-verifiedwikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.Soft-tissue-Sarcoma
Soft-tissue-Sarcoma (STS)
A TCIA collection of 51 patients with histologically proven soft-tissue sarcoma
of the extremities, each imaged with joint pre-treatment FDG-PET/CT and MRI
and contoured by an expert radiation oncologist. Collected at McGill University
Health Centre (Montreal) and published with Vallières et al., Phys Med Biol 2015.
The original study built a radiomics model predicting lung metastases from
joint PET/MRI texture features; 19 of the 51 patients developed… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/Soft-tissue-Sarcoma.harmless_alpaca_jaJapanese auto-translation of mlabonne/harmless_alpacausing llmfan46/gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4-GGUF
SoftDis
SoftDis dataset
SoftDis is a dataset for the exploration of disordered regions in
protein structures, and their relations with interacting sites.
The concept of soft disorder was introduced in (Seoane and Carbone, 2021),
as a general term for regions in a protein identified as flexible
(characterized by high B-factor) or intermittently missing across different
X-ray crystal structures of the same sequence. The definition is derived from
an extensive analysis of clusters of… See the full description on the dataset page: https://huggingface.co/datasets/CQSB/SoftDis.CMPR_LTS_SOFT_TOME_ADJ_CTDHarmful-Harmless-100Pairs-JA-HighIntensity
Harmful-Harmless-100Pairs-JA-HighIntensity
This is a small-scale dataset consisting of 100 pairs of high-intensity Harmful / Harmless contrastive data written in Japanese.
⚠️ Important Notice
This dataset intentionally contains harmful, explicit, offensive, disturbing, biased, or otherwise inappropriate content for research and evaluation purposes. Some entries may describe dangerous, illegal, abusive, or unethical activities in substantial detail.
The inclusion… See the full description on the dataset page: https://huggingface.co/datasets/OS-Software/Harmful-Harmless-100Pairs-JA-HighIntensity.chess-soft-sf19
avewright/chess-soft-sf19
Official Stockfish 19 MultiPV soft targets. This release supersedes the 25k
pilot. It is not a filter of chess-soft-multipv-lichess or
chess-soft-100m-disagreements.
2,010,006 rows. Source id 4. Vocab compact (1968).
Mix (as generated)
origin
rows
note
self-play (origin=1)
0
SF19 vs SF19, ε=0.20, book + 4 random legal
relabel (origin=0)
0
existing local boards, new SF19 labels
frozen eval
10,000
split=1 in… See the full description on the dataset page: https://huggingface.co/datasets/avewright/chess-soft-sf19.Python-Code-Solutions
Python Code Solutions
Features
1000k of Python Code Solutions for Text Generation and Question Answering
Python Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
training-embeddingssoft-prompt-experiments-archive-20260918
Soft prompt 实验归档
用于查阅和恢复的历史研究记录,涵盖数学与代码任务。共 45 个运行目录,包含教师生成数据、评测输出、原始配置和已有 prompt 检查点。部分目录仅有评测、复核或失败记录,不能将目录数量理解为成功实验数量。
快速查阅
实验总览:模型系列、任务、规模与记录状态。
CSV 索引 / JSON 索引:便于筛选和定位。
archives/:按实验分别压缩的原始文件。
manifests/:各文件 SHA-256 与归档路径。
系列包括 AReaL Boba2、GPT-OSS/Swallow、MiMo、X-Coder/Qwen3、Nemotron、Klear、Mellum2、OLMo3、Polaris 和 Poro2。页面不展开具体方法或实现细节;原始配置仍保留在归档内供恢复。
状态与注意事项
上传完成以 ARCHIVE_COMPLETE.json 为准;文件不存在时表示仍在上传。 每个归档都经过完整下载的 SHA-256 校验。… See the full description on the dataset page: https://huggingface.co/datasets/namezz/soft-prompt-experiments-archive-20260918.software_requirementsSoftware-Engineering-Dataset_90_10stockfish-19-soft-targets
avewright/stockfish-19-soft-targets
Official Stockfish 19 MultiPV soft targets, mined from the Lichess ECO
opening set. In-progress snapshot toward 1M unique positions.
600,000 rows in this upload. Source id 4. Vocab compact (1968).
How positions are chosen
Games start from the Lichess Chess Openings dataset
(lichess-org/chess-openings): 3,810 named
leaves (HF card still lists 3,704) plus
book prefixes, 7,852 unique starts.
ECO volumes: A 817 / B 772 /
C 1,250 / D… See the full description on the dataset page: https://huggingface.co/datasets/avewright/stockfish-19-soft-targets.arxiv-softwares-2021software_slackschess-soft-multipv-lichess
avewright/chess-soft-multipv-lichess
Soft MultiPV policy targets for chess transformers (compact move vocab).
Built from Lichess cloud evaluations + local harvests. Each row is a position with
an 8-wide soft move distribution (soft_indices / soft_probs) plus hard best move.
Fields
board_array (64): piece encoding
turn, castling, ep_square
move_idx, cp, mate
soft_indices[8], soft_probs[8]
label_depth, phase, source, cache_name
catalan-dictionary
Dataset Card for ca-text-corpus
Descripció (ca)
En aquest repositori s'apleguen llistes de paraules etiquetades amb la categoria gramatical, usades per a construir eines com correctors ortogràfics i gramaticals.
Dataset Summary
Catalan word lists with part of speech labeling curated by humans. Contains 1 180 773 forms including verbs, nouns, adjectives, names or toponyms. These word lists are used to build applications like Catalan spellcheckers or… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/catalan-dictionary.United_States_State_Legislation_with_SummariesTest Push
datapoints_round1_dpsk_software_engineering_shard1_daytona_n100k1OmniGIRLThis repository contains the data presented in OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution. OmniGIRL is a GitHub issue resolution benchmark that is multilingual, multimodal, and multi-domain. It includes 959 task instances collected from repositories across four programming languages (Python, JavaScript, TypeScript, and Java) and eight different domains.
italic-softkd-pool
italic-softkd-pool
The exact training data of idealab-cs2/zagreus-0.4B-italic-softkd: 21,606 Italian multiple-choice questions with committee soft labels. One soft-KD training run from mii-llm/zagreus-0.4B-ita on the train split reaches 0.4787 on the full ITALIC 10K (official harness, 5-shot fast, temperature 0), from a 0.2802 base.
train is the full pool; the other three splits partition it by provenance:
split
rows
contents
train
21,606
the full training file (union… See the full description on the dataset page: https://huggingface.co/datasets/idealab-cs2/italic-softkd-pool.datapoints_round1_dpsk_software_engineering_shard2_daytona_n100k1khazri-corpus
Khazri Corpus
A clean, high-quality Azerbaijani text corpus built from 570+ books.
570+ kitabdan hazırlanmış təmiz, yüksək keyfiyyətli Azərbaycan dili mətn korpusu.
English · Azərbaycanca
Language / Dil
Azerbaijani (az)
Total rows / Ümumi sətir sayı
~840K
Source books / Mənbə kitablar
570+ (300+ fiction · 200+ scientific · 70+ political)
Format
Parquet / JSONL, single text field
License / Lisenziya
CC BY-NC-ND 4.0
Used to train / Təlimdə istifadə olunub… See the full description on the dataset page: https://huggingface.co/datasets/softyugroup/khazri-corpus.us-k12-schools-directory
US K-12 Schools Directory
A directory of 124,613 US K-12 schools covering all 50 states, DC, and US territories,
compiled from federal and state government sources. Each record carries directory
information (address, phone, website), enrollment and demographics, and, where a source
supplied it, a principal name and email.
This is a compilation of public government data. It is not a survey, and no field was
independently verified against the school itself.
Loading… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/us-k12-schools-directory.2023-1000-software-release-notessoftware-strategist-v1
Software Fundamentals — Strategy Knowledge Base
A language-agnostic knowledge base of software engineering fundamentals, paired with a synthetic instruction-tuning dataset (~13,500 examples) for training small language models (SLMs) as software engineering strategists.
The trained model takes a description of a coding situation and routes it to relevant concepts, outputting synthesized strategic guidance as structured JSON.
Dataset Summary
This dataset provides ~13… See the full description on the dataset page: https://huggingface.co/datasets/jtregunna/software-strategist-v1.arxiv-software-repo-links
arXiv Software Repository Links
A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering.
Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models
Quick Start
from datasets import load_dataset
# Load DOI-to-repo links
links = load_dataset("cometadata/arxiv-software-repo-links", "links")
#… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.k12-standards-instruction-tasks
K-12 Curriculum Tasks (generated)
2,489 generated instruction/input/output records covering five curriculum tasks:
assessment creation, learning objective generation, misconception detection, standard
explanation, and standards Q&A. Content is predominantly mathematics.
Important: the name is misleading
Despite the name, this dataset contains no school directory data. There are four
columns - task, input, output, metadata - and no staff, principal, or school… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-standards-instruction-tasks.olmocr_science_pdfs-software_developmenthttps://huggingface.co/datasets/allenai/dolma3_pool/tree/main/data/olmocr_science_pdfs-software_development
