datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Expert-Sudoku-100kChinese-Qwen3-235B-Thinking-2507-Distill-100k
📌 Note: The English translation of this dataset card is provided below.
Chinese-Qwen3-235B-Thinking-2507-Distill-100k
Dataset Summary
Chinese-Qwen3-235B-Thinking-2507-Distill-100k 是一个包含约 100k 条高质量中文推理与指令数据的数据集,由 Qwen-3-235B-A22B-Thinking-2507(官方 Thinking 模式,上下文长度 32K)蒸馏生成。
该数据集覆盖了多个重要领域:
数学与工程任务(Mathematics, Applied Math, Advanced Math)
通用知识与写作(General Knowledge, Language & Writing)
技术与编程(Technology & Programming)
商业与经济(Business & Economics)… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k.pump-fun-sentiment-100k
Pump.fun Token Sentiment & Risk Analysis
AI-agent-generated sentiment analysis and quantitative risk labels for Solana memecoins on Pump.fun, collected via Pump Studio.
Dataset Description
Each row is a validated analysis submission from an AI agent operating on the Pump Studio platform. Agents observe real-time token data (price, market cap, holders, volume, bonding curve) and produce:
Sentiment label — bullish / bearish / neutral with 0-100 confidence score
Risk… See the full description on the dataset page: https://huggingface.co/datasets/Pumpdotstudio/pump-fun-sentiment-100k.inzynierka-synth40-100k Architektura synth40. tu jest 40
parametrów i inny tryb cech pann_mfcc_hilbert, embedding PANN/Cnn14 zamiast
log-mel. Primary key to midi_48
zamiast midi_60 jak w synth37.
inzynierka-synth40-100k
100 004 presetów syntezatora synth40, stokenizowane pod trening Transformer-VAE.
Rekordy leżą na gładkich i spójnych brzmieniowo trajektoriach w przestrzeni brzmienia, cel modelu to przestrzeń latentna,
po której da się organicznie ewoluować brzmienie z seeda w duchu Synplant2 od… See the full description on the dataset page: https://huggingface.co/datasets/Amourman/inzynierka-synth40-100k.Persona-100k
🧠 Synthetic Persona Dataset (100K Personas)
📦 Overview
The Synthetic Persona Dataset is a large-scale, open-source dataset containing 100,000 uniquely generated personas. Each persona is structured as a JSON object with over 70 fields, combining both narrative-style descriptions and structured data. These synthetic personas simulate realistic human profiles, supporting a wide range of AI training and evaluation use cases — including personalization, recommendation… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Persona-100k.ogbench-block-double-hermite-100k
ogbench-block-double-hermite-100k
Unofficial reproduction of the scripted policies described in https://seohong.me/blog/behavioral-cloning-mystery/
using random piecewise Hermite splines as the backbone, and randomized control points, grasping angles/directions,
contact points, motion speed, gripper yaw/roll/pitch, mistakes and retries, etc.
This might not be the exact setup used by the study, but I tried to infer the parameters from
"How exactly did you script the policies?"… See the full description on the dataset page: https://huggingface.co/datasets/Yassine/ogbench-block-double-hermite-100k.FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details
Dataset Card for Evaluation run of FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
Dataset automatically created during the evaluation run of model FlofloB/100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/FlofloB__100k_fineweb_continued_pretraining_Qwen2.5-0.5B-Instruct_Unsloth_merged_16bit-details.unifiedor-100k
UnifiedOR-100K
Unified Operations Research foundation dataset combining eight heterogeneous OR benchmarks into a single schema with multi-layer representations.
Source Datasets
Source
Hub Reference
OR Layer
FrontierCO
alirezaaminzadeh/frontierco-instance-features
Combinatorial optimization + solver performance
Text2Opt-Bench
alirezaaminzadeh/opticoder-binding-cases
NL → MILP binding
OptMATH
nvidia/OptiMATH-Train
Math word problems
Learn2Zinc… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/unifiedor-100k.traj_sft_bench_100k100k-finewebedu-samples-8ktmt-dialogues-100k-v2
Multi-turn medical dialogues — V2 (pass@k)
39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup
as V1, but
generated with pass@k=4: a case is regenerated from scratch, graded with
Inspect's model_graded_fact, until the conclusion is graded correct or 4
attempts are spent. 30,234 dialogues (76%) end on a graded-correct
conclusion, roughly 2 attempts per case on average thanks to stopping as soon
as one succeeds.
Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.100k-finewebedu-samples4096tTotal tokens in matching entries: 275_639_417
Average tokens per entry: 2756.39
fineweb-portuguese-100k
FineWeb2 Portuguese 100k - Safety Classified
A 100,000-sample subset of FineWeb2 Portuguese web text, classified for content safety using Cohere Command A.
Dataset Description
Each record contains the original FineWeb2 text and metadata, plus a classification field with:
Field
Description
safety_rating
"safe" or "unsafe"
category
List of applicable harm categories (null if safe)
reason
Brief explanation of the classification
Safety Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/safety-aya/fineweb-portuguese-100k.100k-finewebedu-samples2048tTotal tokens in matching entries: 140_782_625
Average tokens per entry: 1407.83
judgelm-train-100k-dpoinzynierka-transformer-100k
inzynierka-transformer-100k
100 000 presetów syntezatora, stokenizowane pod trening Transformer-VAE.
Rekordy leżą na gładkich i spójnych brzmieniowo trajektoriach w przestrzeni brzmienia, cel modelu to przestrzeń latentna,
po której da się organicznie ewoluować brzmienie z seeda w duchu Synplant2 od Sonic Charge.
Każdy fragment kodu który jest w tym README jest od codexa, żeby nakierować na poprawne uzycie tego repo, bo jest zagmatwane troche
tl;dr
100 000… See the full description on the dataset page: https://huggingface.co/datasets/Amourman/inzynierka-transformer-100k.100k-finewebedu-samples256tTotal tokens in matching entries: 19672962
Average tokens per entry: 196.73
wildchat-100k-qwen
WildChat 100k Qwen cleaned
Danish WildChat prompt generations with a cleaned response set. This revision merges regenerated responses for high-refusal target rows, removes high-precision unwanted refusal rows, drops extreme over-length samples, and removes a detected system-prompt leak row.
The dataset keeps the same row schema as the previous synquid/wildchat-100k-qwen upload.
Cleaning summary:
Source rows: 99,983
Kept rows: 99,688
Replaced responses: 3,683
Dropped rows: 295… See the full description on the dataset page: https://huggingface.co/datasets/synquid/wildchat-100k-qwen.wildchat-100k-qwen-messages
WildChat Qwen 100k Messages
Messages-format version of the cleaned WildChat Qwen dataset.
Each row has a messages list:
69688 rows have the original two-message conversation: user, assistant.
30000 rows additionally include a simulated follow-up user turn and a generated assistant response: user, assistant, user, assistant.
Total rows: 99688.
The system prompt used for assistant generation is not included in the dataset.
See metadata/generation_report.json for generation metadata.
KCC-Sample-Dataset-100klaion2b_100k_religionFastMath_100K
Это учебная версия датасета.
Другие версии датасета вы можете найти в коллекции на HBB-Community
100k_labeleddanish-wildchat-100k
Danish WildChat 100k
Deduplicated Danish prompt sample from danish-foundation-models/danish-wildchat4.8M.
The train split contains 100,000 Danish initial user prompts from translated_first_user_message.
The sample keeps short and messy WildChat-style prompts, and removes exact and conservative near duplicates.
See metadata/dedupe_report.json for run statistics.
gen_select_100k_PT100k-finewebedu-samples512tvietquill-qcpg-100k-synthesis-questionvietquill-qcpg-100k-synthesis-sentence
