datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
employee-burnout-turnover-prediction-800k
Synthetic Employee Dataset
800,000+ employee records with real-world distributions for burnout prediction, turnover analysis, and HR analytics
What You Get
This isn't just another CSV dump. You're looking at 800K+ carefully engineered employee profilesthat mirror actual workforce dynamics, complete with performance metrics, burnout indicators, skill matrices, and behavioral personas. Think of it as a production-ready HR database that never existed but feels like it… See the full description on the dataset page: https://huggingface.co/datasets/Mystic777/employee-burnout-turnover-prediction-800k.sockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "lekiwi_client",
"total_episodes": 1,
"total_frames": 1793,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MysticShirou/sock.Mystery_Arena_Results
MysteryArena Results
This dataset stores public MysteryArena run summaries, match records, and compressed episode trajectories for the read-only Streamlit frontend.
The frontend reads index/runs.json first, then loads the unified matches/all_matches.jsonl.gz split. Per-run summaries and compressed trajectory files are kept for metadata and replay.
identifiersmath-hard-traces
Dataset Card for "math-hard-traces"
More Information needed
nemotron-math-v2-truly-hard-tool
Nemotron-Math-v2 Truly Hard Tool Subset
Problems where ALL 6 reasoning regimes score <= 3/8 (37.5%). These are the hardest problems in Nemotron-Math-v2.
Splits
Split
Accuracy
Before (dedup)
After (truly hard)
pass0of8
0/8
2
2
pass1of8
1/8
1,195
689
pass2of8
2/8
652
168
pass3of8
3/8
875
107
Total
2,724
966
Filter
All 6 regimes must have accuracy <= 0.375:
reason_high_with_tool / reason_high_no_tool
reason_medium_with_tool /… See the full description on the dataset page: https://huggingface.co/datasets/akh-mysterio/nemotron-math-v2-truly-hard-tool.example_dataset
example_dataset
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
nemotron-math-v2-hard-tool-dedup
Nemotron-Math-v2 Hard Tool-Use Subset (Deduplicated)
Filtered and deduplicated from nvidia/Nemotron-Math-v2.
Dedup stats
Split
Before
After
Removed
0_8
8
2
6
1_8
1,205
1,195
10
2_8
1,164
652
512
3_8
2,206
875
1,331
Total
4,583
2,724
1,859
Deduplication: dropped duplicate problem texts within each split (kept first occurrence).
nemotron-math-v2-hard-tool
Nemotron-Math-v2 Hard Tool-Use Subset
Filtered from nvidia/Nemotron-Math-v2 — only the hardest problems with tool-use trajectories.
Filters applied
Source: only high reasoning effort splits
changed_answer_to_majority == False
metadata.reason_high_with_tool.accuracy <= 0.375 (model gets <= 3/8 correct)
tools field non-empty (Python TIR tool definition present)
messages contain actual tool_calls (model used the tool)
expected_answer is non-null
Splits (by pass… See the full description on the dataset page: https://huggingface.co/datasets/akh-mysterio/nemotron-math-v2-hard-tool.mysterynailong_grabmy_storageClueMeister_EvaluatedClues_The_DArblay_Mystery_cluesClueMeister_EvaluatedClues_The_DArblay_Mystery_clustersMy_stock_dataset
