datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataset-with-standalone-yamlThis is a test dataset used in the datasets library CI
yamlab_datasetsthe-stack-yaml-k8s
Dataset Card for The Stack YAML K8s
This dataset is a subset of The Stack dataset data/yaml. The YAML files were
parsed and filtered out all valid K8s YAML files which is what this data is about.
The dataset contains 276520 valid K8s YAML files. The dataset was created by running
the the-stack-yaml-k8s.ipynb
Notebook on K8s using substratus.ai
Source code used to generate dataset: https://github.com/substratusai/the-stack-yaml-k8s
Need some help? Questions? Join our Discord server:… See the full description on the dataset page: https://huggingface.co/datasets/substratusai/the-stack-yaml-k8s.yam_lego_taxiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "yam_bimanual",
"total_episodes": 347,
"total_frames": 937993,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:347"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/yam_lego_taxi.yam_lego_taxi_s200_rlt_annotrl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-php-larggorilla_openfunctions_yaml_trainrl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_stac-self-largYAM_lerobot_formatThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "yam",
"total_episodes": 50,
"total_frames": 101695,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/leokswang/YAM_lerobot_format.rl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt_trai-data_exp_rpt_pyme-largrl_rl-conf_24GP_base-yaml_mode-path_r2eg-nl2b-stac-bugs-fixt-agai_trai-data_exp_rpt_stac-rustxinworen-yaml-mode-switch
xinworen-yaml-mode-switch
免重启热切换 YAML 配置 · Hot-toggle YAML config without restart
bash + sed · 零外部依赖 · 幂等 · 可回滚
由 XinWoRen 出品 · 面向企业级部署的通用「配置热开关」工程底座 ·
github.com/XinWoRen-Global · 项目官网 xinworen.com
本工具是一个通用、安全、幂等的配置热开关 CLI,纯 bash + sed 实现、零依赖,可挂载到任意运行时作为「配置热加载」的入口。不绑定任何平台、模型或密钥 — 从生产环境沉淀、脱敏后开源。
A generic, dependency-free shell CLI to hot-toggle boolean/value keys in a YAML config. Extracted from production and open-sourced without any… See the full description on the dataset page: https://huggingface.co/datasets/XinWoRen/xinworen-yaml-mode-switch.rl_rl-conf_24GP_base_noth-yaml_mode-path_r2eg-nl2b-stac-bugs_trai-data_exp_rpt_stac-self-largrl_rl-conf_24GP_base_noth-yaml_mode-path_r2eg-nl2b-stac-bugs_trai-data_exp_rpt_pyme-largterminal_bench_2_rl_rl_conf_qwen_8b_ll_lr1e_5_bs64_yaml_mode_path_r2eg_nl2b_sta63ce4c6aCloudEval-YAMLterminal_bench_2_rl_rl_conf_qwen_8b_ll_lr1e_5_bs64_yaml_mode_path_r2eg_nl2b_stab3b297f3structeval-t-sft-hq-yaml-cleaned
StructEval-T SFT HQ YAML (Cleaned)
このデータセットは、daichira/structeval-t-sft-hq-yaml をベースに、厳密なフォーマット検証とノイズ除去(クリーニング)を行ったものです。
StructEval-T等の構造化データ生成タスク(SFT向け)に最適化されています。
クリーニング統計情報
本データセットの構築時に、以下のクリーニング結果が得られました。
オリジナルレコード数: 2000 件
クリーニング後レコード数: 2000 件
除去されたCoTノイズ: 1528 件 ( Approach: ... Output: を物理的に切除 )
削減された無駄な文字列の総量: 705728 文字
最終YAMLパース成功率: 100%
データセット構築パイプライン(クリーニング手法)
不要テキストの物理的除去: Approach: ... Output: といった思考プロセスや、マークダウンのコードフェンス (```yaml) を正規表現で完全に削除しました。… See the full description on the dataset page: https://huggingface.co/datasets/daichira/structeval-t-sft-hq-yaml-cleaned.marin-starcoderdata_yamlk8s-yaml
k8s-yaml
Kubernetes resource manifests in YAML across many resource kinds. One example per manifest file; the filename prefix encodes the resource kind (e.g. deployment_001750.yaml).
A corpus of 246,439 configuration files, packaged as a single parquet split for
convenient loading. Part of a collection of schema/config corpora.
Format
One row per source file. Columns:
column
type
description
filename
string
original file name (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-tct-authors/k8s-yaml.synthetic-parsed-names-yaml
Dataset Card for Synthetic Parsed Names (YAML)
This dataset contains approximately 500,000 synthetic examples of complex, unstructured historical names paired with their structured YAML equivalents. It is designed to fine-tune small open-source large language models (LLMs) to accurately parse cultural heritage name strings into isolated components (first names, last names, middle names, dates, titles, etc.) for de-duplication and structured data ingestion.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/synthetic-parsed-names-yaml.rl_rl-conf_24GP_base_noth-yaml_mode-path_r2eg-nl2b-stac-bugs_trai-data_exp_rpt_unit-pyth-largdoc-yaml-3
[doc] manual configuration 3
This dataset contains two csv files in the data/ directory and one csv file in the holdout/ directory, and a YAML field configs that specifies the data files and splits, using glob expressions.
doc-yaml-4
[doc] manual configuration 4
This dataset contains two csv files at the root, and a YAML field configs that specifies the data files and configs.
terminal_bench_2_rl_rl_conf_20GP_base_yaml_mode_path_r2eg_nl2b_stac_bugs_fixt_tb67d097ddoc-yaml-2
[doc] manual configuration 2
This dataset contains two csv files in the data/ directory and one csv file in the holdout/ directory, and a YAML field configs that specifies the data files and splits.
terminal_bench_2_rl_rl_conf_20GP_base_yaml_mode_path_r2eg_nl2b_stac_bugs_fixt_t14fd5792terminal_bench_2_rl_rl_conf_24GP_base_noth_yaml_mode_path_r2eg_nl2b_stac_bugs_t32920c19test_yaml
Dataset Card for "test_yaml"
More Information needed
environment-yaml-sorted
